mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 20:34:57 +02:00
67616dd7ea725de7818c735af4cd8e807ec633de
3150
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
67616dd7ea |
Rework secrets dialog and add in-sheet agent access (#9797)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents and the credentials those agents need for work. > - The secrets UI is responsible for making credential creation understandable and for showing where credentials are used. > - The create flow previously exposed an editable generated key too early, used an unnatural field order, and rendered uneven provider tab rows. > - The detail sheet also lacked an in-context way to see and manage which agents reference a selected secret. > - This pull request makes the create dialog follow a predictable name-to-value flow and adds agent access management directly to the secret details sheet. > - The benefit is a clearer secrets workflow with fewer accidental key edits and less navigation when granting or revoking agent access. ## Linked Issues or Issue Description **Subsystem affected:** `ui/` — React + Vite board UI **Problem or motivation:** Creating a secret currently makes its generated key look immediately editable, places fields in an awkward keyboard order, and applies provider-tab overrides that produce uneven rows. After creation, operators cannot inspect or change agent access from the selected secret's detail sheet. **Proposed solution:** Generate the key from the path-style name and keep it read-only until an explicit Edit action; place Value directly after Name; use the standard tab sizing; and add an Agent access section that reads and updates `secret_ref` / `user_secret_ref` environment bindings through agent adapter configuration. **Alternatives considered:** Keeping access management only on agent configuration screens was rejected because it hides a secret-centric question—“which agents can use this?”—and requires repetitive navigation. Keeping the key always editable was rejected because the generated value should be the safe default. **Roadmap alignment:** No matching item was found in `ROADMAP.md`; this is a focused usability and access-management improvement to the existing secrets surface. **Additional context:** GitHub search found no duplicate PR for this dialog and in-sheet access change. PR #9321 also mentions user-secret resolution but addresses unrelated skills-route behavior. ## What Changed - Auto-generate the create-secret key from Name, keep it read-only by default, and expose an explicit Edit action. - Use a path-style Name placeholder (`/dev/foo/bar`) and place Value immediately after Name for natural keyboard navigation. - Remove tab sizing/whitespace overrides that caused uneven provider-tab row heights. - Add an Agent access section to the Details tab that lists referencing agents and grants or revokes `secret_ref` / `user_secret_ref` environment bindings in place. - Cover company secrets and each-user definitions with focused render tests. ## Verification - `pnpm exec vitest run ui/src/pages/Secrets.render.test.tsx` — 17/17 passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — changed files are clean; the repository-wide command currently reports five pre-existing `#9627` violations in unrelated files (`Sidebar.tsx`, `Inbox.tsx`, `IssueDetail.tsx`, `Issues.tsx`, and `Routines.tsx`). - Additional secrets verification completed during implementation: 35/35 adjacent secrets tests passed, and the flow was checked in a real browser in light and dark themes. ## Risks - Agent access mutations update adapter environment configuration, so malformed legacy env entries could affect how a binding is displayed or revoked. - Each-user definitions use `user_secret_ref` rather than `secret_ref`; focused tests cover selecting the correct binding type. - No database or API contract changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Fable 5 assisted with the implementation using repository-aware code editing and test execution. - OpenAI GPT-5.5 via Codex CLI assisted with PR preparation, branch hygiene, focused verification, GitHub operations, and review/check loops. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.718.0-canary.0 |
||
|
|
83765f08d1 |
fix(ui): reap React 19.2 performance-track measures (memory leak) (#9827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its web UI is a long-lived single-page app; users leave tabs open for hours/days > - We chased multi-GB tab memory growth through several fixes (#9569, #9624, #9627, #9701) that cut real churn — but the footprint kept climbing > - A heap-snapshot diff on a 12-hour tab finally showed the true cause: of 13.2M heap nodes, **12.1M were `PerformanceMeasure` objects** — React 19.2 emits a `performance.measure()` per component render for its DevTools "Performance Tracks" and never clears them > - These are native objects, so `performance.memory` never reported them (it read a flat ~74 MB while the real heap was ~308 MB), which is why our earlier heap/DOM sampling looked stable while the footprint ballooned > - This pull request periodically clears the User Timing measure buffer, since nothing in the app consumes it > - The benefit is that long-lived tabs stop accumulating millions of native `PerformanceMeasure` objects, eliminating the remaining unbounded growth ## Linked Issues or Issue Description No public GitHub issue exists; describing inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR", following the bug report template. This is the root-cause fix for the memory growth chased in #9569 / #9624 / #9627 / #9701. **What happened?** Long-lived browser tabs grew to multiple GB of memory footprint over hours/days. A heap snapshot of a 12h tab showed **13.2M nodes, of which 12.1M were `PerformanceMeasure`** (vs ~1.0M total nodes on a fresh tab). Live capture showed **~340 `performance.measure()` calls/sec**, named after React components with a `detail.devtools` payload — React 19.2's "Performance Tracks". Nothing ever clears them, so they accumulate without bound. **Expected behavior** A long-lived tab should not accumulate millions of `PerformanceMeasure` entries. Idle/long-running tabs should hold a bounded footprint. **Steps to reproduce** Open the app (React 19.2 production build), leave a tab open with normal activity, and run `performance.getEntriesByType('measure').length` periodically — it climbs unbounded (12.3M after ~12h). `performance.clearMeasures()` drops it to near-zero and frees the memory. **Paperclip version or commit** Branch `fix/react-perf-measure-leak`, off `master` (after #9701). **Deployment mode** Local dev (`pnpm dev` build served by the dev instance), web UI. React 19.2.7. Not adapter-specific. ## What Changed - **`ui/src/lib/perf-measure-reaper.ts` (new)** — `startPerfMeasureReaper(intervalMs = 10_000)` clears `performance.clearMeasures()` on a timer and returns a stop function. It never calls `getEntriesByType('measure')` (which would materialize the huge buffer). Honors a `window.__paperclipKeepPerfMeasures = true` opt-out so a developer recording a React Performance Track in DevTools can keep the entries. - **`ui/src/main.tsx`** — start the reaper at app boot. - Only *measures* are cleared — React's tracks pass explicit start/end times and leave no marks, and the app doesn't use the User Timing API at all (verified), so nothing else is affected. ## Verification - **Root cause proven** by heap-snapshot diff: fresh tab ~1.0M nodes vs 12h tab 13.2M nodes, 12.1M of them `PerformanceMeasure`; live capture showed ~340 measures/sec with `detail.devtools` and React component names. - **Fix proven live**: `performance.clearMeasures()` on the aged tab dropped the buffer from **12,418,266 → 1,500** and reclaimed the memory (heap `perf.memory` 308 → 269 MB, plus the ~12M native objects, which are the bulk of the footprint). - `vitest`: `perf-measure-reaper.test.ts` — interval clearing, `stop()`, opt-out flag, and no-API no-op. All pass. - `tsc -b` clean. ## Risks Very low, client-only. - Clearing the User Timing measure buffer only affects the DevTools Performance panel's React track *history*; normal users never consume it. Developers who want to record it can set `window.__paperclipKeepPerfMeasures = true`. - Only `performance.clearMeasures()` is called (not `clearMarks`), so any mark-based timing elsewhere is untouched; a grep confirmed the app makes no `performance.mark()/measure()` calls of its own. - Adds exactly one 10s interval (negligible), and `clearMeasures()` does not materialize the buffer. Note: this is a React 19.2 upstream behavior (its performance tracks are emitted in the production build and never cleared). If React later gates or clears them, this reaper can be removed. ## Model Used - **Provider:** Anthropic, via the Claude Code CLI. - **Model:** Claude Opus 4.8 (`claude-opus-4-8`). - **Reasoning mode:** Extended thinking enabled. - **Capabilities used:** tool use (shell, file editing), the Chrome DevTools MCP to reproduce and profile, and a streaming heap-snapshot parser to identify the `PerformanceMeasure` accumulation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (continues #9701; no duplicates) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [ ] I have updated relevant documentation to reflect my changes (N/A — no user-facing docs; rationale documented inline) - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
14da75dfc7 | feat(codex-local): outbound auth copy-back as home-asset restore contribution (#9788) | ||
|
|
936215693c |
feat(codex-local): add outbound direction to codex-auth-merge decision predicate (#9787)
## Thinking Path > - Paperclip runs Codex as a sandboxed agent with its own CODEX_HOME; at the start of each run Paperclip merges the sandbox's auth.json with the host's so the agent can authenticate to the Codex API > - The merge uses `codex-auth-merge-decision.cjs`, a small CJS script that reads two `auth.json` snapshots and exits `10` (use source copy) or `20` (keep destination copy) > - Until now the script only handled the **inbound** direction (host → sandbox) with hard-coded sandbox/host semantics > - Sandbox-to-host copy-back (**outbound**) requires the same freshness guard but with swapped roles — and exposing both directions via a flag requires hard-coding knowledge about which role is "sandbox" and which is "host" into the predicate itself > - A cleaner design: the predicate is **direction-agnostic** — it accepts a `source` path and a `destination` path; the caller decides direction by which argument it passes first. The script answers one question: "should destination be replaced by source?" Exit `10` = yes (use source); exit `20` = no (keep destination) > - This PR refactors `codex-auth-merge-decision.cjs` to this source/destination model, removes the `--direction` flag, and extends coverage with a 15-row source/destination predicate matrix plus updated extract integration tests ## Linked Issues or Issue Description No public GitHub issue exists. The underlying problem: **Context** — Paperclip merges `auth.json` files between sandbox and host to keep Codex credentials in sync across agent runs. The existing decision predicate (`codex-auth-merge-decision.cjs`) handled the inbound direction (host→sandbox) with hard-coded sandbox/host naming. A copy-back (outbound) direction was needed, but rather than add a `--direction` flag that hard-codes role knowledge into the predicate, the design was simplified: the predicate is now **direction-agnostic**, taking `source` and `destination` paths and answering "replace destination with source?" The freshness guard is unified: exit `10` only when source and destination share the same `account_id`, same subscription-kind, both carry a parseable `last_refresh`, and `source.last_refresh > destination.last_refresh` (strict). Ties and null freshness keep the destination copy so a spent single-use refresh token (openai/codex#15502) is never written back over a valid one. Related PRs: - Refs #9785 — Phase 2: relocated the decision predicate from `adapter-utils` into `codex-local`; this PR builds on that merged change - Refs #9621 — community PR targeting Codex sandbox auth sync-back via a different approach; operates on different files (runtime layer vs. decision script), non-conflicting ## What Changed - `packages/adapters/codex-local/src/server/codex-auth-merge-decision.cjs` — refactored from direction-parameterized (sandbox/host with `--direction`) to **direction-agnostic source/destination** semantics. Args: first = source path, second = destination path. Exit `10` (use source) only when: same `account_id`, same subscription-kind, both carry a parseable `last_refresh`, and `source.last_refresh > destination.last_refresh` (strict). Every other case exits `20` (keep destination). Unknown/extra flags are still rejected with exit `1`. No token bytes are emitted to stdout/stderr. - `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` — rewritten to source/destination API: 15-row predicate matrix (strictly-newer→`10`; tie/older/identity-mismatch/null-refresh/apikey/kind-mismatch/unusable→`20`) plus token-bytes non-emission guard. The inbound extract integration suite updated to match. - Runtime callers (`codex-auth-merge-extract.sh`) already invoke `node decision <sandbox> <host>` = `(source, destination)` and branch on exit `10` → keep sandbox, so **no caller changes are needed**. **Behavioral note:** unifying the two directions under a single "use source only when strictly fresher" rule means a tie (exact equal `last_refresh`) now keeps the **destination** for both the inbound and outbound frames. For inbound restore, this means an exact timestamp tie keeps the host copy. On a tie, both credentials are the same generation — preferring the canonical destination (host) copy is more conservative. This is an intentional consequence of the unified guard. ## Verification ```sh # From the codex-local package root: pnpm vitest run src/server/codex-auth-merge.test.ts # Expected: all predicate rows green (15-row source/destination matrix + token-leak guard) # Plus: codex-auth-merge-extract integration rows green ``` The test suite drives the `.cjs` as a child process over temporary `auth.json` fixtures: - Source/destination predicate matrix (15 rows): strictly-newer→`10`; tie→`20`; older→`20`; identity-mismatch→`20`; source last_refresh unparseable→`20`; destination last_refresh unparseable→`20`; apikey either side→`20`; kind-mismatch→`20`; unusable either side→`20` - Token-bytes non-emission guard: asserts no auth token bytes appear in `stdout`/`stderr` - Inbound extract integration: `codex-auth-merge-extract.sh` end-to-end rows including tie/identity-mismatch/sandbox-newer cases ## Risks - **Tie behavior change**: exact `last_refresh` ties now always keep the destination copy (was: keep-sandbox for inbound). On a tie both credentials are the same generation; keeping the canonical destination (host) copy is the more conservative choice. Covered by test assertions. - **Token non-emission**: `stdout`/`stderr` of the `.cjs` must not echo credential bytes; covered by the token-bytes non-emission guard test. - **Unknown/extra flags**: handled with `process.exit(1)` rather than a silent default, so misconfigured callers surface errors immediately. - **Caller compatibility**: the only runtime caller, `codex-auth-merge-extract.sh`, already calls `node decision <sandbox> <host>` with no flags — compatible with the new source/destination API with no changes required. - **Low risk overall**: the script is exercised end-to-end via the extract integration tests; all paths covered. ## Model Used Claude claude-sonnet-4-6 (Anthropic, `claude-sonnet-4-6`), 200 k context window, tool use and agentic code execution enabled. Automated Paperclip multi-agent system. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.7 |
||
|
|
7ffeafad1f |
refactor(codex-local): relocate Codex auth-merge scripts + decision predicate into the adapter (#9785)
## Thinking Path > - Paperclip is an open source platform for managing AI agent teams, with adapter packages providing concrete runtime environments (e.g. `codex-local` runs a local OpenAI Codex sandbox) > - `adapter-utils` is the shared utilities package — it should hold only generic, adapter-agnostic primitives (sandbox lifecycle helpers, merge logic, type definitions) usable by every adapter > - Three files lived in `adapter-utils/src/` that are entirely Codex-specific: `codex-auth-merge-extract.sh` (shell script that extracts auth tokens), `codex-auth-merge-decision.cjs` (CJS decision helper), and `codex-auth-merge-scripts.ts` (TypeScript factory that wires them into the inbound provision seam added in PR #9778) > - Their presence in a "generic" package violates the adapter isolation principle and requires a cross-package build step to copy `.sh`/`.cjs` files into `codex-local/dist/server/` at build time > - This PR completes Phase 2 of the inbound-seam refactor: move all three files to `packages/adapters/codex-local/src/server/`, rebase the TypeScript imports, update `execute.ts` to import locally, move the Codex-auth tests into a new `codex-local` test file, and fix packaging so `codex-local` copies its own scripts to `dist/server/` > - Script bytes are identical after the move; no change to inbound auth-merge behavior or to which bytes cross the sandbox boundary > - The benefit is a clean ownership boundary: `adapter-utils` retains only generic runtime code, and the structural "Codex-free core" test in the adapter-utils suite validates this invariant going forward ## Linked Issues or Issue Description No pre-existing public GitHub issue. Describing the underlying problem inline: **Problem:** `packages/adapter-utils/src/` contains three files (`codex-auth-merge-extract.sh`, `codex-auth-merge-decision.cjs`, `codex-auth-merge-scripts.ts`) consumed exclusively by the `codex-local` adapter. Their presence in a generic utilities package violates adapter isolation and requires a cross-package build step (copy `.sh`/`.cjs` into `codex-local/dist/server/`). The inbound provision seam landed in PR #9778 routed these through `adapter-utils`; this PR finishes the relocation. **Related:** Refs #9778 (Phase 1 — generic asset-lifecycle-seam, now merged). ## What Changed - Moved `codex-auth-merge-extract.sh`, `codex-auth-merge-decision.cjs`, and `codex-auth-merge-scripts.ts` from `packages/adapter-utils/src/` → `packages/adapters/codex-local/src/server/` (script bytes are unchanged; `codex-auth-merge-scripts.ts` imports rebased to `adapter-utils` subpaths for `shellQuote` and `SandboxManagedRuntimeAssetProvision`) - `packages/adapters/codex-local/src/server/execute.ts`: updated import of `buildCodexAuthInboundProvision` from `adapter-utils` → local `./codex-auth-merge-scripts` - New `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` — 5 Codex-auth test rows extracted from `adapter-utils/src/workspace-restore-merge.test.ts`; the generic test file stays and no longer references Codex - `packages/adapters/codex-local/package.json`: build now copies `.sh`/`.cjs` from `src/server/` into `dist/server/` directly - `packages/adapter-utils/package.json`: removed the now-obsolete cross-package copy step for those files ## Verification ```sh # Type-check both affected packages pnpm --filter @paperclipai/adapter-utils --filter @paperclipai/adapter-codex-local typecheck # Codex auth-merge suite (moved tests — 5/5 pass) pnpm --filter @paperclipai/adapter-codex-local test -- --reporter=verbose codex-auth-merge # Full adapter-utils suite including the structural "core free of Codex literals" test (243 passed, 4 pre-existing skips) pnpm --filter @paperclipai/adapter-utils test # Confirm build copies scripts to dist/server pnpm --filter @paperclipai/adapter-codex-local build ls packages/adapters/codex-local/dist/server/*.sh packages/adapters/codex-local/dist/server/*.cjs ``` All of the above were run locally and passed before this PR was opened. ## Risks Low risk — pure file relocation: - No change to `.sh` or `.cjs` script bytes; the same content reaches the sandbox boundary as before - No behavioral change to inbound auth-merge logic from the sandbox's perspective - `adapter-utils`' structural "core free of Codex literals" test now validates that the relocation is complete and will catch any regression - Only `codex-local` consumed these files from `adapter-utils`; no other package in the monorepo imported them from there ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) — Anthropic, 200k context window, tool use, agentic reasoning. Used to author the refactor. `Co-authored-by: Paperclip <noreply@paperclip.ing>` trailer present on all commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cf5ba4bbea |
feat(adapter-utils): generic per-asset lifecycle-contribution seam (#9778)
## Thinking Path
> - Paperclip's sandbox managed runtime is responsible for provisioning
the agent's execution environment — it extracts a home directory asset
into the sandbox before the adapter runs.
> - The sandbox runtime core was directly branching on the adapter key
(`codex`) to decide which merge scripts to stage and which merge-extract
command to run, coupling generic infrastructure to a specific adapter's
credential-merge protocol.
> - This makes it harder to add, remove, or modify per-adapter asset
provisioning without touching the runtime core; it also prevents other
adapters from contributing staged files or a custom extract command at
all.
> - The fix is to move the adapter-specific knowledge into the adapter
itself: the asset descriptor gains optional `provision` (stageFiles +
extractCommand) and `restore` contribution fields that any adapter can
populate, and the runtime core consumes them generically.
> - This pull request introduces those contribution fields, wires the
Codex adapter's inbound credential-merge as a `provision` contribution,
and removes the adapter-specific branching from the runtime core.
> - The benefit is a clean seam: the runtime core is now
adapter-agnostic for asset provisioning, the inbound behavior is
unchanged (same merge matrix, same scripts), and other adapters can
attach custom staged files or extract commands without modifying shared
infrastructure.
## Linked Issues or Issue Description
No pre-existing public GitHub issue. Describing the problem inline per
the feature template:
**Problem or motivation**
The sandbox managed-runtime asset provisioning in
`sandbox-managed-runtime.ts` branched directly on the adapter key
(`codex`) to decide which merge scripts to stage and which shell command
to use during asset extraction. This tight coupling prevents other
adapters from customizing their provisioning without modifying the
runtime core, and it means the runtime core must import and know about
adapter-specific merge scripts.
**Proposed solution**
Add an optional `provision` contribution (array of `stageFiles` entries
+ an `extractCommand` string) and an optional `restore` contribution to
the asset descriptor returned by adapters. The runtime core now consumes
these generically — if a `provision` contribution is present, it stages
those files and uses the supplied command; otherwise it falls back to
the default `tar -xf` extraction. The Codex adapter populates the
`provision` contribution where it previously depended on core branching.
**Alternatives considered**
Keeping the adapter-specific logic in the core as a documented
exception; rejected because it makes the seam inextensible.
**Roadmap alignment**
Decoupling — removes a latent coupling between the runtime core and a
specific adapter.
## What Changed
- Added `provision` contribution field (`stageFiles: Array<{src, dest}>`
+ `extractCommand: string`) to the `SandboxManagedRuntimeAsset`
descriptor type in `adapter-utils`.
- Added `restore` contribution field (hook for post-restore logic,
populated in a later phase) to the descriptor.
- Removed adapter-key branching (`if adapterKey === 'codex'`) from the
runtime core in `sandbox-managed-runtime.ts`; the core now reads
`provision.stageFiles` and `provision.extractCommand` generically.
- Extracted Codex-specific merge-script paths and the merge-extract
command into `codex-auth-merge-scripts.ts` in `adapter-utils`; the Codex
adapter's `execute.ts` now attaches them as a `provision` contribution
when it builds its managed-home asset descriptor.
- Updated `execution-target.ts` to pass the extended asset type through
to the adapter call site so the new fields are load-bearing end-to-end.
- Added seam-proving unit tests in `sandbox-managed-runtime.test.ts`:
contribution-less asset uses the default path; a non-adapter asset
round-trips the generic provision+restore seam; a structural assertion
verifies the runtime core carries no Codex-specific string literals.
- Added one test in `workspace-restore-merge.test.ts` confirming the
inbound merge matrix is unaffected.
## Verification
```bash
# Unit tests (20 pass):
npx vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/adapter-utils/src/workspace-restore-merge.test.ts
# Type-check both affected packages:
cd packages/adapter-utils && npx tsc --noEmit
cd packages/adapters/codex-local && npx tsc --noEmit
# Structural: runtime core carries no adapter string literals
grep -n 'codex\|auth\.json' packages/adapter-utils/src/sandbox-managed-runtime.ts
# Expected: zero matches
```
## Risks
**Low risk.** This is a behavior-preserving refactor: the inbound
provisioning output (which files get staged, which command runs) is
identical to before, now driven by the adapter-supplied contribution
instead of core branching. The existing inbound merge matrix tests are
the regression guard. No change to which bytes cross the sandbox
boundary. The SSH transport is untouched.
## Model Used
- **Provider:** Anthropic
- **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`)
- **Context window:** 200 K tokens
- **Capabilities used:** tool use (file read/edit, bash execution,
Paperclip API), extended reasoning over multi-file TypeScript refactor
- **Mode:** agentic (Paperclip ACPX platform)
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Harold Kim <harold@paperclip.ing>
canary/v2026.717.0-canary.6
|
||
|
|
051ae4d102 |
feat: restore decision training library and inspector (#9779)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and supervise governed work. > - Decisions capture high-value operator judgment, and the decision-training foundation merged in #9702 freezes that evidence for later evaluation and learning. > - Operators still need the UI from the closed stacked PR #9718 to intentionally capture examples and inspect the resulting dataset. > - GitHub automatically closed #9718 when its stacked base branch was deleted after #9702 merged, leaving the server foundation on `master` without the corresponding UI. > - This pull request restores the final UI and its still-required supporting API fields directly on current `master`, while excluding the obsolete migration and duplicated server-foundation diffs. > - The benefit is a reviewable replacement PR that preserves the completed decision-training workflow without replaying stale stack history. ## Linked Issues or Issue Description - Refs #9718 - Refs #9702 ## What Changed - Restored the top-level `/training` library and record inspector with search, filters, JSONL export, notes editing, and evidence tabs. - Restored the Decisions-row training affordance and capture drawer, including preview, provenance, deletion, cache refresh, and approval consistency behavior. - Restored the shared types and focused server support needed by the UI without reintroducing decision-training migrations or the already-merged server foundation. - Restored focused UI and attention-service tests from the final #9718 state. - Credit to the authors and reviewers of #9718; this recovery transplants their final reviewed delta after the stacked base deletion. ## Verification - `pnpm exec vitest run ui/src/pages/Training.test.tsx ui/src/components/DecisionTrainingDrawer.test.tsx ui/src/components/AttentionQueueRow.test.tsx server/src/__tests__/attention-service.test.ts server/src/__tests__/decision-training.test.ts` — 5 files, 48 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/openapi-routes.test.ts` — 3 tests passed; confirms exact route/OpenAPI parity for the restored preview endpoint. - `pnpm check:token-gates` — the restored files are clean; the repository-wide command currently reports five pre-existing false positives where comments reference GitHub issue `#9627` as if it were a color literal. ## Risks - Low migration risk: this PR contains no database migrations and is based directly on current `master`. - The main behavioral risk is cache invalidation across Decisions and Training views; focused tests cover capture, update, deletion, row state, and approval refresh behavior. - The token-gate baseline remains red on unrelated `#9627` comment references; this PR does not modify those files. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3-Codex, reasoning with repository/tool access and code execution. Context window not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f7e511a3e5 |
fix(ui): stop the inbox unread dot from indenting rows (#9767)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Inbox lists issues as rows; unread rows show a small blue mark-read dot at the leading edge > - The dot was rendered as an `order-first` flex column that existed only on unread rows, so it changed each row's leading layout depending on read state > - On parent (chevron) rows the dot stacked on top of the collapse chevron and pushed the whole row's status/title to the right, so unread rows sat indented relative to their read siblings > - This PR gives every inbox row a reserved leftmost dot slot so the dot never changes horizontal alignment > - The benefit is a stable, non-shifting inbox: read and unread rows line up at the same x, and the dot no longer crowds the collapse caret ## Linked Issues or Issue Description No public issue exists. Describing as a bug report: - **What happened:** In the Inbox, an unread issue row's blue mark-read dot pushed the row's status/title to the right, so unread rows were indented relative to read siblings. On parent (collapsible) rows the dot also overlapped the collapse chevron. - **Expected behavior:** The unread dot should not change a row's horizontal alignment. Read and unread rows should line up identically, and the dot should sit clear of the collapse caret. - **Steps to reproduce:** Open the Inbox with a mix of read/unread issues, including a parent issue with children. Observe unread rows indented vs. read rows, and the dot crowding the chevron on parent rows. - **Deployment mode:** UI-only (frontend). Related PR (same inbox area, in flight): https://github.com/paperclipai/paperclip/pull/9685 ## What Changed - `ui/src/components/IssueRow.tsx`: Add a reserved leftmost dot slot (`w-4`, desktop) that is present on **all** inbox rows — read and unread alike — so the mark-read dot renders far-left ahead of any leading control without shifting content. The dot button markup is now shared between the desktop reserved slot and the mobile in-flow rendering. The row is `relative` for slot positioning. - `ui/src/pages/Inbox.tsx`: Always reserve the leading spacer for non-chevron rows (removed the special-case that dropped it for unread rows), since the dot now has its own dedicated slot to the left of the spacer. This removes the double-indent. - Mobile keeps the dot in flow as the leading item (no reserved desktop gutter there). - Tests: `IssueRow.test.tsx` and `Inbox.test.tsx` updated to assert the reserved dot slot exists on read and unread rows and that alignment no longer differs by read state. ## Verification - `cd ui && NODE_ENV=development npx vitest run src/components/IssueRow.test.tsx src/pages/Inbox.test.tsx` → 28 tests pass. - `cd ui && npm run typecheck` → clean. - Note: running the full UI suite with `NODE_ENV=production` surfaces pre-existing `act is not a function` failures unrelated to this change (React's production bundle strips `act`); they pass under a normal test env. ## Risks Low risk. UI-only, scoped to inbox row layout. The reserved slot is `hidden ... sm:inline-flex`, so mobile layout is unchanged (dot stays in flow). No API, schema, or telemetry changes. ## Model Used Claude Opus 4.8 (claude-opus-4-8), 1M context window, extended thinking + tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no docs affected) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.5 |
||
|
|
a090c09ee5 |
feat: add decision training snapshot foundation (#9702)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - Human approvals, issue interactions, and execution decisions already capture high-value decision moments > - Those moments are currently transient and cannot be reused as stable evaluation or training examples > - Reusable examples need a server-owned, immutable snapshot so later comments or runs cannot leak into the recorded state > - Human notes need to remain editable and auditable without changing the captured state > - This pull request adds the database model, snapshot capture service, API, export format, and attention-feed enrichment for decision training > - The benefit is a durable, inspectable foundation for evaluating whether agents can reproduce good human decisions from only the context available at decision time ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`server/`, `packages/db`, and `packages/shared`). ### Problem or motivation Paperclip has no durable dataset for converting human decisions into evaluation-ready examples. Teams need to capture pending or resolved decisions with the exact issue context, comments, runs, and repository evidence available at a cutoff, while preventing future context from leaking into the example. ### Proposed solution Store immutable, schema-versioned snapshots anchored to durable interaction, approval, or execution-decision records; keep notes separately editable with history; expose human-only CRUD, list, and JSONL export APIs. ### Alternatives considered Client-generated snapshots were rejected because they duplicate cutoff logic and cannot reliably enforce no-leakage boundaries. Automatic outcome backfill was deferred so captured examples remain faithful to what was known at capture time. ### Roadmap alignment Supports the roadmap direction of turning completed work and decision patterns into reusable organizational knowledge. ### Additional context The implementation records explicit commit-resolution confidence (`exact`, `nearest_run`, `workspace`, or `none`) so downstream evaluation can distinguish evidence quality. ## What Changed - Added the `decision_training_examples` schema and idempotent migration with company, issue, and source/author indexes. - Added shared types for decision-training records, notes history, and versioned snapshots. - Added a single server-side snapshot capture path with inclusive comment cutoffs, pre-cutoff run capture, durable decision payloads, and explicit commit-resolution confidence. - Added create, list, detail, notes-only update, delete, and JSONL export routes with human-only write authorization and activity logging that skips no-op note submissions. - Added per-user `trainingExampleId` enrichment to attention items. - Added focused embedded-Postgres tests for cutoff boundaries, post-cutoff leakage, immutable snapshots, human-only writes, duplicate prevention, notes history, attention enrichment, and export shape. - Updated UI test and Storybook attention-item factories for the new required `trainingExampleId` contract. ## Verification - `pnpm exec vitest run server/src/__tests__/decision-training.test.ts` — 10 tests passed. - `pnpm --filter @paperclipai/db typecheck` — passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. ## Risks - The migration adds a new table and indexes only; it does not rewrite existing rows or install resolve-time hooks. - Snapshot JSON can grow with long comment threads and run histories; v1 intentionally favors complete, inspectable examples over aggressive truncation. - Commit SHA resolution is evidence-based and records `exact`, `nearest_run`, or `none` so downstream consumers can account for confidence. - The API is additive, but future UI work must continue to treat the snapshot as immutable and use notes-only updates. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.3-codex`, with repository tool use, terminal execution, and code-editing capabilities; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.4 |
||
|
|
1f1f545238 |
feat: add built-in summarizer and summary slots (#9713)
## Thinking Path > - Paperclip is the open source control plane people use to organize, govern, and understand AI-agent work > - Operators need concise, current status views across projects and execution workspaces without manually reading every issue and run > - Paperclip already has auditable issues, documents, built-in agents, routines, and live run events, but no first-class summary-slot workflow connecting those systems > - A built-in Summarizer can generate status prose through ordinary governed tasks while summary slots provide stable, revisioned destinations for that output > - The UI needs to show current summaries, generation progress, failures, revisions, and streaming draft status in the places operators already work > - This pull request adds the end-to-end summary-slot data, API, agent, orchestration, and UI surfaces behind an experimental setting > - The benefit is decision-oriented status context that remains company-scoped, auditable, retryable, and inexpensive by default ## Linked Issues or Issue Description No public GitHub issue exists for this feature. **Problem** Operators currently have to reconstruct project and workspace status by reading many issues, runs, and comments. This makes it hard to identify decisions, review queues, recent work, and the next event worth watching. **Proposed capability** Add an experimental summary system with revisioned summary slots for projects and workspaces, a paused-by-default built-in Summarizer agent, governed generation tasks, live draft status, and reusable UI cards. **Expected behavior** - Summary data remains company-scoped and revisions remain auditable. - Generation runs through normal issue/agent orchestration and deduplicates active requests. - Only the linked built-in Summarizer generation task can author a slot revision. - Operators can generate, retry, inspect revisions, and follow draft progress from project and workspace views. - The feature remains opt-in and background generation remains paused by default. ## What Changed - Added summary-slot schema, idempotent migrations, shared contracts, validators, API paths, and service tests. - Added company-scoped summary-slot routes for reading revisions, requesting generation, and guarded Summarizer writes with activity logging. - Added terminal generation finalization, failure reasons, assignment wakeups, and orchestration integration. - Added the paused-by-default built-in Summarizer bundle, low-cost runtime defaults, status-summarization skill, and stale-summary routine. - Added summary cards, revision selection, retry/configuration states, live draft streaming, transcript chunk handling, and project/workspace integrations. - Updated Claude local parsing for streamed status output and expanded server, adapter, shared, database, catalog, and UI coverage. ## Verification - `pnpm -r typecheck` - `pnpm exec vitest run packages/db/src/summary-slots-schema.test.ts packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slot-routes.test.ts server/src/__tests__/summary-slots.test.ts server/src/__tests__/built-in-agents.test.ts ui/src/components/SummarySlotCard.test.tsx ui/src/components/SummarySlotCard.status.test.tsx ui/src/components/useSummaryDraftStream.test.tsx ui/src/lib/summary-draft-stream.test.ts ui/src/lib/run-log-chunks.test.ts ui/src/context/LiveUpdatesProvider.hook.test.tsx` — 113 tests passed - `pnpm test:run` — server and UI suites passed; one CLI AWS doctor test was affected by inherited `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, and passed when those host credentials were removed - `pnpm exec vitest run cli/src/__tests__/secrets.test.ts` with inherited AWS credential variables removed — 8 tests passed - `pnpm build` - `pnpm check:token-gates` currently reports nine `#9627` comment references introduced by current `master`; none are in this PR diff ## Risks - Database risk is limited by incrementally ordered, idempotent migrations and migration safety checks. - Summary generation creates normal issues/runs, so misconfiguration can produce failed slots; the UI exposes retryable failure reasons and agent configuration entry points. - Streaming draft parsing depends on the documented `STATUS:` protocol; final persisted revisions remain the source of truth. - The feature is experimental, opt-in, and its built-in routine is paused with no background token spend by default. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.3 Codex with reasoning, repository tool use, code execution, GitHub CLI, and Paperclip control-plane integration. Earlier branch commits also record Claude model co-authorship where applicable. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.717.0-canary.3 |
||
|
|
009410164f |
Fix inbox unread badge alignment (#9685)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Inbox lets operators scan task state and distinguish unread activity at a glance > - Read and unread rows should keep the same status-and-title alignment so the list remains easy to scan > - The mark-read dot previously participated in flex layout, which added an extra leading column and visibly indented unread rows > - Parent rows also needed special handling because their collapse chevron occupies the normal leading gutter > - This pull request moves the unread dot out of desktop flex flow, preserves a consistent leading spacer, and positions parent-row dots after the chevron > - The benefit is consistent alignment for read, unread, leaf, nested, and collapsible parent rows without hiding tree controls ## Linked Issues or Issue Description ### What happened? In the Inbox, an unread row rendered its status icon and title farther right than an equivalent read row because the mark-read dot occupied its own flex column. On a collapsible parent row, the dot also competed with the leading chevron. ### Expected behavior Read and unread Inbox rows keep identical status/title alignment while preserving the unread affordance and any tree-expansion controls. ### Steps to reproduce 1. Run Paperclip from `master` and open the Inbox at desktop width. 2. Compare otherwise-equivalent read and unread leaf rows. 3. Expand a task tree containing an unread depth-zero parent. 4. Observe the unread row indentation and the dot competing with the parent chevron. ### Paperclip version or commit Reproduced against the pre-fix `master` parent of this PR. - **Deployment mode:** Local dev (`pnpm dev`). - **Installation method:** Built from source. - **Agent adapters involved:** Not adapter-specific; this is a core Inbox UI bug. - **Database mode:** Not database-related. - **Access context:** Board (human operator). - **Additional context:** No logs or configuration are involved; the regression is visual layout behavior covered by focused component/page tests. ## What Changed - Render the desktop unread dot as an absolute overlay so it does not consume row width; retain the existing in-flow behavior on mobile. - Add an `unreadDotPlacement` option so depth-zero collapsible parents place the dot after the leading chevron. - Reserve the same leading spacer for read and unread non-chevron Inbox rows. - Expand `IssueRow` and Inbox tests to cover leaf alignment, nested rows, parent chevrons, fading state, and mobile behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueRow.test.tsx ui/src/pages/Inbox.test.tsx` — 2 files, 30 tests passed. - `pnpm check:token-gates` — all three token gates clean across 630 scanned files. - Browser QA (desktop 1280px and mobile 375px) — PASS: read/unread desktop content measured at identical x positions; nested guides, parent chevron hit target, fading state, and mobile mark-read hit targets verified. ## Risks - Low risk: the change is isolated to Inbox row presentation and has focused regression coverage. - The main visual risk is breakpoint-specific placement of the mark-read dot; tests explicitly cover desktop absolute positioning and mobile in-flow positioning. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex `gpt-5.6-sol`, high reasoning effort, with repository/tool execution. The runtime did not expose a context-window size. - Original implementation commit also records assistance from Claude Opus 4.8. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.2 |
||
|
|
5d42382df4 |
feat(ui): stabilize workspace service controls (#9705)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent work. > - Execution workspaces provide the local service loop where operators start, stop, restart, inspect, and open workspace services. > - The existing header exposed those actions through separate controls whose position and labeling changed across runtime states. > - That movement made the most common development actions harder to scan and easier to misclick, especially with multiple services or long URLs. > - The design therefore uses one fixed-geometry, state-aware control bar and keeps service-specific detail behind a compact disclosure. > - This pull request adds that control surface, maps existing runtime data and every pending mutation into it, and integrates it into the execution-workspace header without changing server contracts. > - The benefit is a calmer, predictable service-control loop across stopped, transitional, running, unhealthy, failed, multi-service, and narrow-width states. ## Linked Issues or Issue Description ### Subsystem affected `ui/ — React + Vite board UI` ### Problem or motivation Execution-workspace service actions move and change shape as runtime state changes, while URLs and multi-service status compete for header space. During bulk actions, operators also need every targeted service to show its transitional state immediately. ### Proposed solution Use one fixed-geometry, state-aware service control bar in the workspace header. Map existing runtime records into a stable status, URL, and actions model, and track each in-flight bulk request independently until it settles. ### Alternatives considered Keeping the separate quick-control buttons was rejected because their geometry changes by state. Showing every service inline was rejected because it makes the header too wide; per-service detail remains in a compact disclosure and the Services tab. ### Roadmap alignment Reviewed `ROADMAP.md`; this focused execution-workspace UI improvement does not duplicate a listed roadmap initiative. ### Additional context The published design and state viewer is available at https://pages.paperclip.ing/pap-14233-workspace-service-controls/. ## What Changed - Added `WorkspaceServiceControlBar`, a fixed-geometry responsive control for single- and multi-service runtime states. - Added 15 Storybook states covering running, stopped, transitions, unhealthy, failed, disabled, long-URL, mobile, and multi-service behavior. - Replaced `WorkspaceRuntimeQuickControls` in the execution-workspace header with adapters that map live services and all pending requests into the new control model. - Added focused unit coverage for service-entry construction, bulk pending overlays, request resolution, clipboard feedback, and header integration. ## Verification - `cd ui && NODE_ENV=development pnpm vitest run src/components/WorkspaceServiceControlBar.test.tsx src/components/WorkspaceRuntimeControls.test.tsx src/pages/ExecutionWorkspaceDetail.test.tsx` — 29 tests passed. - `NODE_ENV=development pnpm --dir ui typecheck` — passed. - `pnpm check:token-gates` — all token gates clean. - Reviewed the Storybook captures for all primary states. ## Risks - Low-to-moderate UI risk: service controls depend on adapter mapping from existing runtime records; focused tests cover single-service and bulk-action mapping and integration paths. - Multi-service bulk actions intentionally apply to all eligible services, while per-service actions remain in the disclosure. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.3-codex`; runtime-managed context window; coding/reasoning mode with repository, terminal, Git, GitHub CLI, and test execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.1 |
||
|
|
b07b2994cc |
feat: stamp responsible users on activity logs (#9731)
## Thinking Path > - Paperclip is the control plane people use to manage AI-agent companies and their work > - The activity log is the generic audit spine for mutations across the control plane > - Activity rows identify agents and runs, but they do not persist the responsible human upstream > - Relying only on run joins loses attribution after run pruning and misses agent API-key actions outside a run > - This pull request resolves responsible-user attribution when each activity row is written and stores it directly > - The benefit is durable, queryable agent audit feeds without rewriting historical provenance ## Linked Issues or Issue Description ### Problem or motivation Agent activity records do not persist the responsible user, so attribution can disappear when runs are pruned and no-run API-key mutations cannot be attributed correctly. ### Proposed solution Resolve attribution for each new activity row from the run, related issue, active agent API key, or company default, in that order, and persist the result directly. ### Alternatives considered Read-time joins alone were rejected because pruned runs lose durable attribution and out-of-run agent-key actions have no run to join. Historical backfill was rejected because it would invent provenance. ### Roadmap alignment This strengthens the durable audit-trail direction described in `ROADMAP.md` without adding a new product surface. ## What Changed - Added nullable `activity_log.responsible_user_id` plus company/agent/time and company/responsible-user/time indexes. - Added an idempotent forward-only migration with no historical backfill. - Added centralized write-time resolution: heartbeat run → issue attribution → active agent API key → company default. - Propagated authenticated API-key IDs through existing request-backed `logActivity()` calls. - Added unit coverage for every fallback and an embedded-Postgres assertion for the no-run API-key stamping path. ## Verification - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/activity-log-responsible-user.test.ts src/__tests__/authz-company-access.test.ts` - Initial focused verification: 25 tests passed; migration safety passed. - Follow-up regression verification: 54 focused attribution/company-skill/environment/issue-tree/tool-gateway tests passed; server typecheck passed. - GitHub: full build, typecheck, server shards, serialized suites, e2e, security, and policy checks passed. ## Risks - Adding two indexes to an existing large table can hold a write lock while the transactional migration runs. The migration safety suppressions document why `CONCURRENTLY` is unavailable under the current Drizzle migration runner. - Historical rows remain nullable by design; this avoids inventing provenance and keeps the migration forward-only. - API-key attribution requires request-backed activity call sites to pass the authenticated key ID; this PR mechanically updates the existing actor-based activity calls and covers the no-run path with integration testing. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.4`; context-window size was not exposed by the runtime. Medium reasoning with repository editing, terminal execution, and test execution capabilities was used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.717.0-canary.0 |
||
|
|
53f09cb818 |
fix: prevent duplicate task creation and recovery loops (#9648)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents, routines, productivity review, and recovery services can all create or re-trigger work > - Repeated heartbeats or catch-up cycles can produce duplicate tasks or repeat recovery actions before prior work is visible > - The base duplicate-create and routine catch-up coalescing work has now landed on `master` via related PRs while this PR was being prepared > - This pull request carries the remaining hardening: bounded idempotency retention, recovery cooldown/throttle fixes, productivity-review query batching and ordering fixes, and regression coverage > - The benefit is fewer duplicate tasks, safer retries, and enough provenance to diagnose any future recurrence ## Linked Issues or Issue Description Agents can retry issue creation after ambiguous responses or independently recreate the same child title, while recovery and short-interval routine catch-up paths can repeat before prior work settles. This can produce visible duplicate tasks and makes the originating heartbeat difficult to identify. Related work: Refs #8356 for caller-supplied issue-create idempotency and Refs #9224 for plugin-scoped issue-create idempotency. Prior related PR: #6936. The base issue-create deduplication and routine catch-up coalescing pieces have since landed on `master` via #9650 and #9649; this PR remains as the follow-up hardening stack on top of those changes. ## What Changed - Add 7-day retention for issue-create idempotency claims with indexed, batch-limited cleanup so the claim table does not grow forever. - Preserve recovery cooldown intent after terminal recovery actions are closed, and throttle repeated source-scoped recovery work. - Batch productivity-review source-activity checks to avoid repeated per-source queries while keeping the no-action suppression behavior. - Order productivity-review no-action streak windows by review creation time, matching the window semantics even when completion timestamps are out of order. - Preserve generated issue IDs in route mocks used by backlog/assignment contract tests. - Document PR-gardening task deduplication expectations in the company skill. - Add focused regression tests for idempotency retention, liveness recovery cooldowns, and productivity-review batching/suppression/ordering behavior. ## Verification - `pnpm --filter @paperclipai/server typecheck` — passed on latest head. - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/issue-create-deduplication-routes.test.ts --reporter verbose` — 2 files, 23 tests passed on latest head. - `pnpm --filter @paperclipai/adapter-utils build && pnpm exec vitest run --project @paperclipai/adapter-utils --reporter dot` — 30 files passed, 476 tests passed, 8 skipped. - `pnpm build` — passed on latest head. - `pnpm -r typecheck` — passed on the rebased head before the final productivity-review ordering commit; the latest touched server code is covered by the server typecheck above. - `pnpm check:token-gates` — passed. - `pnpm test:run` — progressed through server, UI, CLI, shared, skills-catalog, and DB sections, then exposed an adapter-utils compiled-test fingerprint mismatch before rebuilding adapter-utils; the adapter-utils project passed after rebuild, and the GitHub split PR checks passed on the pushed head. ## Risks - Caller-supplied idempotency replay is now bounded to 7 days; reusing an old key after retention can create new work, which matches retry-oriented idempotency semantics. - Recovery and productivity-review timing changes may suppress redundant follow-up work; focused tests cover the intended boundaries. - Advisory locking and idempotency cleanup rely on PostgreSQL-compatible transaction semantics already used by the production data layer. - The migration extends the private claim table indexes without rewriting existing issue rows. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5, with reasoning, repository editing, shell/tool execution, and test execution. Exact model ID and context-window size are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.21 |
||
|
|
a015ac7a57 |
fix(ui): allow interrupting queued issue runs (#9725)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue detail chat lets operators send messages while an issue run is live > - Messages sent during a live run are shown as queued and can expose an interrupt action > - The UI only treated running runs as interruptible, even though queued runs can also own the pending message > - That mismatch hid the interrupt button after sending a message while the agent run was still queued > - This pull request treats queued and running issue runs as interruptible and preserves the exact target run on optimistic and persisted comments > - The benefit is that operators can immediately interrupt the queued run their message is waiting behind ## Linked Issues or Issue Description - **What happened:** Sending a message while an issue-owned agent run was in `queued` state showed the message as pending but did not make the interrupt action available. - **Expected behavior:** A message queued behind either a queued or running issue run should retain that run as its interrupt target and expose the interrupt control. - **Steps to reproduce:** Open an in-progress issue with an issue-owned run still queued, send a chat message, and inspect the queued message actions. - **Version/commit:** Reproduced against the pre-change `master` UI behavior. - **Deployment mode:** Paperclip board UI with a queued issue execution run. ## What Changed - Generalized issue-run resolution from running-only to queued-or-running interruptible runs. - Used the interruptible run consistently for optimistic queue metadata, persisted comment decoration, cancel controls, and targeted interruption. - Added a regression test that sends a message behind a queued run and verifies the exact run is cancelled. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx` — 44 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — repository-wide gate currently reports nine pre-existing `#9627` comment literals; this patch adds no token literals or gate violations. ## Risks - Low risk: the behavior change is limited to selecting queued issue-owned runs as valid interrupt targets in the existing chat flow. - Cancellation remains targeted by run ID, and the regression test verifies the queued run ID is preserved through optimistic and persisted comment states. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` via Codex CLI; context-window size is not exposed by this runtime; reasoning-enabled with repository, shell, GitHub CLI, and code-execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.20 |
||
|
|
59fb27ff79 |
feat(inbox): let agents safely tidy user inboxes (#9724)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - The inbox is a per-user attention view, so archiving an item must not alter the underlying issue, assignment, or status > - Agents can help responsible users tidy resolved work only when the action is company-scoped, reversible, policy-controlled, and fully attributable > - The database and authorization foundations landed in #9654 and #9658, but the end-to-end archive routes, audit details, agent workflow guidance, and operator UI still need to ship together > - Separate stacked PRs #9659 and #9661 made the complete behavior harder to review and land as one coherent capability > - This pull request consolidates the remaining server, shared-contract, documentation, skill, and UI work on top of current master > - The benefit is a single reviewable change that lets agents safely archive responsible-user inbox items and lets users control or undo that behavior ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting inbox management across shared contracts, server authorization/routes/services, shipped agent skills, and the board UI. ### Problem or motivation Agents may complete work whose issue remains in the responsible user's Mine inbox. Existing board-user archive behavior does not provide the agent-facing policy endpoints, target resolution, heartbeat-run attribution, typed denials, conservative workflow guidance, or UI needed for safe agent-managed cleanup. ### Proposed solution Allow authorized agents to archive or unarchive responsible-user inbox items under the user's open, allowlist, or disabled policy; preserve actor/agent/run attribution in issue detail and activity records; expose policy controls and agent archive attribution in the UI; and document conservative cleanup rules for agents and PR gardening. ### Alternatives considered - Reuse generic issue mutation permissions: rejected because inbox state belongs to a target user and requires user-scoped authorization. - Automatically archive every completed or closed item: rejected because completion signals can still require human review or a decision. - Keep the backend and UI as separate stacked PRs: superseded by this consolidated PR so the complete user-visible behavior can be reviewed and verified together. ### Related work - Builds on merged foundations #9654 and #9658. - Supersedes the remaining stacked changes in #9659 and #9661. - `ROADMAP.md` has no overlapping inbox archive or inbox authorization initiative. ## What Changed - Added shared inbox-agent policy types and validators plus company-scoped self-service policy routes and OpenAPI coverage. - Enabled agent archive/unarchive mutations with responsible-user targeting, policy enforcement, typed failures, attribution, idempotency, and detailed activity auditing. - Returned agent archive attribution in issue detail and documented reversible inbox cleanup semantics in the implementation spec and Paperclip skill. - Added conservative PR-gardening inbox tidy guidance that keeps GitHub access read-only and avoids archiving work that still needs human action. - Added the Profile settings policy control and Issue Properties attribution/unarchive UI with focused component coverage and narrow-pane handling. ## Verification - `pnpm exec vitest run server/src/__tests__/inbox-archive-routes.test.ts server/src/__tests__/inbox-agent-policy-routes.test.ts server/src/__tests__/authorization-service.test.ts server/src/__tests__/openapi-routes.test.ts ui/src/components/InboxAgentPolicyControl.test.tsx ui/src/components/IssueProperties.test.tsx` — 110 passed. - `pnpm --filter @paperclipai/db exec vitest run src/inbox-archive-agent-policies-migration.test.ts` — 1 passed. - `node --test .agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 9 passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — changed files are clean; the repository-wide command currently reports nine unrelated pre-existing `#9627` literals outside this PR's diff. ## Risks - Agent inbox mutations broaden an existing endpoint path, so authorization and target resolution must remain fail-closed; focused route and authorization tests cover allowed and denied paths. - Archive state affects only the responsible user's inbox presentation and remains reversible; it does not mutate issue status, assignment, or visibility. - The UI policy defaults to the existing open behavior, while allowlist and disabled modes can reduce agent access. - This PR intentionally builds on #9654 and #9658 and contains no new migration number or modification to an already-applied migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4, medium reasoning, repository tool use, shell execution, code review, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`feat/inbox-agent-archive-complete`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
89af58b6dc |
Clarify task-level model overrides in issue properties (#9710)
## Thinking Path > - Paperclip is the control plane where operators configure and inspect AI-agent work. > - An issue can override the assigned agent's primary model for that task. > - The issue-properties UI labeled that per-task setting as `Custom · <model>`, which could be read as a property of the model rather than a replacement of the agent default. > - Operators need the UI to distinguish the agent's primary model from an issue-specific override at both the collapsed summary and selection point. > - This pull request renames the lane presentation to `Override` and adds concise provenance text without changing the stored lane value or adapter configuration behavior. > - The benefit is that operators can immediately understand which model will run and why it differs from the agent default. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related implementation work: #9700 moved Codex ACPX model configuration to startup and is already on `master`; #9355 also addresses ACP session-config rejection behavior. Those PRs concern execution behavior, while this PR is limited to clarifying the issue-level override UI. Bug report details: - **What happened?** The issue-properties Model row displayed `Custom · <model>` for a task-level `assigneeAdapterOverrides.adapterConfig.model`. Operators could misread `Custom` as describing the model itself and could not see that the value replaced the agent's primary model for this issue. - **Expected behavior:** The UI should explicitly identify a task-level model as an override and explain that it replaces the agent's primary model for the issue. - **Steps to reproduce:** Configure an agent with a primary model, set a different model on an issue, and inspect the Model row and model picker in Issue Properties. - **Paperclip version or commit:** Reproduced on the pre-change branch derived from current `master`. - **Deployment mode:** Local development or any deployment using the board UI. - **Installation method:** Built from source. - **Agent adapters involved:** Adapter-agnostic; any adapter exposing model selection. - **Database mode:** Not database-related. - **Access context:** Board operator viewing issue properties. - **Relevant logs or output:** Not applicable; this is a presentation ambiguity. - **Relevant config:** An issue-level `assigneeAdapterOverrides.adapterConfig.model` differing from the assigned agent's primary model. - **Additional context:** The internal lane identifier remains `custom`; only user-facing copy and explanatory text change. - **Privacy checklist:** No private instance links, internal ticket IDs, secrets, usernames, or local paths are included. ## What Changed - Renamed the collapsed issue model label from `Custom · <model>` to `Override · <model>` and added a provenance tooltip. - Renamed the model-picker lane from `Custom` to `Override` and added explanatory subtext at the selection point. - Updated the Issue Properties component test to assert the new label. - Added an isolated Storybook fixture for the task model override state so visual review has a stable target. ## Verification - `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` — 43 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `git diff --check public-gh/master...HEAD` — passed. - `pnpm check:token-gates` — the changed files are clean, but the repo-wide command currently reports nine unrelated `#9627` comment references already present on `master` as color literals. ## Risks - Low risk: this changes display strings and explanatory copy only; override storage, lane identifiers, API contracts, and execution behavior are unchanged. - The global token-gate false positive may also appear in CI until the unrelated `#9627` references on `master` are allowlisted or the scanner ignores issue-number comments. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent assisted this PR preparation using medium reasoning, repository-aware shell execution, Git/GitHub tooling, and local test/typecheck execution. The hosted runtime did not expose an exact model ID or context-window size to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.19 |
||
|
|
52aea90263 |
feat: organize skills with nested folders and My Skills (#9633)
## Thinking Path > - Paperclip is the open source control plane people use to organize and govern AI-agent companies > - Company skills are durable resources that users browse, import, assign, and maintain over time > - A flat skill list plus tags does not provide a stable location or hierarchy for personal, company, project-imported, and bundled skills > - Folder paths need to be canonical, company-scoped, safe to move, and preserved across re-imports without changing skill IDs > - The `/skills` UI also needs traversal, breadcrumbs, move/create flows, and a dedicated My Skills namespace that work on desktop and mobile > - This pull request adds the folder data model and APIs, reserved-root lifecycle, project import behavior, and the folder-first skills experience > - The benefit is a predictable filesystem-like organization model while tags remain available for cross-cutting classification ## Linked Issues or Issue Description Refs #9619 — the reviewed folder foundation was intentionally closed and folded into this combined feature PR. Refs #9026 — earlier flat-folder attempt superseded by this integrated implementation. Refs #3281 — related skill organization proposal; this PR uses canonical persisted folders rather than deriving groups from skill keys, and does not add hidden-skill behavior. **Feature request** - **Problem:** Skills currently lack a canonical hierarchical location, making personal skills, project imports, bundled skills, and company-authored skills difficult to traverse and manage at scale. - **Proposed behavior:** Add nested company-scoped folders with stable paths, reserved My/Projects/Bundled roots, subtree queries, safe move/create operations, and a folder-first `/skills` library UI. - **Import behavior:** New project scans file skills under `projects/<project-slug>`; later imports update content without overriding a user-selected folder. - **Alternatives considered:** Tags alone remain useful for cross-cutting classification, but they do not provide canonical location, nesting, reserved namespaces, or stable import placement. - **Roadmap alignment:** Extends the completed Skills Manager and Scheduled Routines capabilities without duplicating an active roadmap item. ## What Changed - Adds `folders` persistence for routine and skill folders, nested canonical paths, parent/slug/system-key fields, migration backfills, and reapply-safe migrations `0174`–`0175` after current master migrations. - Adds company-scoped folder CRUD, cycle/depth/namespace validation, reserved My/Projects/Bundled lifecycle, item moves, subtree filtering, and folder paths on skill results. - Preserves project-import placement: first import files into the project folder, while re-import keeps user-owned placement and stable skill IDs. - Adds the `/skills` folder tree rail, tags facet, breadcrumbs, subfolder browser, move/new-folder dialog, canonical detail location, inline tag editing, and folder-aware Studio creation. - Keeps bundled skills read-only even when their source metadata is incomplete by detecting the reserved Bundled folder and hiding selection/move actions. - Extends routine folder UI and OpenAPI coverage, and adds regression tests across migrations, services, routes, tree helpers, pages, and Studio creation. ## Verification - `pnpm exec vitest run packages/db/src/nested-skill-folders-migration.test.ts server/src/__tests__/folders-routes.test.ts server/src/__tests__/folders-service.test.ts server/src/__tests__/company-skills-service.test.ts server/src/__tests__/routines-service.test.ts ui/src/components/folders/FolderControls.test.tsx ui/src/components/folders/SkillFolderTree.test.tsx ui/src/components/folders/skill-folder-tree.test.ts ui/src/pages/CompanySkills.test.tsx ui/src/pages/Routines.test.tsx ui/src/pages/SkillStudio.test.tsx ui/src/lib/company-skill-routes.test.ts ui/src/lib/skill-create.test.ts` — 13 files, 192 tests passed. - `pnpm exec vitest run ui/src/pages/CompanySkills.test.tsx ui/src/components/folders/SkillFolderTree.test.tsx` — 2 files, 20 tests passed after preserving the existing PR's bundled-skill fixes. - `pnpm -r typecheck` — passed for all workspace packages. - `pnpm test:run` — passed in an isolated CI-like environment with inherited Paperclip runtime identity and static AWS credential variables removed. - `pnpm build` — production build passed for all workspace packages. - Greptile iteration 2 — 5/5 confidence with zero unresolved threads on commit `ff2d67aa71`. - Latest-head GitHub checks — all success, neutral, or skipped; PR is mergeable with a clean merge state. - `pnpm check:token-gates` — reports nine existing `#9627` comment false positives already present on `master`; this PR introduces no new token violation. ## Risks - **Migration/backfill:** `0174` creates the foundation and `0175` adds nested/reserved semantics. Both are ordered after current master migration `0173`, are covered by numbering/safety checks, and are designed to be reapply-safe. - **Reserved namespaces:** My, Projects, and Bundled roots are service-managed. Regression coverage prevents namespace squatting, cross-company folder use, bundled writes, cycles, and excessive depth. - **Behavioral change:** Project scans choose a project folder only on initial creation; existing skills deliberately retain their current folder during refresh. - **UI scope:** The folder rail applies to the Installed library; Catalog retains the discovery-oriented category sidebar. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.5` in Codex CLI, medium reasoning mode; runtime did not expose a context-window value. Used repository/file tools, terminal execution, Git/GitHub operations, test execution, and code editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.18 |
||
|
|
1d2b6af5ac |
fix(tests): stabilize heartbeat cleanup for tsx update (#9573)
## Thinking Path > - Paperclip is an open-source platform for orchestrating AI agents, built on an embedded-Postgres server running a heartbeat loop to advance agent work. > - The server test suite exercises heartbeat liveness escalation and retry scheduling logic against a real embedded database; tests create and tear down full database state across every case. > - Dependabot PR #9480 bumps `tsx` from 4.22.4 to 4.23.1. The new version exposed two fragile teardown patterns in the heartbeat tests that caused failures. > - The first problem: `TRUNCATE TABLE "companies" CASCADE` in the liveness-escalation teardown clashes with FK constraints when child tables (e.g. `heartbeat_run_events`, `issue_tree_hold_members`) hold rows that tsx 4.23.1's changed execution order materialises before the CASCADE runs. > - The second problem: the retry-scheduling test duplicated a 10-line delete block inline at two mid-test reset points; one copy deleted `heartbeat_run_events` after `heartbeat_runs` (wrong FK order) and `activityLog` was deleted twice. > - A third concern was identified during review: several `GET /tool-connections/:connectionId` routes called `assertCompanyAccess` before checking whether the actor has access at all, leaking 403 (existence oracle) instead of 404. This is fixed in this PR. > - This PR updates the three `tsx` version pins to `^4.23.1`, replaces the TRUNCATE with explicit child-to-parent deletes, centralises the retry cleanup into a shared `cleanupRetryFixture()` helper, and adds `hasCompanyAccess` pre-checks before the four affected `assertCompanyAccess` calls in `tool-access.ts`. > - The benefit is CI green on tsx 4.23.1, cleaner non-duplicated teardown code across both test files, and no cross-tenant existence leakage on tool-connection routes. ## Linked Issues or Issue Description Refs #9480 (`tsx` 4.22.4 → 4.23.1 dependabot bump whose CI failures this fixes) ## What Changed - **cli/package.json**, **packages/db/package.json**, **server/package.json**: bump `tsx` dev-dependency range from `^4.22.4` to `^4.23.1` so package manifests agree with the lockfile update landing in #9480. `pnpm-lock.yaml` is left untouched — GitHub Actions owns lockfile regeneration. - **heartbeat-issue-liveness-escalation.test.ts**: replace `TRUNCATE TABLE "companies" CASCADE` with explicit FK-ordered deletes. The new chain adds `heartbeatRunEvents`, `issueTreeHoldMembers`, `agentRuntimeState`, and `companySkills` before their respective parent tables. - **heartbeat-retry-scheduling.test.ts**: extract the repeated teardown block into a `cleanupRetryFixture()` helper; call it from `afterEach` and the two mid-test resets; fix `heartbeatRunEvents` deleted before `heartbeatRuns` (parent-child FK order); remove the duplicate `activityLog` delete. - **server/src/routes/tool-access.ts**: add `hasCompanyAccess` pre-checks before `assertCompanyAccess` on four `GET /tool-connections/:connectionId` and `GET /tool-profiles/:profileId/new-tools` routes. Returns 404 instead of 403 when the actor cannot access the resource, closing the cross-tenant existence oracle. ## Verification ```sh # Focused test run (49 tests, all pass) pnpm exec vitest run \ server/src/__tests__/heartbeat-retry-scheduling.test.ts \ server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts pnpm --filter @paperclipai/server typecheck # pass pnpm -r typecheck # pass pnpm build # pass ``` Full `pnpm test:run` was also attempted: server suite (242 files, 2 243 tests) and UI suite (310 files, 2 536 tests) both passed. A backup-dir assertion in `src/__tests__/onboard.test.ts` failed but is unrelated to this diff — it expects a temp `PAPERCLIP_HOME` but receives the global instance path. ## Risks Low risk. Changes are limited to test teardown logic, dev-dependency version pins, and existence-oracle guard additions on read-only tool-connection routes. No new business logic or production data paths are introduced. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`, 200 k context, tool use, agentic coding) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.17 |
||
|
|
db61cc97d3 |
build(deps-dev): bump tsx from 4.22.4 to 4.23.1 (#9480)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.22.4 to 4.23.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.1</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.0...v4.23.1">4.23.1</a> (2026-07-13)</h2> <h3>Bug Fixes</h3> <ul> <li>support tsImport after global preload (<a href="https://github.com/privatenumber/tsx/commit/8d4ffc24f37b396ca2fe3f251aa92c4919f2c1a4">8d4ffc2</a>)</li> <li><strong>watch:</strong> avoid clearing piped output (<a href="https://github.com/privatenumber/tsx/commit/95d0672e0247a829ae4469daa493212967ea768e">95d0672</a>)</li> <li><strong>watch:</strong> treat script and dependency paths literally (<a href="https://github.com/privatenumber/tsx/commit/79fddde523d3bb7d0af66682ce1265f95113a073">79fddde</a>)</li> </ul> <h3>Performance Improvements</h3> <ul> <li>index transform cache lazily (<a href="https://github.com/privatenumber/tsx/commit/e818ad608159a6fb36fb8a0bd59327fec313323d">e818ad6</a>)</li> <li>load esbuild lazily in CLI (<a href="https://github.com/privatenumber/tsx/commit/d0679381b60a55a9b5863603a4022a81db5d13c8">d067938</a>)</li> <li>map Node TypeScript formats directly (<a href="https://github.com/privatenumber/tsx/commit/cdcc6232a3277fb3028b226958b66c49a6d86c17">cdcc623</a>)</li> <li>use sync module hooks on Node v22.22.3+ (<a href="https://github.com/privatenumber/tsx/commit/f8992f1a50213e11b7ef8ab5121c78e0d2f29384">f8992f1</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.1"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.0</h2> <h1><a href="https://github.com/privatenumber/tsx/compare/v4.22.5...v4.23.0">4.23.0</a> (2026-07-03)</h1> <h3>Bug Fixes</h3> <ul> <li>avoid redundant filesystem probes during module resolution (<a href="https://github.com/privatenumber/tsx/commit/257bbbb7eb2784cad6a3bb7a2d9c9747d28d96ec">257bbbb</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a></li> </ul> <h3>Features</h3> <ul> <li>add multi-scenario startup benchmark suite (<a href="https://github.com/privatenumber/tsx/commit/c178197b104d055fd3431f7448982f3156394d12">c178197</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a> <a href="https://github.com/hi/issues/signal">hi#signal</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/145">privatenumber/tsx#145</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a></li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.0"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.22.5</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.22.5">4.22.5</a> (2026-07-02)</h2> <h3>Bug Fixes</h3> <ul> <li>isolate hook state per async module.register() registration (<a href="https://github.com/privatenumber/tsx/commit/a305f365f0cbcc31a44549dcbb0e63dc2883e96d">a305f36</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.22.5"><code>npm package (@latest dist-tag)</code></a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/79fddde523d3bb7d0af66682ce1265f95113a073"><code>79fddde</code></a> fix(watch): treat script and dependency paths literally</li> <li><a href="https://github.com/privatenumber/tsx/commit/e818ad608159a6fb36fb8a0bd59327fec313323d"><code>e818ad6</code></a> perf: index transform cache lazily</li> <li><a href="https://github.com/privatenumber/tsx/commit/cdcc6232a3277fb3028b226958b66c49a6d86c17"><code>cdcc623</code></a> perf: map Node TypeScript formats directly</li> <li><a href="https://github.com/privatenumber/tsx/commit/d0679381b60a55a9b5863603a4022a81db5d13c8"><code>d067938</code></a> perf: load esbuild lazily in CLI</li> <li><a href="https://github.com/privatenumber/tsx/commit/95d0672e0247a829ae4469daa493212967ea768e"><code>95d0672</code></a> fix(watch): avoid clearing piped output</li> <li><a href="https://github.com/privatenumber/tsx/commit/6fd4607e8a99d1efe27f185f749c659138f00ece"><code>6fd4607</code></a> docs: add per-page metadata</li> <li><a href="https://github.com/privatenumber/tsx/commit/f4176d8c6329a12205ed9b8c582e559cafc45018"><code>f4176d8</code></a> docs: generate sitemap</li> <li><a href="https://github.com/privatenumber/tsx/commit/8d4ffc24f37b396ca2fe3f251aa92c4919f2c1a4"><code>8d4ffc2</code></a> fix: support tsImport after global preload</li> <li><a href="https://github.com/privatenumber/tsx/commit/f0e89b244c98849dc5f9483fa33aaaf983ede18f"><code>f0e89b2</code></a> docs: document Node's public type-stripping API vs internal loader path</li> <li><a href="https://github.com/privatenumber/tsx/commit/f8992f1a50213e11b7ef8ab5121c78e0d2f29384"><code>f8992f1</code></a> perf: use sync module hooks on Node v22.22.3+</li> <li>Additional commits viewable in <a href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.23.1">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.716.0-canary.16 |
||
|
|
c4da1c925f |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.52.0 to 0.59.0 (#9484)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.52.0 to 0.59.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.59.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.1...v0.59.0">0.59.0</a> (2026-07-13)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> bump nanoid from 3.3.15 to 3.3.16 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e67dacdcac629b00e5bdcda94b3de57459696b1e">e67dacd</a>)</li> <li><strong>deps:</strong> bump <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.207 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7f5b8fe768507afa28cdffe5cc3c054c56306ad">c7f5b8f</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Add subagent parent tool use attribution (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9cd48c597af68f630774ec00bc3adc85f6b0fd4b">9cd48c5</a>)</li> <li>forward result text when the turn emitted no assistant message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61ae8609d4343da6758c2376c6cf684c7fb0956a">61ae860</a>)</li> <li>hold a turn open while its background subagents are still live (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7a70f82739e085014cad878f08513cdef7b7fe16">7a70f82</a>)</li> <li>Refine tool calls from streamed input (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2c19974d5b80d3178b35e80a4bb0a1050588aab9">2c19974</a>)</li> <li>Seed context window from SDK usage report (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3ba2d367a282fe3053a73804a416a435cb01ee08">3ba2d36</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/596">#596</a></li> <li>Skip synthetic login messages on replay (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7dff3cf7eeb03cb993493efafa96344f18de81b">c7dff3c</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/863">#863</a></li> </ul> <h2>v0.58.1</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.0...v0.58.1">0.58.1</a> (2026-07-09)</h2> <h3>Bug Fixes</h3> <ul> <li>Use valid npm version for publish (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/855">#855</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b366d8e9ad55e9237bf2d618cfc99b57fac570d">8b366d8</a>)</li> </ul> <h2>v0.58.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.57.0...v0.58.0">0.58.0</a> (2026-07-09)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> update to <code>@anthropic-ai/claude-agent-sdk</code> 0.3.205 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/854">#854</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f664ced76df00674d88aedd018a4261da3bc3f3a">f664ced</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Preserve live model on resumed sessions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/848">#848</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f3d8ae3eb389ee9367f48c8762562a55280ade0a">f3d8ae3</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/845">#845</a></li> <li>Report usage for cancelled active turns (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/846">#846</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b03318f59c6165146c8d8e036227f35c3abd0ff4">b03318f</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/844">#844</a></li> <li>tolerate missing text in streamed thinking chunks (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/852">#852</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e944ceddea79950d6bfb6339266862507c4b964f">e944ced</a>)</li> <li>Use SDK guards for elicitation validation (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/850">#850</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/32b93501d7a35a3b405b5022b4caaad492eed11c">32b9350</a>)</li> </ul> <h2>v0.57.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.56.0...v0.57.0">0.57.0</a> (2026-07-07)</h2> <h3>Features</h3> <ul> <li>Update <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.202 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/843">#843</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1612c078953c134e12368a2d968a9bcf14d31b55">1612c07</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.1...v0.59.0">0.59.0</a> (2026-07-13)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> bump nanoid from 3.3.15 to 3.3.16 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e67dacdcac629b00e5bdcda94b3de57459696b1e">e67dacd</a>)</li> <li><strong>deps:</strong> bump <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.207 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7f5b8fe768507afa28cdffe5cc3c054c56306ad">c7f5b8f</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Add subagent parent tool use attribution (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9cd48c597af68f630774ec00bc3adc85f6b0fd4b">9cd48c5</a>)</li> <li>forward result text when the turn emitted no assistant message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61ae8609d4343da6758c2376c6cf684c7fb0956a">61ae860</a>)</li> <li>hold a turn open while its background subagents are still live (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7a70f82739e085014cad878f08513cdef7b7fe16">7a70f82</a>)</li> <li>Refine tool calls from streamed input (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2c19974d5b80d3178b35e80a4bb0a1050588aab9">2c19974</a>)</li> <li>Seed context window from SDK usage report (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3ba2d367a282fe3053a73804a416a435cb01ee08">3ba2d36</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/596">#596</a></li> <li>Skip synthetic login messages on replay (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7dff3cf7eeb03cb993493efafa96344f18de81b">c7dff3c</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/863">#863</a></li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.58.0...v0.58.1">0.58.1</a> (2026-07-09)</h2> <h3>Bug Fixes</h3> <ul> <li>Use valid npm version for publish (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/855">#855</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b366d8e9ad55e9237bf2d618cfc99b57fac570d">8b366d8</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.57.0...v0.58.0">0.58.0</a> (2026-07-09)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> update to <code>@anthropic-ai/claude-agent-sdk</code> 0.3.205 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/854">#854</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f664ced76df00674d88aedd018a4261da3bc3f3a">f664ced</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Preserve live model on resumed sessions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/848">#848</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f3d8ae3eb389ee9367f48c8762562a55280ade0a">f3d8ae3</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/845">#845</a></li> <li>Report usage for cancelled active turns (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/846">#846</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b03318f59c6165146c8d8e036227f35c3abd0ff4">b03318f</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/844">#844</a></li> <li>tolerate missing text in streamed thinking chunks (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/852">#852</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e944ceddea79950d6bfb6339266862507c4b964f">e944ced</a>)</li> <li>Use SDK guards for elicitation validation (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/850">#850</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/32b93501d7a35a3b405b5022b4caaad492eed11c">32b9350</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.56.0...v0.57.0">0.57.0</a> (2026-07-07)</h2> <h3>Features</h3> <ul> <li>Update <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.202 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/843">#843</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1612c078953c134e12368a2d968a9bcf14d31b55">1612c07</a>)</li> </ul> <h3>Bug Fixes</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/30b7c06f7640fb6a0530ba18f85e26fe2bc08882"><code>30b7c06</code></a> chore(main): release 0.59.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/860">#860</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7f5b8fe768507afa28cdffe5cc3c054c56306ad"><code>c7f5b8f</code></a> feat(deps): bump <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.207 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/874">#874</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7a70f82739e085014cad878f08513cdef7b7fe16"><code>7a70f82</code></a> fix: hold a turn open while its background subagents are still live (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/870">#870</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e67dacdcac629b00e5bdcda94b3de57459696b1e"><code>e67dacd</code></a> feat(deps-dev): bump nanoid from 3.3.15 to 3.3.16 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/875">#875</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61ae8609d4343da6758c2376c6cf684c7fb0956a"><code>61ae860</code></a> fix: forward result text when the turn emitted no assistant message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/858">#858</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2c19974d5b80d3178b35e80a4bb0a1050588aab9"><code>2c19974</code></a> fix: Refine tool calls from streamed input (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/867">#867</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c7dff3cf7eeb03cb993493efafa96344f18de81b"><code>c7dff3c</code></a> fix: Skip synthetic login messages on replay (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/869">#869</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3ba2d367a282fe3053a73804a416a435cb01ee08"><code>3ba2d36</code></a> fix: Seed context window from SDK usage report (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/868">#868</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f4e750decf6057cf984d15afe3a25a42f16f3911"><code>f4e750d</code></a> feat(deps): bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/862">#862</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9cd48c597af68f630774ec00bc3adc85f6b0fd4b"><code>9cd48c5</code></a> fix: Add subagent parent tool use attribution (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/859">#859</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.52.0...v0.59.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d62c5adb7a |
fix(ui): stop recreating markdown mention observers (#3809)
Fixes #3759 ## Thinking Path > - Paperclip orchestrates AI agents and issue workflows, so the comment composer has to stay stable under normal typing. > - The affected subsystem is the shared markdown comment editor used across issue and workflow surfaces. > - That editor decorates mention links after Lexical updates the editable DOM. > - The current mention decoration effect recreates its `MutationObserver` whenever `value` changes, which happens on every keystroke. > - That observer also reacts to the DOM mutations produced by mention decoration itself, creating unnecessary observer churn in Chrome. > - This pull request keeps one observer instance alive, batches decoration work into `requestAnimationFrame`, and disconnects the observer while decoration writes run. > - The benefit is lower observer churn while preserving the existing mention chip behavior. ## What Changed - Removed `value` from the mention-decoration observer effect dependency list so the observer is not recreated on every external value update. - Batched mention decoration with `requestAnimationFrame` and temporarily disconnected the observer while DOM decorations are applied to avoid self-triggered feedback loops. - Added a regression test that verifies external value changes do not recreate the mention decoration observer. ## Verification - `pnpm install --frozen-lockfile` - `pnpm --filter @paperclipai/ui exec tsc --noEmit` - `pnpm --filter @paperclipai/ui exec vitest run src/components/MarkdownEditor.test.tsx` currently fails before test collection on the existing repo baseline with `TypeError: undefined is not an object (evaluating 'z.string')` from `packages/shared/src/adapter-type.ts` ## Risks - Low risk. The change is scoped to the mention decoration observer lifecycle and keeps the existing decoration logic intact. - The main behavior change is deferring decoration to the next animation frame instead of running immediately on every observed mutation. ## Model Used - OpenAI Codex, GPT-5-based coding agent with local tool use in the Codex CLI environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.716.0-canary.15 |
||
|
|
8b04147ca4 |
perf(ui): stop polling the event-sourced company live-runs list (#9701)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its web UI keeps live views fresh with a live-events websocket plus React Query, and #9627 began replacing polling with event-sourcing (pushing data into the cache) > - After the churn fixes (#9624) and #9627 shipped, a long-lived tab's memory footprint was still climbing, so I re-profiled with the Chrome DevTools MCP > - The dominant remaining churn is React Query re-arming a polled query's `refetchInterval` timer on **every** observer notification — and our live-event handlers `setQueryData(liveRuns(companyId))` on nearly every event, so each pushed update re-arms every `liveRuns` observer's timer (the sidebar is always mounted) > - #9627 event-sourced the live-runs data but left the now-redundant `refetchInterval` in place, so we did half the fix — the poll is pure waste and the thing re-arming timers > - This pull request removes `refetchInterval` from the event-sourced company live-runs queries so the frequent cache writes have no timer to re-arm > - The benefit is that the steady-state timer churn on the most-observed resource collapses, so the off-heap footprint stops climbing ## Linked Issues or Issue Description No public GitHub issue exists; describing inline per CONTRIBUTING.md → "Link Issues or Describe Them In-PR", following the bug report template. Continues #9569 / #9624 / #9627. **What happened?** With the earlier fixes deployed, a browser tab left open on the app kept growing its memory footprint. MCP profiling showed ~100+ `setInterval` create/clear cycles per 5s on an aged tab (vs ~16 fresh), all from React Query's refetch-interval timers being re-armed on every `setQueryData` to the frequently-written `liveRuns(companyId)` query. **Expected behavior** A resource whose data is pushed (event-sourced) should not also poll; cache writes should not repeatedly re-arm interval timers. Idle tabs should hold a bounded footprint. **Steps to reproduce** Open a tab with agents streaming, leave it open, and instrument `setInterval`/`clearInterval`: the churn rate climbs and traces to `QueryObserver.updateTimers` (`refetchInterval`) for `liveRuns`, re-armed by every live-event cache write. **Paperclip version or commit** Branch `perf/drop-live-runs-refetch-interval`, off `master` (after #9627). **Deployment mode** Local dev (`pnpm dev`), web UI. Core UI live-updates plumbing; not adapter-specific. ## What Changed - Set `refetchInterval: false` on every site that polls the **plain** `queryKeys.liveRuns(companyId)` query (event-sourced by #9627): `Sidebar`, `SidebarAgents`, `Issues`, `Inbox`, `IssueDetail` (companyLiveRuns), `ProjectDetail` (×2), `Routines`, `ExecutionWorkspaceDetail`, `bridge-init`. - Removed the now-unused `useVisibilityRefetchInterval` interval vars/imports in `Issues`, `Inbox`, `IssueDetail`. - **Left variant-key sites polling on purpose** — `Agents` page (`[...liveRuns, "agents-page"]`) and `ActiveAgentsPanel` (`[...liveRuns, scope, …]`) are NOT event-sourced by #9627 (different exact cache key), so dropping their poll would make them stale. Those are a later phase. - Freshness for the converted queries now comes from event-sourcing (#9627) + its reconnect reconcile; the initial mount fetch and cross-tab publish (`usePublishSharedQueryData`) still happen. ## Verification - MCP profiling identified the churn: the single churning callback is React Query's `refetchInterval` timer, re-armed by `setQueryData(liveRuns)` on live events. - `vitest`: all affected suites pass (`Sidebar`, `SidebarAgents`, `Issues`, `Inbox`, `IssueDetail`, `ProjectDetail`, `Routines`, `ExecutionWorkspaceDetail`, `LiveUpdatesProvider`) — 153 tests. - Updated two `SidebarAgents` linger-window tests: they advanced fake timers to the *exact* linger-expiry boundary and had relied on poll-induced re-renders to flush. The linger self-schedules its own `setTimeout`, so the tests now cross the boundary with a small margin + an explicit flush (no product change). - `tsc -b` clean. - End-to-end footprint reduction should be re-measured against a rebuilt bundle with the same instrumentation. ## Risks Low, client-only. - `liveRuns(companyId)` freshness now depends entirely on event-sourcing + reconnect reconcile (both from #9627). If an event path is missed, the reconnect handler refetches once; durable replay is a planned later phase. - Variant-key run lists (Agents page, ActiveAgentsPanel) are unchanged and still poll, so they don't regress. - Issue-scoped run queries (`issues.liveRuns/activeRun/runs`) are **not** touched here — they aren't event-sourced yet and are a separate phase. ## Model Used - **Provider:** Anthropic, via the Claude Code CLI. - **Model:** Claude Opus 4.8 (`claude-opus-4-8`). - **Reasoning mode:** Extended thinking enabled. - **Capabilities used:** tool use (shell, file editing), and the Chrome DevTools MCP to re-profile the live instance and pinpoint the `refetchInterval` timer churn. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work (a perf/plumbing change) - [x] I have searched GitHub for duplicate or related PRs and linked them above (continues #9627; no duplicates) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [ ] I have updated relevant documentation to reflect my changes (N/A — no user-facing docs; rationale documented inline) - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code)canary/v2026.716.0-canary.14 |
||
|
|
5a5c918705 |
fix(acpx): configure Codex models at startup (#9700)
Move Codex ACPX model, reasoning effort, and fast-mode settings into CODEX_CONFIG startup config so arbitrary model IDs avoid ACP session picker validation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9a92124c63 |
fix(codex-local): raise default output-inactivity timeout to 30m (#9699)
## Thinking Path > - Paperclip is the open source control plane for managing AI-agent companies. > - Agent adapters are responsible for launching, observing, and terminating provider processes safely. > - The local Codex adapter uses an output-inactivity monitor to stop genuinely hung child processes. > - The existing seven-minute default also stopped legitimate long-running tasks that emitted no output while tests or remote checks were still running. > - Explicit per-agent timeout overrides already provide configuration flexibility, so the smallest safe correction is to increase only the default window. > - This pull request raises the default to 30 minutes, retains the existing termination behavior, and adds a regression assertion for the new value. > - The benefit is fewer unnecessary Codex restarts while still bounding genuinely silent processes well below the platform safety limit. ## Linked Issues or Issue Description **Pre-submission checklist** - Searched open and closed GitHub issues and pull requests; no duplicate fix exists. The original inactivity monitor was introduced in #5017. - Reproduces on the current `master` implementation. - The termination originates in Paperclip's Codex adapter inactivity monitor rather than the model provider or local configuration. **What happened?** Long-running `codex_local` tasks were terminated after seven minutes without stdout or stderr, even when the child process was still performing legitimate work such as a quiet test suite or waiting for remote checks. **Expected behavior** The default inactivity window should tolerate common long-running quiet tasks while continuing to terminate processes that remain silent for an extended period. **Steps to reproduce** 1. Start a `codex_local` run with the default `outputInactivityTimeoutMs` configuration. 2. Have the child process perform legitimate work without emitting stdout or stderr for more than seven minutes. 3. Observe the adapter terminate the process at the old default threshold. **Version / deployment** Current `master`, built from source in local/self-hosted deployments, using the Codex adapter. This behavior is not database- or access-context-specific. ## What Changed - Raise `DEFAULT_CODEX_OUTPUT_INACTIVITY_TIMEOUT_MS` from seven minutes to 30 minutes. - Update the adapter configuration documentation to state the new default. - Add a focused regression assertion that pins the default to 30 minutes. ## Verification - `pnpm vitest run packages/adapters/codex-local/src/server/output-inactivity-monitor.test.ts` — 15 tests passed. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed. ## Risks - Low risk: only the fallback default changes; explicit positive timeout values and `null` disablement retain their existing behavior. - A genuinely silent Codex process now remains alive up to 23 minutes longer before the same SIGTERM/SIGKILL cleanup path runs. - No schema, API, migration, UI, or workflow changes are included. > This is a focused bug fix and does not overlap with planned core feature work in `ROADMAP.md`. ## Model Used - Anthropic Claude Fable 5 via the `claude_local` adapter produced the initial investigation and implementation with repository/tool access. - OpenAI Codex via the Codex CLI prepared the PR, added the focused regression assertion, and ran verification; the runtime did not expose a more specific model ID or context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] I have used a public-friendly branch name without internal tracker identifiers (execution-workspace exception: this branch is runtime-managed and cannot be renamed) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge The source branch is execution-workspace managed and cannot be renamed during this run; the PR title and body intentionally contain no internal tracker references. Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.13 |
||
|
|
6ec059ab4e |
fix(server): suppress stale handoff alarms during live continuation (#9695)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The control plane records a successful-run handoff when productive work ends without a durable next-step disposition > - That handoff state was derived only from the latest activity event, without checking whether a corrective run or wake was currently alive > - As a result, actively progressing issues could still show a high-severity missing-disposition alarm and blocked-inbox row > - The same stale required event could also remain indefinitely when a later successful run correctly skipped recovery because another valid continuation path already existed > - This pull request makes the derived state liveness-aware, suppresses attention only while the live path exists, and resolves stale required events on valid-path skips > - The benefit is that productive work stays calm while genuine stalls still resurface automatically when liveness disappears ## Linked Issues or Issue Description - **Bug:** An issue whose latest successful-run handoff event is `required` continues to report a missing disposition even while a heartbeat run, scheduled retry, or queued/deferred/claimed wake is actively targeting that issue. - **Expected behavior:** The API should expose current continuation liveness, the blocked inbox should suppress the alarm only while that path remains live, and a later successful run that skips recovery because a valid path exists should durably resolve the stale event. - **Related but distinct:** #9370 changes disposition freshness at detection time; #8748 adds an explicit policy opt-out. This PR preserves detection/escalation policy and fixes read-time/current-liveness state. ## What Changed - Extended `SuccessfulRunHandoffState` with `hasLiveContinuation` and optional `liveRunId` evidence. - Added bounded liveness hydration for required handoff states using active heartbeat-run and wake-request signals. - Suppressed `missing_disposition` blocked-inbox rows only while a run, scheduled retry, or live wake targets the issue. - Added durable `issue.successful_run_handoff_resolved` logging when handoff detection skips because another valid continuation path owns the next action. - Added focused regressions for live/absent derived state, self-healing attention suppression, valid-path skip classification, and resolved-event logging. - Updated UI normalization and fixtures for the shared contract without changing rendering behavior. ## Verification - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm vitest run server/src/services/recovery/successful-run-handoff.test.ts server/src/__tests__/issue-list-assignee-filter-routes.test.ts server/src/__tests__/issue-blocker-attention.test.ts` — 56 passed - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "queues one finish-handoff wake when a successful run leaves in-progress work without a next action"` — 1 passed - `git diff --check` ## Risks - Low risk: no schema or migration changes, and detection, bounded correction attempts, and escalation behavior are unchanged. - Liveness lookups are limited to issues whose latest handoff state is `required`; blocked-inbox suppression reuses rows already loaded by that query path. - Suppression is read-time and self-healing: when the run or wake stops, the alarm returns on the next fetch. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.4`, tool-enabled software-engineering workflow with repository, shell, test, Git, GitHub, and Paperclip control-plane access. Context-window size is not exposed by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.12 |
||
|
|
a04a77c9d3 |
feat(authz): govern agent inbox archive access (#9658)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - The inbox subsystem must let agents act for a responsible user without silently granting access to every company user's tasks > - Existing authorization had no inbox-specific action, target-user scope, or per-user agent policy > - Inbox archive data also needs company-safe ownership and replay-safe schema changes before API mutations can rely on it > - This pull request adds the database policy foundation and a fail-closed `inbox:manage` authorization decision > - The benefit is a least-privilege core for later inbox archive endpoints, including explicit cross-user grants and low-trust denial ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`packages/db`, `packages/shared`, and `server`). ### Problem or motivation Agents need to manage inbox state for the user responsible for their run, but the control plane lacks an inbox-specific permission model and user-targeted grant scope. A generic mutation path would risk cross-user access or inconsistent policy enforcement. ### Proposed solution Add inbox archive ownership and per-user agent policies, introduce `inbox:manage`, and evaluate responsible-user defaults, disabled/allowlist policies, active membership, low-trust presets, and scoped cross-user grants in one authorization decision. ### Alternatives considered Reusing generic issue mutation permissions was rejected because it cannot express user-targeted inbox scope. Requiring grants for all self-user access was rejected because it would make the responsible-user path closed by default instead of using the requested per-user policy model. ### Roadmap alignment `ROADMAP.md` contains no overlapping inbox archive or inbox authorization item; this is incremental control-plane authorization work. ### Additional context This PR provides the authorization and schema foundation. Route and UI behavior can build on this decision without duplicating access-control rules. ## What Changed - Builds on the merged migration `0172_inbox_archive_agent_policies` (#9654) for company/user-scoped inbox archives and per-user agent policy rows. - Added replay-safe migration `0173_inbox_policy_agent_cleanup` with a GIN allowlist index and GIN-backed database cleanup that removes deleted agent IDs from policy allowlists. - Added Drizzle schema exports for inbox agent policies and responsible-user ownership on inbox archives. - Added the shared `inbox:manage` permission key and `scope.userIds` evaluation for user-targeted grants. - Added fail-closed inbox authorization for unresolved targets, inactive memberships, low-trust agents, disabled policies, allowlist misses, and ungranted cross-user access. - Added migration replay coverage and the full inbox authorization decision matrix. ## Verification - `pnpm exec vitest run packages/db/src/inbox-archive-agent-policies-migration.test.ts server/src/__tests__/authorization-service.test.ts` — 50 tests passed. - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check origin/master...HEAD` ## Risks - The merged `0172` migration changed inbox archive uniqueness from agent-owned to responsible-user-owned rows; `0173` is additive (index + cleanup trigger) and idempotent, and replay coverage verifies both remain safe for databases that already applied an earlier form. - `scope.userIds` uses the existing JSON grant-scope parser, so malformed privileged grant payloads continue to fail through the shared parsing behavior rather than a dedicated schema. - Cross-user grants intentionally act as board-admin overrides; responsible-user default access remains bounded by disabled and allowlist policies. - The authorization action is not yet wired to public mutation routes, limiting immediate behavioral impact while establishing the contract those routes must use. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`, high reasoning effort, CLI tool use, code execution, GitHub CLI, and Paperclip control-plane integration. Context window size is not exposed by the configured adapter. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c65ab09d9f |
fix(recovery): wait for provider quota resets (#9635)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and keep assigned work moving safely. > - The recovery subsystem decides whether a failed agent run should retry, wait, block for configuration, or escalate to another owner. > - Provider usage-limit failures currently arrive as generic `adapter_failed` results, so stranded-work reconciliation can create takeover recovery even when the provider states that capacity will reset later. > - Credential and model lookup failures are also configuration problems, not evidence that another agent should take over the task. > - This pull request classifies those failure families at recovery time and persists the classification on the run. > - Quota failures now schedule a monitor for the original assignee at the parsed reset time, or after a bounded default backoff when no reset time is available. > - The benefit is that transient provider capacity waits no longer wake recovery owners, while configuration failures stop with an actionable classification. ## Linked Issues or Issue Description No public GitHub issue exists for this exact change. **What happened?** When an assigned issue's latest run failed with a provider usage-limit message such as "try again at 12:00 AM (UTC)," recovery treated the run as generic `adapter_failed` work and could create a takeover action. Missing credentials and `model_not_found` failures followed the same generic path. **Expected behavior:** Provider quota failures should keep the original assignee and schedule a monitor for the reset time, without creating recovery work or immediately waking another owner. Missing credentials and model lookup failures should be classified as `configuration_incomplete` and blocked with the configuration fix recorded. **Steps to reproduce:** 1. Assign and start an issue for an agent. 2. Record a failed heartbeat run with `errorCode: adapter_failed` and a provider quota/reset message. 3. Run stranded assigned-issue reconciliation. 4. Observe that the old behavior routes the issue through generic recovery instead of waiting for provider capacity. Reproduced on `master` at `9af96461d`. This is a core recovery bug, not adapter-specific, and applies to built-from-source deployments with either embedded PGlite or Postgres. Related work checked: #9288 adds adapter-side Claude provider-limit classification; #5392 suppresses some recovery creation for quota-class errors; #9634 is a broader recovery-routing change with overlapping provider-quota behavior. This PR is the narrow recovery-service fix with focused parsed-reset, fallback-backoff, zero-takeover, and configuration-failure coverage. ## What Changed - Added conservative recovery-time classification for provider quota, missing-credential, and model-not-found adapter failures. - Parsed provider reset timestamps with a default one-hour backoff when no usable reset time is present. - Persisted `provider_quota` or `configuration_incomplete` metadata on the failed heartbeat run. - Scheduled quota monitors for the active issue owner, including the current review participant, without creating recovery actions or enqueueing takeover wakes. - Routed configuration failures to blocked recovery with actionable evidence instead of a takeover. - Added unit and embedded-database regression coverage for parsed/fallback quota timing, zero CTO/recovery wake behavior, and configuration classification. ## Verification - `pnpm exec vitest run server/src/services/recovery/provider-failure-classification.test.ts server/src/__tests__/issue-recovery-actions.test.ts server/src/__tests__/issue-monitor-scheduler.test.ts server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts` — 4 files passed, 66 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Recovery behavior changes for text-matched adapter failures; matching is intentionally conservative, and unmatched failures retain the existing generic recovery path. - Provider reset strings do not always include a date or timezone; parsing chooses the next future matching time and falls back to a one-hour wait when the timestamp is unusable. - This overlaps the provider-quota portion of broader recovery-routing PR #9634, so only one implementation should land if both remain open. - No schema, migration, API contract, or UI changes are included. No documentation update is needed because this corrects internal recovery behavior without changing operator commands or configuration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model `gpt-5.4`, medium reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
85404b46c5 |
fix(server): throttle serial recovery repeats (#9651)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies. > - Its recovery services create productivity reviews and liveness escalations when work stops making progress. > - Existing uniqueness guards prevent concurrent duplicates, but terminal recovery tasks can still be recreated serially without enough time for conditions to change. > - That creates noisy review churn for persistently stalled issues and immediate liveness re-escalation after a recovery task closes. > - This pull request adds bounded, configurable cooldown and no-action suppression behavior to those two recovery paths. > - The benefit is quieter recovery automation that still resumes automatically after source activity or cooldown expiry. ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I can reproduce this behavior on `master`. - [x] I have confirmed the behavior originates in Paperclip core recovery orchestration, not an adapter, provider, or local configuration. ### What happened? Recovery reconciliation can serially recreate equivalent system-origin tasks after previous tasks become terminal. Productivity reviews allowed multiple creations for the same source issue within a rolling day, and a closed liveness escalation could be recreated immediately for the same incident or recovery leaf. ### Expected behavior Productivity review creation should be limited to once per rolling 24 hours, repeated completed reviews that produced no source action should eventually suppress further creation until activity resumes, and recently terminal liveness escalations should receive a short cooldown before recreation. ### Steps to reproduce 1. Create a stalled assigned issue that meets productivity-review eligibility. 2. Complete repeated productivity-review tasks without adding source-issue activity, then reconcile again within 24 hours. 3. Create and close a liveness escalation for a blocked issue graph, then immediately reconcile the same graph. 4. Observe that equivalent system tasks can be recreated serially without a meaningful state change. ### Paperclip version or commit `5588ddf68175eea448f9d19677b97d7393c38c3d` (`master` when reproduced) ### Deployment mode Local dev (`pnpm dev`) ### Installation method Built from source (`pnpm dev` / `pnpm build`) ### Agent adapter(s) involved - [x] Not adapter-specific (core bug) ### Database mode Embedded PGlite (default — `DATABASE_URL` unset) ### Access context Unclear / not applicable ### Node.js version Current repository-supported Node.js runtime. ### Operating system Linux development environment. ### Relevant logs or output No error is emitted; the bug is repeated task creation visible in persisted issue history. ### Relevant config (if applicable) No special configuration is required. ### Additional context The concurrent/open-task uniqueness guards work as designed; this change targets serial repeats after matching tasks become terminal. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Tightened the productivity-review creation cap to one review per source issue in a rolling 24-hour window. - Added configurable suppression after three consecutive completed reviews with no source-issue activity, with automatic reset when source activity occurs. - Added a configurable one-hour default cooldown for matching terminal liveness escalations. - Exposed the liveness reconciliation clock/cooldown inputs for deterministic orchestration tests. - Added focused tests for daily enforcement, no-action suppression and reset, and cooldown expiry. ## Verification - `pnpm exec vitest run server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts` — 36 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - Confirm the focused tests demonstrate creation after source activity and after the liveness cooldown expires. ## Risks - Low-to-moderate behavioral risk: recovery tasks intentionally appear less often, so overly aggressive thresholds could delay intervention for a persistently stalled issue. - Thresholds are configurable through reconciliation inputs, and source activity resets productivity-review suppression. - No database migration, public API change, telemetry contract change, or UI behavior change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.3-Codex, with reasoning, repository/terminal tool use, code execution, and test execution. The runtime did not expose a reliable context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.11 |
||
|
|
da549123cc |
feat(db): add inbox archive agent policies (#9654)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The inbox tracks issue visibility separately for each responsible user > - Existing inbox archive records only identify the user whose inbox changed, not the actor that made the change > - Agent-managed inbox cleanup needs durable agent and heartbeat-run attribution for auditability > - Users also need a minimal core policy that can remain open, restrict access to an allowlist, or disable agent inbox management > - This pull request adds those database contracts without changing routes or services > - The benefit is a company-scoped, auditable foundation for later agent inbox-management APIs ## Linked Issues or Issue Description ### Subsystem affected `packages/db` — Drizzle schema and migrations. ### Problem or motivation Agent workflows such as PR gardening can reduce inbox noise after work is complete, but existing inbox archive records only identify the user whose inbox changed. They cannot retain which agent acted or which heartbeat run performed the action, and there is no user-specific policy controlling whether agents may manage that inbox. ### Proposed solution Add database-only contracts for agent-managed inbox archiving: actor type, agent ID, and heartbeat-run attribution on archive rows, plus one company-scoped policy per user with `open`, `allowlist`, or `disabled` mode. Existing archive rows retain `user` attribution by default. Route, service, and UI behavior are intentionally deferred. ### Alternatives considered - Store attribution only in activity logs: rejected because archive state needs durable, directly queryable attribution. - Add richer per-agent rules immediately: rejected because advanced policy rules belong in a later or enterprise layer; the core table stays deliberately minimal. - Implement APIs in the same change: rejected to keep this migration-focused PR independently reviewable and safe to deploy. ### Roadmap alignment `ROADMAP.md` does not currently list this capability. This PR establishes only the database foundation and does not overlap a listed roadmap item. ### Additional context Expected behavior is that legacy archive rows upgrade without backfill, new archive rows can reference an agent and heartbeat run, and each company/user pair has at most one agent policy. ## What Changed - Added actor type, agent ID, and heartbeat run ID attribution columns to `issue_inbox_archives` with enum checks and `SET NULL` foreign keys. - Added the company-scoped `user_inbox_agent_policies` table with mode validation, JSONB agent allowlists, timestamps, and unique company/user ownership. - Added migration `0172_inbox_archive_agent_policies.sql` using idempotent DDL for existing installations. - Added an embedded PostgreSQL migration test covering legacy rows, attribution round-trips, policy JSON, and policy uniqueness. ## Verification - `pnpm --filter @paperclipai/db exec vitest run src/inbox-archive-agent-policies-migration.test.ts` - `pnpm --filter @paperclipai/db typecheck` ## Risks - Low migration risk: adding the non-null actor type uses a constant `user` default so legacy rows upgrade without a data backfill. - Agent and run deletions clear attribution foreign keys by design; the actor type remains available for audit interpretation. - No API behavior changes are included, so the new contracts remain unused until follow-up service work lands. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4 with reasoning, repository tools, shell execution, and test execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
054b076b58 |
fix(ui): keep inbox hover and j/k keyboard selection in sync across list reshapes (#9680)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Inbox is the primary triage surface; it supports both mouse hover and `j`/`k` keyboard navigation over a flat, expandable list of rows > - Hover and keyboard selection are meant to be a single shared "cursor": hovering a row and then pressing `j`/`k` should continue from the hovered row, not jump elsewhere > - A prior hover-perf rewrite moved the hovered row into a numeric `hoveredIndexRef` and nulled it whenever the list reshaped; because the inbox polls constantly, any refresh between hovering and pressing a key dropped the hovered index > - With the hovered index dropped, the next keypress fell back to selection index `0`, stranding the cursor at the top of the list instead of continuing from the hovered row > - This pull request tracks the hovered row's stable identity (a nav key) alongside the numeric index and re-anchors it by key across reshapes, mirroring the existing keyboard-selection reconciliation > - The benefit is that mouse hover and keyboard navigation stay in sync exactly as intended, even while the inbox is polling ## Linked Issues or Issue Description **Bug report** - **What happened:** In the Inbox, hovering a row with the mouse and then pressing `j`/`k` (or another keyboard shortcut) does not continue from the hovered row. If the list refreshes (which happens on the inbox's constant polling) in the moment between hovering and pressing a key, keyboard selection snaps back to the top row instead. - **Expected behavior:** The keyboard cursor should be in sync with the hovered row — pressing `j`/`k` after hovering should move relative to the row the mouse is over. - **Steps to reproduce:** 1. Open the Inbox with several rows. 2. Hover the mouse over a row partway down the list. 3. Wait for (or trigger) a background poll/refresh of the list. 4. Press `j` or `k`. 5. Observe selection jumps to the top of the list instead of continuing from the hovered row. - **Root cause:** The `[flatNavItems]` effect nulled `hoveredIndexRef` on every list reshape. Since hover also clears the keyboard selection band to `-1`, the fallback selection index resolved to `0`. ## What Changed - Hoisted `navEntryKey` to module scope so it can compute a stable, index-independent identity for a nav row from both the hover handler and the reshape effect. - Added `hoveredNavKeyRef` to track the hovered row's stable key alongside the existing numeric `hoveredIndexRef`, set whenever the pointer selects a row. - On list reshape, re-anchor the hovered index by key (find the row with the same key) instead of unconditionally dropping it; only drop the hover when the row is actually gone. This mirrors the existing `selectedIndex` key-based reconciliation. - Added a unit test that hovers a row, reshapes the list via a simulated poll, presses `j`, and asserts selection continues from the hovered row. ## Verification - `pnpm --filter ./ui exec vitest run src/pages/Inbox.test.tsx` → 15/15 passing, including the new hover→`j`/`k` sync test and the existing keyboard-nav tests. - The new test explicitly covers the reshape-during-hover path that the prior test suite had deferred to live/e2e verification. ## Risks - Low risk. The change is confined to the Inbox's in-memory hover/selection bookkeeping (two refs and one effect); it adds no new renders (hover still paints via CSS `:hover`) and touches no data fetching, routing, or persistence. Behavior is unchanged when the list does not reshape; when it does, the hover now follows the same row instead of being dropped. ## Model Used - Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, used with extended thinking and tool use (file edits, running the UI test suite locally). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.716.0-canary.10 |
||
|
|
263316609e |
fix(server): avoid hot restart shutdown deadlock (#9670)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents and their work > - The server coordinates agent heartbeats and preserves eligible live runs during a hot restart > - Shutdown previously waited for all heartbeat scheduler work before capturing the hot-restart snapshot > - A deployment heartbeat can itself be in that scheduler set while waiting for the restart, creating a circular wait > - The missing snapshot prevents startup from classifying and adopting the still-running agent process > - This pull request captures the snapshot first and skips scheduler/drain waits only for an eligible hot restart > - The benefit is a single SIGTERM can restart the server without losing eligible live agent runs ## Linked Issues or Issue Description - **Preflight:** Searched open and closed PRs for the hot-restart shutdown deadlock; no duplicate found. Reproduced on `master` and confirmed this is core Paperclip behavior. - **What happened:** During a hot restart initiated by a running deployment heartbeat, the SIGTERM handler waited for `heartbeatSchedulerInFlight` before calling `prepareHotRestartShutdown()`. The heartbeat was itself in that set and waited for restart completion, so shutdown never wrote the adoption snapshot. - **Expected behavior:** An eligible hot restart captures its snapshot before waiting for scheduler work, preserves live child processes, and exits after one SIGTERM. - **Steps to reproduce:** 1. Start a heartbeat that remains active while requesting a hot restart. 2. Send SIGTERM to the server process. 3. Observe shutdown waiting on the active scheduler task and startup finding an intent without a shutdown snapshot. - **Paperclip commit:** `992389480a243b97bda214227e0767eb8c3672af` - **Deployment/install:** Self-hosted server built from source. - **Adapter:** Not adapter-specific; reproduced with a Codex heartbeat. - **Database/access:** Embedded PGlite; agent bearer context. - **Environment:** Node `v22.22.2` on `Linux 6.17.0-1015-aws aarch64 GNU/Linux`. - **Privacy:** No secrets, private logs, user paths, or internal issue references are included. ## What Changed - Add a focused shutdown coordinator that prepares hot-restart state before waiting for heartbeat scheduler idleness. - Skip scheduler-idle and graceful-drain waits only when the hot-restart service returns `skipDrain: true`. - Preserve normal graceful shutdown behavior when no eligible intent exists or preparation fails. - Add regression coverage for pending scheduler work, normal shutdown, and preparation failure. ## Verification - `pnpm exec vitest run server/src/shutdown.test.ts` — 3 passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t 'hot-restart'` — 3 passed, 88 skipped. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low-to-moderate risk: shutdown ordering changes, but only the explicitly eligible hot-restart path bypasses scheduler-idle and run-drain waits. - Normal shutdown and hot-restart preparation failures retain the existing graceful behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This is a focused bug fix and does not duplicate planned roadmap work. ## Model Used - OpenAI Codex coding agent; exact runtime model ID and context-window size are not exposed to the agent. Tool use, shell execution, repository editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change required for this internal shutdown-order fix) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.9 |
||
|
|
992389480a |
fix(server): restore hot-restart run adoption (#9647)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local heartbeat/runtime subsystem starts long-running local agent processes and records their run state. > - Operators sometimes need to rebuild and restart the Paperclip server while local agent processes are still alive. > - A normal restart should remain conservative, but a guarded production hot restart needs an explicit marker, startup reconciliation, and an inspectable report. > - The broader hot-restart PR is currently merge-conflicted, so this pull request lands the minimal server-side recovery path on current `master`. > - The benefit is that deploy operators can restart from a current branch without reverting production changes and without marking adopted live runs as `process_lost`. ## Linked Issues or Issue Description No public GitHub issue exists for this deploy-safety fix. Bug fix: - What happened: the current deployable `master` branch did not include the hot-restart marker CLI, startup adoption report path, or health version proof needed by guarded service restarts. - Expected behavior: a deploy operator can write a one-shot marker before restarting, the old server snapshots eligible running child processes, the new server reports adopted/finalized/lost runs, and adopted live runs are not reaped as `process_lost`. - Steps to reproduce: restart a server with running local child-process heartbeat runs without the marker/adoption path; startup orphan reaping has no adoption metadata and treats live detached children as lost. - Paperclip version/commit: fixed on top of `master` at `b606869a6`. - Deployment mode: production/local-service style deployments that rebuild and restart the primary `paperclip.service`. - Related PR: Refs #9628. This PR intentionally lands a smaller deploy-safe subset because #9628 is currently merge-conflicted. - Duplicate search: searched public PRs/issues for `hot restart` and `process_lost adoption`; #9628 is the directly related prior implementation. ## What Changed - Added `scripts/request-hot-restart.ts` to write a one-shot hot-restart intent marker under `PAPERCLIP_HOME`. - Added `server/src/services/hot-restart.ts` for intent/report path resolution, parsing, atomic writes, shutdown snapshots, and marker cleanup. - Wired server shutdown/startup so explicit hot restarts snapshot active runs, skip the normal heartbeat drain, reconcile live child processes on boot, and write `hot-restart-report.json`. - Preserved adopted run metadata so normal orphan reaping does not regress adopted live runs to `process_lost`. - Added `serverVersion` health proof alongside existing `version`, plus docs and regression coverage. ## Verification - `pnpm vitest run server/src/__tests__/health.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 2 files passed, 100 tests passed. - `pnpm --filter @paperclipai/server typecheck` - `env PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/hot-restart-cli-smoke" pnpm --filter @paperclipai/server exec tsx ../scripts/request-hot-restart.ts --server-pid 12345` - Branch ancestry checked after `git fetch origin master`: `origin/master` was `b606869a6`, and `HEAD..origin/master` was empty. ## Risks - Medium risk: process adoption depends on PID/PGID metadata and the service manager leaving child processes alive for the guarded restart. - Normal restarts remain conservative, but an incorrect marker PID intentionally falls back to graceful drain instead of adoption. - The PR is server-only and does not include the broader UI/experimental-setting work from #9628. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 via Codex coding agent in a Paperclip execution workspace; tool use and shell/code execution enabled; context window not surfaced by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.8 |
||
|
|
cf7711ecc5 |
feat(ui): explain when a message won't reopen a blocked issue (Rule C, PAP-13554) (#9417)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - When an agent's task is `blocked`, humans often comment on the issue
thread expecting it to reopen and resume
> - If the issue stays `blocked` because an unresolved (not-done)
blocker still gates the reopen path, the UI said nothing — the human
"sent a message and nothing happened" (a real silent-failure report)
> - That silence makes the product feel broken even though the gate is
working as designed
> - This pull request adds Rule C copy to the existing
`IssueBlockedNotice` surface: when a comment won't reopen a blocked
issue, it explains why and names the deepest unresolved blocker leaf
with its status
> - The benefit is that the human immediately understands the task is
*paused, not stuck*, and knows exactly which task to act on to unblock
it
## Linked Issues or Issue Description
<!-- No public GitHub issue — describing the underlying problem inline
(path B), following the feature issue template. -->
**Problem or motivation**
A human comments on a `blocked` issue expecting it to move back to
`todo` and resume. When the reopen gate keeps it blocked because an
unresolved (not-`done`) blocker remains, nothing in the UI communicates
this, so the user perceives a dropped message ("I sent a message and
nothing happened").
**Proposed solution**
Reuse the existing amber `IssueBlockedNotice` surface (the designated
blocked/recovery surface — no new component). When a message won't
reopen a blocked issue, the notice states that it won't reopen yet,
names the unresolved blocker leaf with its status (e.g. "Still blocked
by PAP-XXXXX (in progress)"), and reassures that it reopens
automatically once the blocker is done. UI-only; no server change.
**Alternatives considered**
Adding a server signal (a reopen-suppressed reason on the comment/notice
payload) was considered but rejected as unnecessary — the
unresolved-blocker set is already available client-side, so the copy is
derived in the component. Done-but-pending-finalize blockers are `done`,
so they fall out of the unresolved set into the standard reopen (Rule B)
path and are correctly not shown as reopen-suppressed.
## What Changed
- `ui/src/components/IssueBlockedNotice.tsx`: added reopen-suppressed
messaging for `blocked` issues that still have unresolved blockers — a
lead sentence ("a message won't reopen it yet, then it reopens
automatically"), the named unresolved blocker leaf with its status, and
an "and N other task(s)" summarization when multiple blockers remain.
- `ui/src/components/IssueBlockedNotice.test.tsx`: added/updated tests
covering the single, nested-chain (deepest-leaf), multiple-blocker, and
empty-blocker states, plus the not-a-reopen-case (`in_progress`) path.
- `ui/storybook/stories/issue-blocked-notice.stories.tsx`: new Storybook
stories rendering each notice state for copy review.
## Verification
- `cd ui && tsc -b` — typecheck clean (previously failed TS2322 on the
story meta; fixed by a default `args`).
- Vitest: `IssueBlockedNotice.test.tsx` green (single / nested /
multiple / empty / in-progress states).
- Storybook stories rendered at 1440×900 for all four states;
screenshots posted on the tracking issue.
- UXDesigner reviewed the notice copy and signed off (no copy edits
required).
## Risks
Low risk. UI-only, additive copy on an already-shipped amber notice
surface; no server or schema changes. The reopen behavior itself is
unchanged — this only explains the existing gate. Worst case is copy
wording, which had a design review.
## Model Used
Claude — claude-opus-4-8 (Opus 4.8), extended thinking, with tool use /
code execution in the Paperclip harness.
---
- [x] I searched the GitHub PR list for similar/duplicate PRs before
opening this one.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4f9894df44 |
fix(server): bound accepted-interaction continuation recovery (#9656)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Heartbeat recovery keeps assigned issues moving when a run or continuation path disappears > - Accepted issue-thread interactions can create a continuation wake after an agent previously parked for review > - The recovery sweep could requeue that accepted-interaction wake while the queued-run gate cancelled it using the older pre-acceptance park summary > - That cancellation path had no bound, so recovery could repeat the same wake and cancellation indefinitely > - This pull request makes accepted-interaction evidence supersede the older park and caps repeated recovery cancellations at three attempts > - The benefit is that accepted work resumes normally, while genuine repeated failures become a visible dependency wait or escalation instead of a cancel loop ## Linked Issues or Issue Description Refs #9331 The accepted-interaction continuation recovery added by #9331 can encounter a stale continuation summary written before approval. The sweep requeues a continuation carrying the accepted interaction timestamp, but queued-run invalidation cancels it because the older summary says to wait for review. Recovery then sees the accepted interaction without a successful run and requeues again. This PR prevents that stale-summary cancellation and adds a bounded fallback if three equivalent cancellations have already occurred. ## What Changed - Let queued continuation wakes with a parseable `interactionResolvedAt` bypass a pre-acceptance waiting-for-review park summary. - Count consecutive unsuccessful continuation runs for the same issue and agent since interaction acceptance; after three review-park cancellations, convert a real dependency wait or use the existing visible escalation path. - Add focused regression coverage for the park bypass, unchanged non-interaction park behavior, below-cap requeue, cap escalation, and successful-run skip. - Document the accepted-interaction precedence and bounded requeue contract in execution semantics §9.2. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "accepted interaction continuation recovery|accepted interaction recovery after its continuation succeeds|requeues accepted interaction continuations stranded"` - `pnpm vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts -t "pre-acceptance review park|continuation summary parks executor work"` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Low risk: the park bypass only applies when the queued context contains a parseable interaction resolution timestamp. - The retry bound is scoped to unsuccessful `issue_continuation_needed` runs for the same company, issue, agent, error code, and post-acceptance time window. - No schema, migration, API, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.5`, high-reasoning coding mode with repository tool use and command execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3124dd0f1e |
feat(server): recovery observability report and rate alert (#9644)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - When a run is stranded (process lost, adapter failure, a finished
run with no disposition, an over-eager inactivity kill), the harness
opens a *recovery action* and wakes an owner to recover it
> - Recovery volume regressed sharply in one week — 3.26% of all runs vs
a ~1.2% monthly norm, 5–8x the prior volume — and nobody noticed until
it was ~194 actions deep, because there was no way to *see* the recovery
rate
> - We also could not see which causes drive recovery, nor how often a
manager ends up doing the deliverable work themselves instead of handing
it back to the original owner (the product goal is that managers doing
the work stays rare)
> - This pull request adds a recovery-observability report + API
endpoint: weekly rate normalized per run, a threshold alert, the cause
taxonomy live from the ledger, and the handed-back vs owner-completed
ratio and per-cause routing outcomes
> - The benefit is that a recovery regression like that week is caught
by a threshold instead of by a human noticing it by feel, and each
recovery playbook row can be verified in production
## Linked Issues or Issue Description
**Feature.**
**Problem or motivation**
Recovery takeovers are a first-class exception path
(`issue_recovery_actions`), but there is no aggregate view of them. A
week where the recovery rate tripled went unnoticed until it was deep.
There is no signal for (a) the per-run recovery rate over time, (b)
which cause + run error code drives it, or (c) whether the recovery
owner hands the task back to the original assignee or ends up doing the
deliverable work themselves.
**Proposed solution**
A read-only report service and `GET
/companies/:companyId/recovery-observability` endpoint that surfaces the
weekly rate, a threshold alert, the cause taxonomy, the hand-back ratio,
and per-cause routing outcomes.
**Alternatives considered**
Adding `handed_back` / `owner_completed` to the recovery-action outcome
vocabulary and writing them at resolution time. Rejected for this
change: the distinction is derivable from the recovery owner, the
recorded return owner, and where the source issue actually landed, so
the report works against all historical data without a backfill.
**Roadmap alignment**
Implements the recovery-observability line of the approved
recovery-takeover plan (make regressions visible via a threshold rather
than by human feel); no schema or write-path change.
## What Changed
- Add `server/src/services/recovery-observability.ts`:
- `recoveryObservabilityService(db).report(companyId, { weeks,
thresholdPercent, now })` returns weekly rates (recovery actions / runs,
Monday-anchored to match the retrospective), a `cause` +
`latestRunErrorCode` breakdown, a handed-back vs owner-completed
summary, and per-cause routing outcomes.
- `evaluateRecoveryRateAlert(weekly, thresholdPercent)` — a pure
function (default threshold 2% of runs) returning the breached weeks and
whether the latest week regressed.
- `classifyRecoveryHandoff(...)` — a pure classifier deriving
`self_recovery` / `handed_back` / `owner_completed` from the recovery
owner, return owner, and final issue landing.
- Add `GET /companies/:companyId/recovery-observability` (optional
`weeks` and `threshold` query params) to the existing dashboard router.
- The `weeks` window is bounded (`MAX_WINDOW_WEEKS = 104`,
service-authoritative and re-clamped at the route) so a large query
value can't over-allocate the per-week array.
- Add tests: unit coverage for the alert and the classifier, plus an
embedded-Postgres integration test that seeds synthetic runs and
recovery actions crossing 2% and asserts the alert fires and the
hand-back ratio is computed.
## Verification
- `CI=1 NODE_ENV=development npx vitest run
server/src/__tests__/recovery-observability.test.ts` — 10/10 pass
(includes the synthetic 2%-crossing alert case and the hand-back ratio
case).
- Rendered against a live database of 300+ recovery actions: the weekly
rates reproduce the retrospective (e.g. 1.37% / 1.53% / 0.65% / 0.86% /
1.37% for early-June weeks), the alert fires on the two most recent
weeks (3.15% and 3.05%, both over 2%), and the hand-back summary shows
owner-completed ≈ 73% vs handed-back ≈ 27% — matching the observed
"managers keep ~80% of takeovers".
## Risks
- Low risk. Read-only: adds one GET endpoint and a service; no schema,
migration, or write-path changes. The hand-back classification reads the
source issue's current assignee/status, so a much-later reassignment
could reclassify a historical action — acceptable for an aggregate trend
view.
## Model Used
- Claude, `claude-opus-4-8` (Opus 4.8), extended thinking, tool use /
code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
canary/v2026.716.0-canary.7
|
||
|
|
f44a002b8d |
fix(ui): restore prefix-aware company export/import links (#6648)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Each company workspace in the web UI is mounted under a URL prefix (e.g. `/NEU/company/...`), and `Link` from `@/lib/router` applies that prefix automatically via `applyCompanyPrefix` > - Company Settings rendered its Org Chart, Export, Import, and Cloud Upstream buttons as raw `<a href>` anchors, which drop the prefix, so those pages 404 on prefixed instances (#2910); `CompanyExport.filePathFromLocation` also failed to locate `/company/export/files/` inside prefixed URLs, breaking export file previews > - #2951 fixed the Settings links with `<Link to>` plus tests, but the sandbox-settings work in #4415 reverted the links back to `<a href>`, silently reintroducing the bug (#6647) > - This pull request restores the prefix-aware `<Link>` for all four Settings links, normalizes the pathname with `toCompanyRelativePath()` before matching the export-files marker, and adds regression tests covering every route so the fix cannot be lost again > - The benefit is that export/import/org-chart/cloud-upstream navigation and export file previews work on every prefixed deployment ## Linked Issues or Issue Description Fixes: #6647 Refs #2910 (original report: `/company/export` → prefix `COMPANY` → not found) Refs #2951 (original fix with `<Link to>` + tests — merged, then lost) Refs #4415 (sandbox settings PR that reverted Settings back to `<a href>`) ## What Changed - **`ui/src/pages/CompanySettings.tsx`**: use `Link` from `@/lib/router` for the Org Chart, Export, Import, and Cloud Upstream buttons (replacing raw `<a href>`) - **`ui/src/pages/CompanyExport.tsx`**: resolve file paths from prefixed URLs by normalizing with `toCompanyRelativePath()` before matching `/company/export/files/` - **`ui/src/lib/company-routes.test.ts`**: regression tests for export/import/cloud-upstream/org prefix rewriting, double-prefix prevention, and export file URL normalization ## Verification ```bash pnpm vitest run ui/src/lib/company-routes.test.ts ``` Manual: 1. Open `http://localhost:3100/NEU/company/settings` (or your company prefix). 2. Click **Export** / **Import** — URL should stay under `/:prefix/company/...`. 3. On export, select a file — URL should be `/:prefix/company/export/files/...` and preview should load. The change is navigation-target-only (no visual/layout changes), so before/after is shown as the resolved URLs: | Link | Before (prefix dropped → not found) | After | |------|-------------------------------------|-------| | Export | `/company/export` | `/NEU/company/export` | | Import | `/company/import` | `/NEU/company/import` | | Org Chart | `/org` | `/NEU/org` | | Cloud Upstream | `/company/settings/cloud-upstream` | `/NEU/company/settings/cloud-upstream` | ## Risks Low — same approach as #2951; only navigation/parsing, no API changes. ## Model Used - Original implementation: authored by @qbamca in Cursor (agentic editor; the session's exact model ID was not recorded) - Follow-up commit (merge-conflict resolution) and this description update: Claude Fable 5 (Anthropic, `claude-fable-5`, extended thinking, agentic tool use), operated by the Commit Capital triage team ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc changes required — behavior matches documented routing) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending re-review of the conflict-resolution commit) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Jakub Mikiciuk <jmikiciuk@igus.net> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.716.0-canary.6 |
||
|
|
df2404d443 |
fix(codex): count raw child output as activity (#9632)
## Thinking Path > - Paperclip is the open source control plane teams use to manage AI agents and their work > - The Codex local adapter supervises CLI child processes and terminates genuinely silent runs > - The existing inactivity timer only recognized parsed JSONL stdout events as activity > - Long verification commands can emit ordinary stdout or stderr while producing no JSONL, so healthy children could be killed > - This pull request makes the watchdog observe raw child-process output before filtering or parsing > - The benefit is that long typecheck, build, and test phases survive while truly silent children remain bounded ## Linked Issues or Issue Description - **Bug:** The Codex output-inactivity watchdog could terminate healthy runs during long verification phases because non-JSON stdout and stderr did not reset its timer. - **Expected:** Any bytes emitted by the child process count as output activity; only a child with no stdout or stderr for the configured interval is terminated. - **Reproduction:** Configure a short `outputInactivityTimeoutMs`, run a Codex child that periodically emits plain-text verification progress without JSONL events, and observe the old monitor firing despite continued output. - **Deployment mode:** Local `codex_local` adapter execution. - Related implementation: #5017 - Related recovery behavior: #8680 ## What Changed - Reset the Codex inactivity monitor on every non-empty stdout or stderr chunk before stderr noise filtering. - Track raw output chunk and byte counts in monitor diagnostics while retaining parsed JSONL event counts. - Add regression coverage for more than 21 simulated minutes of non-JSON verification output and retain silent-child termination coverage. - Document that `outputInactivityTimeoutMs` observes raw child output and that `null` still disables the monitor. ## Verification - `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/vitest run packages/adapters/codex-local/src/server/output-inactivity-monitor.test.ts packages/adapters/codex-local/src/server/output-inactivity-monitor.integration.test.ts` - `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/tsc -p packages/adapters/codex-local/tsconfig.json --noEmit` ## Risks - Low risk: the monitor becomes more conservative and may allow a noisy-but-stuck child to run longer, but the configured hard timeout and platform silent-run safety net remain unchanged. - No schema, API, migration, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; runtime model ID and context-window size were not exposed to the agent. Used repository/tool access, code execution, and focused test verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.5 |
||
|
|
8eff54bc47 |
[codex] Explain AWS secret creation failures in the UI (#9645)
## Thinking Path > - Paperclip is the open source control plane operators use to manage AI-agent companies. > - Operators can store runtime secrets in provider vaults such as AWS Secrets Manager. > - The server already preserves sanitized AWS failure details, including the failed operation, required IAM capability, region, and safe recovery options. > - The create-secret dialog reduced that structured response to a generic message, leaving operators unable to understand or fix failed AWS writes. > - This pull request keeps the safe structured error through the UI and presents concise, actionable diagnostics without exposing raw AWS principals or account details. > - The benefit is that operators can correct IAM access or link an existing AWS secret immediately instead of debugging an opaque failure. ## Linked Issues or Issue Description No public issue was found for this exact UI gap. Bug report: - What happened: creating a Paperclip-managed value in AWS Secrets Manager could fail with a generic dialog error even though the API returned safe, actionable provider details. - Expected behavior: the dialog should identify the AWS operation, required IAM capability, region, provider vault, and safe alternative while keeping raw cloud-provider details redacted. - Steps to reproduce: configure an AWS Secrets Manager provider vault without `secretsmanager:CreateSecret`, then create a Paperclip-managed secret using that vault. - Paperclip version/commit: current `master` before this PR. - Deployment mode: Paperclip server with an AWS Secrets Manager provider vault. Related prior server-side propagation work: - Refs #9161 ## What Changed - Preserve the structured `ApiError` returned by failed create-secret mutations instead of reducing it to a string. - Render AWS-specific, sanitized diagnostics with the required IAM capability, region, provider vault, operation, and external-reference recovery option. - Add a full dialog render regression test that verifies actionable details appear and raw AWS ARN/account information does not. ## Verification - `pnpm exec vitest run ui/src/pages/Secrets.render.test.tsx` — 1 file passed, 15 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check public-gh/master...HEAD` — passed. ### Visual Verification QA verified both states in Chromium at head `357cd2271` using the real `Secrets` component and confirmed that raw AWS account, ARN, assumed-role, and provider exception details are absent from the rendered DOM. **AWS access-denied diagnostics**  **Generic non-AWS fallback**  ## Risks Low risk. The change only affects failed create-secret presentation in the UI; successful secret creation, API contracts, schema, and migrations are unchanged. The structured details are server-sanitized, and the regression test confirms raw AWS principal/account data is not rendered. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI `gpt-5.4` via Codex CLI, with tool-enabled repository editing, shell execution, Git, GitHub, and Paperclip API access. Reasoning mode and exact context-window size were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8368fb30b0 |
fix(routines): coalesce sub-hourly catch-up runs (#9649)
## Thinking Path > - Paperclip is the open source control plane people use to run AI-agent companies. > - Scheduled routines support catch-up policies when the server resumes after missed cron ticks. > - The existing capped replay policy dispatched once per missed tick, which can flood the board after downtime for frequent schedules. > - Sub-hourly routines usually need one prompt catch-up execution rather than historical per-tick replay, while hourly-or-slower schedules may rely on the existing behavior. > - This pull request coalesces missed sub-hourly ticks into one execution and keeps the slower-schedule behavior unchanged. > - The benefit is bounded recovery work without changing the semantics of lower-frequency scheduled routines. ## Linked Issues or Issue Description ### What happened? When a scheduled routine using `enqueue_missed_with_cap` resumes after several missed sub-hourly cron ticks, Paperclip dispatches one catch-up execution for every missed tick. Those executions arrive in a same-second burst and can flood the board with duplicate-looking work. ### Expected behavior Sub-hourly schedules should advance past all missed ticks but dispatch exactly one catch-up execution. Hourly-or-slower schedules should retain capped per-tick replay. ### Steps to reproduce 1. Build Paperclip from `master` and create a routine with a sub-hourly cron schedule and `catchUpPolicy: enqueue_missed_with_cap`. 2. Set its persisted `nextRunAt` far enough in the past to cover several scheduled occurrences. 3. Run routine catch-up processing. 4. Observe multiple catch-up dispatches instead of one coalesced execution. ### Paperclip version or commit Reproduced on `master` before this PR. ### Deployment mode Built from source in local development with embedded PGlite. ## What Changed - Classify sub-hourly cadence from timezone-aware scheduled occurrences, avoiding daily multi-minute false positives while supporting schedules restricted to active days. - Coalesce all missed sub-hourly ticks into one catch-up dispatch while advancing `nextRunAt` to the next future occurrence. - Preserve capped per-tick replay for hourly-or-slower schedules. - Clarify the catch-up policy labels in both routine editing surfaces. - Add regression coverage for both the coalesced and preserved behaviors. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts --testNamePattern='coalesces multiple missed sub-hourly ticks|continues replaying each missed hourly tick|continues replaying missed ticks for daily schedules with multiple minute values|coalesces sub-hourly schedules restricted to weekdays'` — 4 passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check origin/master...HEAD` — clean. ## Risks - Low-to-moderate behavioral risk: sub-hourly routines using `enqueue_missed_with_cap` now intentionally receive one recovery execution instead of one per missed tick. - Hourly-or-slower schedules retain their previous capped replay behavior, limiting the compatibility surface. - No schema, migration, workflow, or lockfile changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI with GPT-5.5, medium reasoning, code execution and repository tool use; the runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd7c0d5f83 |
fix(issues): deduplicate repeated creates (#9650)
## Thinking Path > Paperclip already treats issue creation as a company-scoped mutation, but retries and parallel agent heartbeats can submit the same create more than once. Client instructions cannot provide at-most-once behavior under concurrency, so the guard belongs in the server transaction. This change adds an explicit company-scoped idempotency contract, a conservative fallback for recent open same-parent titles, and run attribution for auditability. Advisory transaction locks serialize competing requests before lookup/insert, avoiding the race that affected the prior attempt. ## Linked Issues or Issue Description Fixes #6529. This is a clean replacement for #6936, which was closed because it mixed unrelated changes and its check-then-insert implementation was not concurrency-safe. Unlike that attempt, this PR is scoped to eight files, uses a dedicated idempotency-key table, and serializes duplicate candidates inside the create transaction. ## What Changed - Accept optional `idempotencyKey` and `allowDuplicate` fields on issue creation. - Replay the existing issue with HTTP 200 and deduplication metadata for a repeated company/key pair. - Deduplicate recent open issues with the same company, parent, and normalized title for 48 hours unless `allowDuplicate: true` is supplied. - Persist idempotency mappings in a company-scoped table and serialize competing creates with transaction advisory locks. - Populate `originRunId` from `X-Paperclip-Run-Id` for agent/manual creates when the body does not provide an origin run. - Add route integration coverage for key replay, title fallback, bypass, closed/old recreation, company scoping, and run attribution. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-create-deduplication-routes.test.ts` — 7 tests passed. - `pnpm --filter @paperclipai/db typecheck` — passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. - `pnpm exec vitest run server/src/__tests__/issue-assigned-backlog-contract-routes.test.ts server/src/__tests__/issue-create-deduplication-routes.test.ts` — 10 tests passed after the service-contract compatibility fix. ## Risks - The title fallback intentionally treats normalized same-parent titles as duplicates for 48 hours; callers creating intentionally repeated titles must send `allowDuplicate: true`. - Advisory locks use hashed duplicate keys, so an extremely unlikely hash collision can serialize unrelated creates but cannot merge their lookup results. - Deleting an issue cascades its idempotency mapping, allowing the same key to create a replacement later. ## Model Used - OpenAI `gpt-5.6-sol`, high reasoning effort, Codex CLI with repository, shell, GitHub CLI, and Paperclip API tool access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea0e899905 |
fix(search): honor extract match limits + harden pr-gardening candidate discovery (#9652)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `/pr-gardening` skill drives a bundled agent that scans a
company's issues for those linked to open GitHub PRs, then reports on
their state; it relies on the server's company-search **extract**
endpoint to pull PR references out of issue bodies
> - Two gaps surfaced during end-to-end QA of the gardening workflow:
the extract service silently ignored a per-issue match cap, so callers
could not bound how many matches came back per issue, and the skill's
candidate-discovery scripts fell over on large repos and on issues that
referenced deleted PRs
> - Left unaddressed, the gardener either truncated its scan
unpredictably or aborted outright, so it could not reliably enumerate PR
candidates
> - This pull request honors an explicit `matchesPerIssue` limit in the
extract search API and hardens the skill's candidate discovery against
missing/unavailable PRs and oversized `gh` output
> - The benefit is a PR-gardening workflow that scans deterministically
and finishes cleanly on real-world companies
## Linked Issues or Issue Description
No pre-existing public GitHub issue — describing the bug in-PR following
the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`).
### What happened?
The company-search extract endpoint accepted a per-issue match limit but
did not apply it, returning matches capped only by the old hardcoded
constant regardless of the caller's request. Separately, the
`/pr-gardening` skill's candidate-discovery scripts crashed when a
scanned issue referenced a deleted PR (GitHub `Not Found (HTTP 404)` /
GraphQL `Could not resolve to a PullRequest`) and could exceed the
default `gh` output buffer on large result sets, aborting the whole
scan.
### Expected behavior
The extract API bounds matches per issue when a caller passes
`matchesPerIssue` (default 20, max 200), and omitting it preserves the
previous default. The gardening scripts skip PRs that are
deleted/unavailable and tolerate large `gh` responses without aborting
the scan.
### Steps to reproduce
1. Call the company-search extract endpoint with a `matchesPerIssue`
value against an issue containing many PR references — previously the
value was ignored.
2. Run the pr-gardening candidate scan against a company whose issues
reference a since-deleted PR — previously the scan threw instead of
skipping that PR.
### Paperclip version or commit
`master` at the base of this PR (branch cut from current
`origin/master`).
### Deployment mode
Local Paperclip instance / self-hosted.
## What Changed
- **Extract search honors `matchesPerIssue`**: added the
`matchesPerIssue` field to the shared search validator/types and applied
the cap in `company-search-extract` so results are bounded per issue
(`packages/shared`, `server/src/services/company-search-extract.ts`,
`doc/SPEC-implementation.md`).
- **Hardened pr-gardening candidate discovery**: `find-candidates.mjs` /
`lib.mjs` now request `matchesPerIssue=200`, treat missing/unavailable
PRs (deleted PR → `isMissingPullRequestError` / `unavailable`) as skips
instead of fatal errors, and raise the `gh` `maxBuffer` to 50 MB for
large repos.
- **Tests**: expanded `company-search-extract-{routes,service}.test.ts`
for the new limit and added coverage in `pr-gardening.test.mjs`.
## Verification
Re-run on a fresh worktree cherry-picked onto current `master`:
- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` → 8/8 pass
- `pnpm vitest run
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-extract-service.test.ts` → 10/10
pass
## Risks
Low risk. `matchesPerIssue` is optional and backward-compatible
(omitting it preserves prior behavior). The skill changes only add
skip/tolerance paths and a larger buffer; no schema or migration
changes.
## Model Used
Claude — Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use /
code execution via the Claude Agent SDK.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.716.0-canary.4
|
||
|
|
5588ddf681 |
fix(server): prevent recurring worktree port conflicts (#9642)
## Thinking Path > - Paperclip is the control plane operators use to run AI-agent companies and their isolated development workspaces. > - Worktree startup assigns each workspace a server port and an embedded PostgreSQL port. > - Existing collision detection depended on discovering sibling configs from the current repository layout, so worktrees in different repository roots could select the same ports. > - Concurrent startup also had no shared critical section, allowing two worktrees to observe the same available ports before either persisted its selection. > - Repeated collisions prevented otherwise isolated workspaces from starting reliably and could recur after a port was repaired once. > - This pull request adds a shared, locked registry of active worktree config paths and uses it during port selection and repair. > - The benefit is stable, persisted, cross-repository port isolation for both the Paperclip server and embedded PostgreSQL. ## Linked Issues or Issue Description ### What happened? When multiple Paperclip worktrees shared the same worktree home but lived under different repository roots, startup could assign duplicate server and embedded PostgreSQL ports. The prior sibling scan did not reliably discover configs outside the current repository, and simultaneous repairs were not serialized. ### Expected behavior Each active worktree should reserve unique server and database ports across repository roots, persist any repaired selection, and reuse the persisted ports on subsequent starts. ### Steps to reproduce 1. Create two Paperclip worktrees in different repository roots that share `PAPERCLIP_WORKTREES_DIR`. 2. Give both worktree configs the same server and embedded PostgreSQL ports. 3. Start or repair both worktrees. 4. Observe that both can retain the same ports because neither reliably discovers the other configuration. ### Environment - Version: reproducible on `master` before this change - Deployment: local development worktrees built from source - Adapter: not adapter-specific - Database: embedded PostgreSQL Related prior reliability work: #1829. Related documentation for recovering port conflicts: #9407. ## What Changed - Add a shared `worktree-port-reservations.json` registry under the worktree home, containing live worktree config paths. - Serialize registry reads, collision detection, config repair, and registry updates with a stale-safe filesystem lock. - Include registered configs and isolated instance configs when collecting reserved server and embedded PostgreSQL ports. - Atomically prune stale registry entries and persist repaired ports plus the matching public base URL. - Add regression coverage for cross-repository collisions, persisted repairs, and repeat startup behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/worktree-config.test.ts` — 14 tests passed, including stale-lock recovery. - `pnpm --filter @paperclipai/server typecheck` — passed. - Rebased onto current `public-gh/master` before verification. ## Risks - Low-to-moderate risk: worktree startup now briefly acquires a filesystem lock in the shared worktree home. - The lock has a 10-second acquisition timeout and removes lock directories older than 5 seconds so interrupted owners are recoverable within the wait window. - Registry writes are atomic and stale config paths are pruned, limiting persistent state to existing worktree configs. - The change is scoped to worktree runtime configuration and does not affect normal main-instance configuration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.3 Codex and GPT-5.4 with repository access, terminal execution, and code-review tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.3 |
||
|
|
f1508a7929 |
feat(decisions): expand image rows into gallery (#9532)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and review work that needs human attention. > - The Decisions page condenses approvals, failed runs, reviews, and issue interactions into a scannable attention queue. > - Some decision rows already include screenshot evidence, but the collapsed thumbnail stack is too small for meaningful inspection. > - Image-only rows were not expandable because expansion was previously reserved for inline decision resolvers. > - Reviewers therefore had to leave the Decisions page before they could understand the visual evidence attached to a row. > - This pull request makes rows with images expandable and presents a readable, linked gallery while preserving the compact collapsed view. > - The benefit is faster evidence review without sacrificing queue density or existing inline-resolution behavior. ## Linked Issues or Issue Description No public GitHub issue currently tracks this feature. ### Subsystem affected `ui/` — React + Vite board UI ### Problem or motivation Screenshot evidence in Decisions rows is only visible as small overlapping thumbnails, and non-inline rows cannot expand to show it. This makes visual review unnecessarily slow and forces reviewers to navigate away from the queue. ### Proposed solution Treat active rows with images as expandable, render the first three images at a readable size in an expanded gallery, and link images plus any remaining-image affordance to the related issue. ### Alternatives considered Always rendering large images would make the queue difficult to scan; opening the issue immediately preserves density but prevents in-context review. An explicit expandable gallery keeps both behaviors available. ### Roadmap alignment `ROADMAP.md` does not list overlapping Decisions image-gallery work. This is a focused improvement to the existing review surface rather than a new product area. ### Additional context The Storybook variants document both the collapsed thumbnail treatment and deterministic expanded gallery state for reviewer inspection. ## What Changed - Allow active Decisions rows with screenshot evidence to expand even when they have no inline resolver. - Keep compact thumbnails in collapsed rows and render up to three larger, linked images when expanded. - Add an accessible remaining-image tile that links to the related issue when more screenshots exist. - Add component coverage for image-only expansion and the remaining-image issue link. - Add collapsed and expanded image-gallery Storybook variants, including deterministic initial expansion. ## Verification - `pnpm exec vitest run ui/src/components/AttentionQueueRow.test.tsx` - `pnpm check:token-gates` - `pnpm --dir ui typecheck` - `pnpm --dir ui build-storybook` - QA visual verification (light + dark, no defects): https://github.com/paperclipai/paperclip/pull/9532#issuecomment-4964769485 ## Risks - Low risk: the change is isolated to Decisions row rendering and Storybook fixtures. - Rows with images gain a new expansion interaction, but existing inline resolver behavior and deep links remain intact. - The gallery intentionally limits the in-row preview to three images to avoid unbounded row height. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 assisted with the original implementation and tests. - OpenAI `gpt-5.4` via Codex CLI assisted with current-master rebase integration, verification, and PR preparation. The runtime context-window size is not exposed; capabilities used include reasoning, repository tool use, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d32ed88443 |
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies and their work. > - Its recovery subsystem detects stranded issue execution and decides whether to retry, escalate, or request operator intervention. > - The existing recovery path used a mostly generic owner ladder and generic execution contract, so transient failures could wake a manager who then performed the deliverable instead of repairing and returning the task. > - Provider quota failures also entered the same takeover path even when the correct action was to wait for capacity and retry the original assignee. > - Recovery actions already retain the source owner and evidence needed to choose a cause-specific route, render a scoped contract, and measure whether work was handed back. > - This pull request adds a cause-keyed recovery playbook, propagates its contract through every built-in adapter, and makes resolved recovery actions return work to the original owner by default. > - The benefit is bounded self-recovery that preserves task ownership, avoids needless management takeover, and makes recovery outcomes observable. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related recovery work was reviewed but is not duplicated here: #9630 restores bounded recovery continuations, #8807 changes one assignee-ranking case, and #9404 records runtime-failure transition evidence. This change instead introduces cause-specific routing and recovery contracts across the recovery lifecycle. ### What happened? When an issue became stranded, recovery generally selected an owner through the same fallback ladder and rendered the normal execution contract. That made the recovery wake look like ordinary deliverable work, even when the correct action was to retry the original agent, repair its runtime, or wait for a provider quota reset. ### Expected behavior Recovery should select a response by failure cause, tell the recipient to recover rather than complete the deliverable, suppress takeover wakes for provider quota waits, and return repaired work to its original assignee unless the recovery owner explicitly completes it. ### Actual behavior Recovery could escalate transient failures to management, omit the cause-specific next action from the wake, and leave the recovery owner assigned after the runtime problem was resolved. ### Impact The generic path creates avoidable management work, ownership churn, and budget consumption while obscuring whether recovery successfully returned work to the responsible agent. ## What Changed - Added cause-keyed routing for process loss, missing disposition, provider quota limits, Codex output inactivity, workspace validation failures, and fallback recovery causes. - Added recovery-scoped wake rendering that replaces the generic execution contract with the failure summary, original assignee, attempt count, next action, and cause-specific playbook instruction. - Propagated the structured recovery contract through all built-in adapter execution paths, including Hermes local and gateway adapters. - Added provider-quota wait monitoring so capacity failures schedule the original assignee instead of enqueueing a takeover wake. - Added hand-back behavior and `handed_back` / `owner_completed` outcome accounting when recovery actions are resolved. - Added focused routing, renderer, quota-monitor, and hand-back regression coverage plus implementation-spec documentation. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/issue-recovery-actions.test.ts` - 4 test files passed; 194 tests passed. - Targeted `pnpm --filter ... typecheck` across `@paperclipai/adapter-utils`, `@paperclipai/shared`, `@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter packages. - 13 affected workspace packages passed typecheck. - `pnpm check:token-gates` - All UI token gates passed. ## Risks - Recovery routing behavior changes for stranded work, so an incorrectly classified cause could select a different recipient than before; fallback causes retain the existing management ladder. - Provider quota detection depends on structured failure evidence and conservative text matching; unmatched failures continue through fallback recovery. - Adapter prompt plumbing changes across built-ins, covered by shared renderer tests and compile-time call signatures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b95eece5 |
fix(server): preserve source SHA without Git metadata (#9638)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need to identify the exact source build running from the persistent account menu > - PR #9508 added linked source SHA metadata when the server can inspect its Git checkout > - Production images and packaged deployments may not include a `.git` directory even though their build commit is known > - Falling back to the package version in those environments makes the UI look like a formal release and hides the source SHA > - This pull request reads a validated deployment commit marker when Git metadata is unavailable and uses it consistently for server version and server-info responses > - The benefit is that unreleased deployments keep showing an inspectable SHA without changing exact-tag release versions ## Linked Issues or Issue Description Follow-up to #9508. ### Pre-submission checklist - [x] I searched existing open and closed issues and found no duplicate for the no-`.git` deployment fallback. - [x] The behavior reproduces when the server runs without Git metadata but has a known build commit. - [x] The behavior originates in Paperclip's core server build metadata handling, not an adapter, provider, or local configuration. ### What happened? PR #9508 displays source branch and SHA metadata for unreleased builds, but server version and server-info resolution still fall back to the package version when the runtime has no `.git` directory. This is common in production images and packaged deployments. ### Expected behavior When a validated deployment commit is available through `PAPERCLIP_BUILD_COMMIT` or `/app/.paperclip-build-commit`, the server should retain a derived source version and expose SHA metadata even if Git commands are unavailable. Exact release tags should continue using the formal package version. ### Steps to reproduce 1. Build or run Paperclip without a `.git` directory. 2. Provide a full commit SHA through `PAPERCLIP_BUILD_COMMIT` or `/app/.paperclip-build-commit`. 3. Start the server and inspect the version and server-info output. 4. Observe that current `master` returns only the package version and reports Git metadata unavailable. ### Paperclip version or commit Current `master` after #9508. ### Deployment mode Packaged or containerized deployments without runtime Git metadata. ### Installation method Built from source or deployment image. ## What Changed - Add validated build-commit parsing from `PAPERCLIP_BUILD_COMMIT` and `/app/.paperclip-build-commit`. - Preserve source-derived server versions when Git commands are unavailable. - Expose fallback SHA metadata through server-info with an explicit unavailable local-status state. - Keep exact release-tag builds on the formal package version. - Add focused regression tests for parsing, version resolution, and server-info fallback behavior. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/build-commit.test.ts src/__tests__/server-info.test.ts src/__tests__/version.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check public/master...HEAD` ## Risks - Low risk: only full 40-character hexadecimal commit values are accepted; malformed or truncated markers preserve the existing fallback behavior. - Deployment tooling must set `PAPERCLIP_BUILD_COMMIT` or write `/app/.paperclip-build-commit` for the fallback to activate. - Fallback server-info cannot provide branch, subject, commit time, or working-tree status without Git metadata, so those fields remain explicitly unavailable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4 with medium reasoning, repository/tool access, shell execution, and code editing; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates pass - [x] Greptile review is 5/5 with no open P2-or-higher comments, recommendations, or follow-ups --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.2 |
||
|
|
b606869a6a |
feat(skills): add active PR gardening workflow (#9510)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its skills layer gives agents repeatable operational workflows without embedding every procedure in core orchestration > - Pull requests referenced across active issue threads currently require expensive manual discovery and inconsistent readiness checks > - Candidate extraction and current-head verification can be deterministic, read-only code paths instead of LLM scanning > - This pull request adds an active PR gardening skill with discovery, readiness, reporting, and originating-issue follow-up instructions > - The benefit is a token-efficient, auditable way to identify PRs that are ready or need attention while preserving a strict no-merge guardrail ## Linked Issues or Issue Description - **Problem:** Paperclip operators lack a fleet-level workflow to discover PRs mentioned by recently active issues, verify their exact current-head readiness, and route actionable failures back to the issue that owns the work. - **Desired behavior:** Use the company extract-search API plus read-only GitHub inspection to deduplicate open PRs, classify readiness and confidence, avoid nagging drafts, and render an inspectable report. - **Safety requirements:** The workflow must never merge, approve, or close PRs; must never issue mutating GitHub calls; and must suppress Paperclip comments in dry-run mode. - No duplicate PR was found. Related historical PR #3725 concerns skill endpoint permissions rather than PR gardening. ## What Changed - Added `.agents/skills/pr-gardening/SKILL.md` with discover, verify, comment, monitor, and report stages plus cooldown and max-round guidance. - Added `find-candidates.mjs` to page through extract-search results, normalize/deduplicate PRs, map source issues and work products, and drop closed GitHub PRs. - Added `check-readiness.mjs` to inspect current-head checks, Greptile freshness, conflicts, reviews, and base distance with machine-readable reasons. - Added `render-report.mjs` to group open PRs into High, Medium, and Low merge-confidence sections. - Added focused Node tests, including an explicit assertion that scripts contain no mutating GitHub commands. ## Verification - `node --test .agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 8 tests passed. - Live read-only GitHub dry run at current heads: PR #9493 classified High/ready; PR #9473 classified Medium because it is two commits behind `master`; merged PR #9470 is excluded from readiness candidates. - The deployed control plane currently returns 404 for `/api/companies/:companyId/search/extract`; direct Stage A live execution will be repeated after the separately prepared extract-search endpoint is deployed. ## Risks - Low runtime risk: this adds a repo skill and read-only scripts, not a server execution path. - The discovery script intentionally fails closed when extract-search reports truncated matches, preventing silently incomplete reports. - Readiness reflects GitHub state at execution time and must be rerun after any head update. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4 with medium reasoning, repository tool use, shell execution, and live GitHub/Paperclip API verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.716.0-canary.1 |
||
|
|
ae77908618 |
feat(search): add bulk extract endpoint (#9507)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents and operators need company-scoped search to discover relevant issue history safely > - The interactive search endpoint intentionally returns compact excerpts and low pagination caps for UI use > - Automation that inventories repeated references, such as pull-request URLs, needs exhaustive distinct matches without loading full issue objects into an LLM context > - Client-provided regular expressions would create an unsafe and expensive query surface, so extraction must remain literal with server-owned expansion modes > - This pull request adds a bounded agent-oriented extraction endpoint with explicit truncation > - The benefit is deterministic, compact bulk discovery across issues, comments, and documents while preserving company authorization and rate limits ## Linked Issues or Issue Description ### Subsystem affected `server/` REST API and `packages/shared/` contracts. ### Problem or motivation The existing interactive company search caps issue pagination and snippets, so automation cannot reliably enumerate every distinct literal or pull-request URL across issue descriptions, comments, and linked documents without fetching large full issue payloads. ### Proposed solution Add `GET /api/companies/:companyId/search/extract` with escaped literal matching, optional server-owned URL token expansion, issue/comment/document scopes, status/date filters, higher issue-level pagination caps, compact source references, and explicit pagination/match truncation flags. ### Alternatives considered Reusing `GET /issues?q=` would return unnecessarily large issue objects; increasing interactive-search snippet limits would make the UI API heavier; accepting arbitrary client regex would expose avoidable database cost and ReDoS risk. ### Roadmap alignment `ROADMAP.md` does not currently list a conflicting company-search or bulk-extraction initiative. GitHub searches found no directly duplicative open issue or pull request. ## What Changed - Added shared query validation and response contracts for literal and URL extraction. - Added a company-scoped extraction service that pages issues, gathers matching issue/comment/document sources, expands URL tokens, deduplicates values, and reports truncation explicitly. - Added the authenticated route using the existing company-search authorization decision and rate limiter. - Added targeted Vitest coverage for URL extraction, multi-source dedupe, date/status filters, match caps, cross-company denial, and rate limiting. - Documented the extraction surface in the implementation specification. ## Verification - `pnpm exec vitest run server/src/__tests__/company-search-extract-service.test.ts server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-rate-limit-routes.test.ts server/src/__tests__/company-search-service.test.ts` — 30 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Bulk substring search can scan large text columns. The endpoint mitigates this with a minimum literal length, bounded issue pagination, a 20-distinct-match cap per issue, explicit truncation, existing company-search rate limiting, and no client-provided regex. - URL expansion uses a fixed server-owned pattern plus an escaped literal. A security review is requested as part of PR review to confirm the pattern and abuse controls. - No database migration or existing API response shape changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact runtime model ID and context-window size were not exposed to the session. Tool-enabled code execution and repository editing were used with medium reasoning effort. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.716.0-canary.0 |
||
|
|
3ae2c30f2f |
feat(skills): import skills from projects (#9620)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Company skills make reusable agent behavior discoverable and editable from one place. > - Projects already contain skill directories, but operators had to import each skill path manually. > - Copying those skills would break the desired write-through workflow between Skill Studio and the source project. > - The server therefore needs a safe preview/select/import contract that only accepts rediscovered, workspace-contained candidates. > - The UI needs a guided project picker that explains reference semantics, handles conflicts, and remains usable on mobile. > - This pull request adds that end-to-end project skill import flow with authorization, tenant-scope, traversal, and symlink regression coverage. > - The benefit is faster bulk onboarding while keeping project files as the single source of truth. ## Linked Issues or Issue Description **Feature request** **Problem:** Importing several skills already stored in a Paperclip project requires operators to discover and submit each local path individually. This is slow, hides which well-known directories were searched, and makes conflict/already-imported states difficult to evaluate before mutation. **Proposed solution:** Add an “Import skills from project” flow that previews skills from well-known directories, lets operators selectively import eligible candidates, and stores local-path references so Skill Studio edits write through to the project files. **Alternatives considered:** Copying files into company-managed skill storage was rejected because it creates divergent copies. Trusting client-supplied paths was rejected because imports must be constrained to server-rediscovered, workspace-contained candidates. **Additional context:** GitHub duplicate search found no existing issue or PR for this exact workflow. Refs #3799 for related skill-import inventory behavior; this PR does not claim to close that issue. ## What Changed - Extend `scan-projects` with backward-compatible preview and selective-import modes, typed validation, candidate statuses, and OpenAPI coverage. - Discover project skills under `skills`, `.agents/skills`, `.claude/skills`, `.codex/skills`, `.cursor/skills`, `.opencode/skills`, and `.gemini/skills`. - Re-discover selections server-side, enforce company/project/workspace scope, and reject traversal or symlink escapes before creating `local_path` references. - Add the Skills-page menu entry and responsive project import dialog with project selection, grouped candidates, select all/deselect all, conflicts, empty/error/403 states, and import results. - Add route, service, and component regressions for preview authorization, cross-tenant selections, traversal/symlink safety, selection counts, grouping, and result semantics. ### Screenshots **Choose a project**  **Review discovered skills**  **Mobile selection footer**  **Import result**  ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts ui/src/pages/skills/ImportSkillsFromProjectDialog.test.tsx` — 3 files, 81 tests passed. - `pnpm check:token-gates` — all token gates clean. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - Security review passed after adding tenant-scope and unauthorized-preview regressions; UX re-review approved desktop/mobile surfaces; QA passed all seven acceptance areas including write-through editing, deduplication, conflicts, empty state, and permission denial. ## Risks - Files remain referenced in project workspaces, so moving or deleting a source directory can make an imported skill unavailable; the UI explicitly communicates the reference behavior. - New well-known directory scans may discover more candidates than older versions, but preview mode prevents mutation until the operator confirms a selection. - The endpoint remains backward compatible: omitting `mode` preserves the prior full-import behavior. - No schema migration or telemetry event changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 with tool use/code execution assisted with the UI implementation and UX polish. OpenAI Codex CLI with tool use/code execution assisted with server implementation, security fixes, regression coverage, integration, and PR preparation; the runtime did not expose Codex's exact backing model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.715.0-canary.6 |
||
|
|
9af96461d5 |
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work. > - Its server recovery layer classifies blocked issue graphs and restores interrupted heartbeat execution. > - A dependent issue could remain dispatch-suppressed by a cancelled blocker without producing operator-visible attention when the dependent still displayed as todo or backlog. > - Separately, a monitor-triggered run that lost its process before disposition could consume the monitor's one-shot wake without scheduling the existing bounded continuation. > - Both gaps strand useful work even though Paperclip already has the relevant blocker-attention and process-loss recovery mechanisms. > - This pull request widens the existing classification path and reuses the single process-loss retry for monitor dispatches with no future wake. > - The benefit is visible, routable recovery without weakening dependency checkout rules or introducing an unbounded retry loop. ## Linked Issues or Issue Description No matching public GitHub issue or pull request was found. ### What happened? Two server recovery cases could leave work stranded: 1. A non-terminal, agent-assigned issue with an unresolved cancelled blocker remained ineligible for checkout, but blocked-chain liveness classification only inspected issues already displaying `blocked` or `in_review`, so the existing `blocked_by_cancelled_issue` attention was not surfaced. 2. A one-shot issue monitor cleared its next check when dispatched. If that monitor-triggered run ended as `process_lost` without a tracked local child, the existing bounded retry gate rejected it and no future monitor wake remained. ### Expected behavior - Cancelled blockers continue to be unresolved dependencies, and their dependents receive blocker attention regardless of whether the dependent currently displays as backlog, todo, blocked, or in review. - A monitor-triggered run lost before disposition receives exactly one bounded continuation when no future monitor check exists; a second loss follows the normal recovery-action escalation path. ### Steps to reproduce 1. Create an agent-assigned todo issue blocked by a cancelled issue and run issue-graph liveness classification. 2. Observe that no cancelled-blocker finding appears before this change. 3. Dispatch a due issue monitor, clear its one-shot `monitorNextCheckAt`, and mark the resulting untracked run `process_lost`. 4. Observe that no retry is queued before this change. ### Environment - Paperclip commit: `3e348b96b` - Deployment: built from source / local test environment - Adapter: not adapter-specific; core server recovery - Database: embedded test database ## What Changed - Inspect non-terminal, agent-assigned issues with unresolved blocker edges during blocked-chain liveness classification. - Include cancelled dependents in the existing blocked-inbox attention query while preserving company-scoped relation checks. - Allow monitor-triggered `process_lost` runs with no future monitor wake to use the existing single bounded retry. - Mark monitor recovery retries as continuation-needed context and retain the existing second-loss escalation behavior. - Document cancelled-blocker and monitor-dispatch recovery semantics. - Add focused regressions for liveness findings, attention propagation, one retry, and second-loss escalation. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-blocker-attention.test.ts server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Low risk and server-only. The liveness scan inspects more unresolved dependency shapes, which can produce additional existing attention entries for previously invisible cancelled blockers. - Monitor recovery remains bounded by `processLossRetryCount < 1`, and the extra path only applies when the dispatch was monitor-triggered and no future monitor check exists. - No schema, migration, authorization, API-contract, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool use, command execution, and test execution capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.715.0-canary.5 |