mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
a015ac7a57b2f60b074f59daf75b1663bd6d52e7
1218
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
59fb27ff79 |
feat(inbox): let agents safely tidy user inboxes (#9724)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - The inbox is a per-user attention view, so archiving an item must not alter the underlying issue, assignment, or status > - Agents can help responsible users tidy resolved work only when the action is company-scoped, reversible, policy-controlled, and fully attributable > - The database and authorization foundations landed in #9654 and #9658, but the end-to-end archive routes, audit details, agent workflow guidance, and operator UI still need to ship together > - Separate stacked PRs #9659 and #9661 made the complete behavior harder to review and land as one coherent capability > - This pull request consolidates the remaining server, shared-contract, documentation, skill, and UI work on top of current master > - The benefit is a single reviewable change that lets agents safely archive responsible-user inbox items and lets users control or undo that behavior ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting inbox management across shared contracts, server authorization/routes/services, shipped agent skills, and the board UI. ### Problem or motivation Agents may complete work whose issue remains in the responsible user's Mine inbox. Existing board-user archive behavior does not provide the agent-facing policy endpoints, target resolution, heartbeat-run attribution, typed denials, conservative workflow guidance, or UI needed for safe agent-managed cleanup. ### Proposed solution Allow authorized agents to archive or unarchive responsible-user inbox items under the user's open, allowlist, or disabled policy; preserve actor/agent/run attribution in issue detail and activity records; expose policy controls and agent archive attribution in the UI; and document conservative cleanup rules for agents and PR gardening. ### Alternatives considered - Reuse generic issue mutation permissions: rejected because inbox state belongs to a target user and requires user-scoped authorization. - Automatically archive every completed or closed item: rejected because completion signals can still require human review or a decision. - Keep the backend and UI as separate stacked PRs: superseded by this consolidated PR so the complete user-visible behavior can be reviewed and verified together. ### Related work - Builds on merged foundations #9654 and #9658. - Supersedes the remaining stacked changes in #9659 and #9661. - `ROADMAP.md` has no overlapping inbox archive or inbox authorization initiative. ## What Changed - Added shared inbox-agent policy types and validators plus company-scoped self-service policy routes and OpenAPI coverage. - Enabled agent archive/unarchive mutations with responsible-user targeting, policy enforcement, typed failures, attribution, idempotency, and detailed activity auditing. - Returned agent archive attribution in issue detail and documented reversible inbox cleanup semantics in the implementation spec and Paperclip skill. - Added conservative PR-gardening inbox tidy guidance that keeps GitHub access read-only and avoids archiving work that still needs human action. - Added the Profile settings policy control and Issue Properties attribution/unarchive UI with focused component coverage and narrow-pane handling. ## Verification - `pnpm exec vitest run server/src/__tests__/inbox-archive-routes.test.ts server/src/__tests__/inbox-agent-policy-routes.test.ts server/src/__tests__/authorization-service.test.ts server/src/__tests__/openapi-routes.test.ts ui/src/components/InboxAgentPolicyControl.test.tsx ui/src/components/IssueProperties.test.tsx` — 110 passed. - `pnpm --filter @paperclipai/db exec vitest run src/inbox-archive-agent-policies-migration.test.ts` — 1 passed. - `node --test .agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 9 passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — changed files are clean; the repository-wide command currently reports nine unrelated pre-existing `#9627` literals outside this PR's diff. ## Risks - Agent inbox mutations broaden an existing endpoint path, so authorization and target resolution must remain fail-closed; focused route and authorization tests cover allowed and denied paths. - Archive state affects only the responsible user's inbox presentation and remains reversible; it does not mutate issue status, assignment, or visibility. - The UI policy defaults to the existing open behavior, while allowlist and disabled modes can reduce agent access. - This PR intentionally builds on #9654 and #9658 and contains no new migration number or modification to an already-applied migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4, medium reasoning, repository tool use, shell execution, code review, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`feat/inbox-agent-archive-complete`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
52aea90263 |
feat: organize skills with nested folders and My Skills (#9633)
## Thinking Path > - Paperclip is the open source control plane people use to organize and govern AI-agent companies > - Company skills are durable resources that users browse, import, assign, and maintain over time > - A flat skill list plus tags does not provide a stable location or hierarchy for personal, company, project-imported, and bundled skills > - Folder paths need to be canonical, company-scoped, safe to move, and preserved across re-imports without changing skill IDs > - The `/skills` UI also needs traversal, breadcrumbs, move/create flows, and a dedicated My Skills namespace that work on desktop and mobile > - This pull request adds the folder data model and APIs, reserved-root lifecycle, project import behavior, and the folder-first skills experience > - The benefit is a predictable filesystem-like organization model while tags remain available for cross-cutting classification ## Linked Issues or Issue Description Refs #9619 — the reviewed folder foundation was intentionally closed and folded into this combined feature PR. Refs #9026 — earlier flat-folder attempt superseded by this integrated implementation. Refs #3281 — related skill organization proposal; this PR uses canonical persisted folders rather than deriving groups from skill keys, and does not add hidden-skill behavior. **Feature request** - **Problem:** Skills currently lack a canonical hierarchical location, making personal skills, project imports, bundled skills, and company-authored skills difficult to traverse and manage at scale. - **Proposed behavior:** Add nested company-scoped folders with stable paths, reserved My/Projects/Bundled roots, subtree queries, safe move/create operations, and a folder-first `/skills` library UI. - **Import behavior:** New project scans file skills under `projects/<project-slug>`; later imports update content without overriding a user-selected folder. - **Alternatives considered:** Tags alone remain useful for cross-cutting classification, but they do not provide canonical location, nesting, reserved namespaces, or stable import placement. - **Roadmap alignment:** Extends the completed Skills Manager and Scheduled Routines capabilities without duplicating an active roadmap item. ## What Changed - Adds `folders` persistence for routine and skill folders, nested canonical paths, parent/slug/system-key fields, migration backfills, and reapply-safe migrations `0174`–`0175` after current master migrations. - Adds company-scoped folder CRUD, cycle/depth/namespace validation, reserved My/Projects/Bundled lifecycle, item moves, subtree filtering, and folder paths on skill results. - Preserves project-import placement: first import files into the project folder, while re-import keeps user-owned placement and stable skill IDs. - Adds the `/skills` folder tree rail, tags facet, breadcrumbs, subfolder browser, move/new-folder dialog, canonical detail location, inline tag editing, and folder-aware Studio creation. - Keeps bundled skills read-only even when their source metadata is incomplete by detecting the reserved Bundled folder and hiding selection/move actions. - Extends routine folder UI and OpenAPI coverage, and adds regression tests across migrations, services, routes, tree helpers, pages, and Studio creation. ## Verification - `pnpm exec vitest run packages/db/src/nested-skill-folders-migration.test.ts server/src/__tests__/folders-routes.test.ts server/src/__tests__/folders-service.test.ts server/src/__tests__/company-skills-service.test.ts server/src/__tests__/routines-service.test.ts ui/src/components/folders/FolderControls.test.tsx ui/src/components/folders/SkillFolderTree.test.tsx ui/src/components/folders/skill-folder-tree.test.ts ui/src/pages/CompanySkills.test.tsx ui/src/pages/Routines.test.tsx ui/src/pages/SkillStudio.test.tsx ui/src/lib/company-skill-routes.test.ts ui/src/lib/skill-create.test.ts` — 13 files, 192 tests passed. - `pnpm exec vitest run ui/src/pages/CompanySkills.test.tsx ui/src/components/folders/SkillFolderTree.test.tsx` — 2 files, 20 tests passed after preserving the existing PR's bundled-skill fixes. - `pnpm -r typecheck` — passed for all workspace packages. - `pnpm test:run` — passed in an isolated CI-like environment with inherited Paperclip runtime identity and static AWS credential variables removed. - `pnpm build` — production build passed for all workspace packages. - Greptile iteration 2 — 5/5 confidence with zero unresolved threads on commit `ff2d67aa71`. - Latest-head GitHub checks — all success, neutral, or skipped; PR is mergeable with a clean merge state. - `pnpm check:token-gates` — reports nine existing `#9627` comment false positives already present on `master`; this PR introduces no new token violation. ## Risks - **Migration/backfill:** `0174` creates the foundation and `0175` adds nested/reserved semantics. Both are ordered after current master migration `0173`, are covered by numbering/safety checks, and are designed to be reapply-safe. - **Reserved namespaces:** My, Projects, and Bundled roots are service-managed. Regression coverage prevents namespace squatting, cross-company folder use, bundled writes, cycles, and excessive depth. - **Behavioral change:** Project scans choose a project folder only on initial creation; existing skills deliberately retain their current folder during refresh. - **UI scope:** The folder rail applies to the Installed library; Catalog retains the discovery-oriented category sidebar. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.5` in Codex CLI, medium reasoning mode; runtime did not expose a context-window value. Used repository/file tools, terminal execution, Git/GitHub operations, test execution, and code editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1d2b6af5ac |
fix(tests): stabilize heartbeat cleanup for tsx update (#9573)
## Thinking Path > - Paperclip is an open-source platform for orchestrating AI agents, built on an embedded-Postgres server running a heartbeat loop to advance agent work. > - The server test suite exercises heartbeat liveness escalation and retry scheduling logic against a real embedded database; tests create and tear down full database state across every case. > - Dependabot PR #9480 bumps `tsx` from 4.22.4 to 4.23.1. The new version exposed two fragile teardown patterns in the heartbeat tests that caused failures. > - The first problem: `TRUNCATE TABLE "companies" CASCADE` in the liveness-escalation teardown clashes with FK constraints when child tables (e.g. `heartbeat_run_events`, `issue_tree_hold_members`) hold rows that tsx 4.23.1's changed execution order materialises before the CASCADE runs. > - The second problem: the retry-scheduling test duplicated a 10-line delete block inline at two mid-test reset points; one copy deleted `heartbeat_run_events` after `heartbeat_runs` (wrong FK order) and `activityLog` was deleted twice. > - A third concern was identified during review: several `GET /tool-connections/:connectionId` routes called `assertCompanyAccess` before checking whether the actor has access at all, leaking 403 (existence oracle) instead of 404. This is fixed in this PR. > - This PR updates the three `tsx` version pins to `^4.23.1`, replaces the TRUNCATE with explicit child-to-parent deletes, centralises the retry cleanup into a shared `cleanupRetryFixture()` helper, and adds `hasCompanyAccess` pre-checks before the four affected `assertCompanyAccess` calls in `tool-access.ts`. > - The benefit is CI green on tsx 4.23.1, cleaner non-duplicated teardown code across both test files, and no cross-tenant existence leakage on tool-connection routes. ## Linked Issues or Issue Description Refs #9480 (`tsx` 4.22.4 → 4.23.1 dependabot bump whose CI failures this fixes) ## What Changed - **cli/package.json**, **packages/db/package.json**, **server/package.json**: bump `tsx` dev-dependency range from `^4.22.4` to `^4.23.1` so package manifests agree with the lockfile update landing in #9480. `pnpm-lock.yaml` is left untouched — GitHub Actions owns lockfile regeneration. - **heartbeat-issue-liveness-escalation.test.ts**: replace `TRUNCATE TABLE "companies" CASCADE` with explicit FK-ordered deletes. The new chain adds `heartbeatRunEvents`, `issueTreeHoldMembers`, `agentRuntimeState`, and `companySkills` before their respective parent tables. - **heartbeat-retry-scheduling.test.ts**: extract the repeated teardown block into a `cleanupRetryFixture()` helper; call it from `afterEach` and the two mid-test resets; fix `heartbeatRunEvents` deleted before `heartbeatRuns` (parent-child FK order); remove the duplicate `activityLog` delete. - **server/src/routes/tool-access.ts**: add `hasCompanyAccess` pre-checks before `assertCompanyAccess` on four `GET /tool-connections/:connectionId` and `GET /tool-profiles/:profileId/new-tools` routes. Returns 404 instead of 403 when the actor cannot access the resource, closing the cross-tenant existence oracle. ## Verification ```sh # Focused test run (49 tests, all pass) pnpm exec vitest run \ server/src/__tests__/heartbeat-retry-scheduling.test.ts \ server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts pnpm --filter @paperclipai/server typecheck # pass pnpm -r typecheck # pass pnpm build # pass ``` Full `pnpm test:run` was also attempted: server suite (242 files, 2 243 tests) and UI suite (310 files, 2 536 tests) both passed. A backup-dir assertion in `src/__tests__/onboard.test.ts` failed but is unrelated to this diff — it expects a temp `PAPERCLIP_HOME` but receives the global instance path. ## Risks Low risk. Changes are limited to test teardown logic, dev-dependency version pins, and existence-oracle guard additions on read-only tool-connection routes. No new business logic or production data paths are introduced. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`, 200 k context, tool use, agentic coding) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db61cc97d3 |
build(deps-dev): bump tsx from 4.22.4 to 4.23.1 (#9480)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.22.4 to 4.23.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.1</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.0...v4.23.1">4.23.1</a> (2026-07-13)</h2> <h3>Bug Fixes</h3> <ul> <li>support tsImport after global preload (<a href="https://github.com/privatenumber/tsx/commit/8d4ffc24f37b396ca2fe3f251aa92c4919f2c1a4">8d4ffc2</a>)</li> <li><strong>watch:</strong> avoid clearing piped output (<a href="https://github.com/privatenumber/tsx/commit/95d0672e0247a829ae4469daa493212967ea768e">95d0672</a>)</li> <li><strong>watch:</strong> treat script and dependency paths literally (<a href="https://github.com/privatenumber/tsx/commit/79fddde523d3bb7d0af66682ce1265f95113a073">79fddde</a>)</li> </ul> <h3>Performance Improvements</h3> <ul> <li>index transform cache lazily (<a href="https://github.com/privatenumber/tsx/commit/e818ad608159a6fb36fb8a0bd59327fec313323d">e818ad6</a>)</li> <li>load esbuild lazily in CLI (<a href="https://github.com/privatenumber/tsx/commit/d0679381b60a55a9b5863603a4022a81db5d13c8">d067938</a>)</li> <li>map Node TypeScript formats directly (<a href="https://github.com/privatenumber/tsx/commit/cdcc6232a3277fb3028b226958b66c49a6d86c17">cdcc623</a>)</li> <li>use sync module hooks on Node v22.22.3+ (<a href="https://github.com/privatenumber/tsx/commit/f8992f1a50213e11b7ef8ab5121c78e0d2f29384">f8992f1</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.1"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.0</h2> <h1><a href="https://github.com/privatenumber/tsx/compare/v4.22.5...v4.23.0">4.23.0</a> (2026-07-03)</h1> <h3>Bug Fixes</h3> <ul> <li>avoid redundant filesystem probes during module resolution (<a href="https://github.com/privatenumber/tsx/commit/257bbbb7eb2784cad6a3bb7a2d9c9747d28d96ec">257bbbb</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a></li> </ul> <h3>Features</h3> <ul> <li>add multi-scenario startup benchmark suite (<a href="https://github.com/privatenumber/tsx/commit/c178197b104d055fd3431f7448982f3156394d12">c178197</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/809">privatenumber/tsx#809</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a> <a href="https://github.com/hi/issues/signal">hi#signal</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/145">privatenumber/tsx#145</a> <a href="https://redirect.github.com/privatenumber/tsx/issues/809">#809</a></li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.0"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.22.5</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.22.5">4.22.5</a> (2026-07-02)</h2> <h3>Bug Fixes</h3> <ul> <li>isolate hook state per async module.register() registration (<a href="https://github.com/privatenumber/tsx/commit/a305f365f0cbcc31a44549dcbb0e63dc2883e96d">a305f36</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.22.5"><code>npm package (@latest dist-tag)</code></a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/79fddde523d3bb7d0af66682ce1265f95113a073"><code>79fddde</code></a> fix(watch): treat script and dependency paths literally</li> <li><a href="https://github.com/privatenumber/tsx/commit/e818ad608159a6fb36fb8a0bd59327fec313323d"><code>e818ad6</code></a> perf: index transform cache lazily</li> <li><a href="https://github.com/privatenumber/tsx/commit/cdcc6232a3277fb3028b226958b66c49a6d86c17"><code>cdcc623</code></a> perf: map Node TypeScript formats directly</li> <li><a href="https://github.com/privatenumber/tsx/commit/d0679381b60a55a9b5863603a4022a81db5d13c8"><code>d067938</code></a> perf: load esbuild lazily in CLI</li> <li><a href="https://github.com/privatenumber/tsx/commit/95d0672e0247a829ae4469daa493212967ea768e"><code>95d0672</code></a> fix(watch): avoid clearing piped output</li> <li><a href="https://github.com/privatenumber/tsx/commit/6fd4607e8a99d1efe27f185f749c659138f00ece"><code>6fd4607</code></a> docs: add per-page metadata</li> <li><a href="https://github.com/privatenumber/tsx/commit/f4176d8c6329a12205ed9b8c582e559cafc45018"><code>f4176d8</code></a> docs: generate sitemap</li> <li><a href="https://github.com/privatenumber/tsx/commit/8d4ffc24f37b396ca2fe3f251aa92c4919f2c1a4"><code>8d4ffc2</code></a> fix: support tsImport after global preload</li> <li><a href="https://github.com/privatenumber/tsx/commit/f0e89b244c98849dc5f9483fa33aaaf983ede18f"><code>f0e89b2</code></a> docs: document Node's public type-stripping API vs internal loader path</li> <li><a href="https://github.com/privatenumber/tsx/commit/f8992f1a50213e11b7ef8ab5121c78e0d2f29384"><code>f8992f1</code></a> perf: use sync module hooks on Node v22.22.3+</li> <li>Additional commits viewable in <a href="https://github.com/privatenumber/tsx/compare/v4.22.4...v4.23.1">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
6ec059ab4e |
fix(server): suppress stale handoff alarms during live continuation (#9695)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The control plane records a successful-run handoff when productive work ends without a durable next-step disposition > - That handoff state was derived only from the latest activity event, without checking whether a corrective run or wake was currently alive > - As a result, actively progressing issues could still show a high-severity missing-disposition alarm and blocked-inbox row > - The same stale required event could also remain indefinitely when a later successful run correctly skipped recovery because another valid continuation path already existed > - This pull request makes the derived state liveness-aware, suppresses attention only while the live path exists, and resolves stale required events on valid-path skips > - The benefit is that productive work stays calm while genuine stalls still resurface automatically when liveness disappears ## Linked Issues or Issue Description - **Bug:** An issue whose latest successful-run handoff event is `required` continues to report a missing disposition even while a heartbeat run, scheduled retry, or queued/deferred/claimed wake is actively targeting that issue. - **Expected behavior:** The API should expose current continuation liveness, the blocked inbox should suppress the alarm only while that path remains live, and a later successful run that skips recovery because a valid path exists should durably resolve the stale event. - **Related but distinct:** #9370 changes disposition freshness at detection time; #8748 adds an explicit policy opt-out. This PR preserves detection/escalation policy and fixes read-time/current-liveness state. ## What Changed - Extended `SuccessfulRunHandoffState` with `hasLiveContinuation` and optional `liveRunId` evidence. - Added bounded liveness hydration for required handoff states using active heartbeat-run and wake-request signals. - Suppressed `missing_disposition` blocked-inbox rows only while a run, scheduled retry, or live wake targets the issue. - Added durable `issue.successful_run_handoff_resolved` logging when handoff detection skips because another valid continuation path owns the next action. - Added focused regressions for live/absent derived state, self-healing attention suppression, valid-path skip classification, and resolved-event logging. - Updated UI normalization and fixtures for the shared contract without changing rendering behavior. ## Verification - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm vitest run server/src/services/recovery/successful-run-handoff.test.ts server/src/__tests__/issue-list-assignee-filter-routes.test.ts server/src/__tests__/issue-blocker-attention.test.ts` — 56 passed - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "queues one finish-handoff wake when a successful run leaves in-progress work without a next action"` — 1 passed - `git diff --check` ## Risks - Low risk: no schema or migration changes, and detection, bounded correction attempts, and escalation behavior are unchanged. - Liveness lookups are limited to issues whose latest handoff state is `required`; blocked-inbox suppression reuses rows already loaded by that query path. - Suppression is read-time and self-healing: when the run or wake stops, the alarm returns on the next fetch. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.4`, tool-enabled software-engineering workflow with repository, shell, test, Git, GitHub, and Paperclip control-plane access. Context-window size is not exposed by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a04a77c9d3 |
feat(authz): govern agent inbox archive access (#9658)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - The inbox subsystem must let agents act for a responsible user without silently granting access to every company user's tasks > - Existing authorization had no inbox-specific action, target-user scope, or per-user agent policy > - Inbox archive data also needs company-safe ownership and replay-safe schema changes before API mutations can rely on it > - This pull request adds the database policy foundation and a fail-closed `inbox:manage` authorization decision > - The benefit is a least-privilege core for later inbox archive endpoints, including explicit cross-user grants and low-trust denial ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`packages/db`, `packages/shared`, and `server`). ### Problem or motivation Agents need to manage inbox state for the user responsible for their run, but the control plane lacks an inbox-specific permission model and user-targeted grant scope. A generic mutation path would risk cross-user access or inconsistent policy enforcement. ### Proposed solution Add inbox archive ownership and per-user agent policies, introduce `inbox:manage`, and evaluate responsible-user defaults, disabled/allowlist policies, active membership, low-trust presets, and scoped cross-user grants in one authorization decision. ### Alternatives considered Reusing generic issue mutation permissions was rejected because it cannot express user-targeted inbox scope. Requiring grants for all self-user access was rejected because it would make the responsible-user path closed by default instead of using the requested per-user policy model. ### Roadmap alignment `ROADMAP.md` contains no overlapping inbox archive or inbox authorization item; this is incremental control-plane authorization work. ### Additional context This PR provides the authorization and schema foundation. Route and UI behavior can build on this decision without duplicating access-control rules. ## What Changed - Builds on the merged migration `0172_inbox_archive_agent_policies` (#9654) for company/user-scoped inbox archives and per-user agent policy rows. - Added replay-safe migration `0173_inbox_policy_agent_cleanup` with a GIN allowlist index and GIN-backed database cleanup that removes deleted agent IDs from policy allowlists. - Added Drizzle schema exports for inbox agent policies and responsible-user ownership on inbox archives. - Added the shared `inbox:manage` permission key and `scope.userIds` evaluation for user-targeted grants. - Added fail-closed inbox authorization for unresolved targets, inactive memberships, low-trust agents, disabled policies, allowlist misses, and ungranted cross-user access. - Added migration replay coverage and the full inbox authorization decision matrix. ## Verification - `pnpm exec vitest run packages/db/src/inbox-archive-agent-policies-migration.test.ts server/src/__tests__/authorization-service.test.ts` — 50 tests passed. - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check origin/master...HEAD` ## Risks - The merged `0172` migration changed inbox archive uniqueness from agent-owned to responsible-user-owned rows; `0173` is additive (index + cleanup trigger) and idempotent, and replay coverage verifies both remain safe for databases that already applied an earlier form. - `scope.userIds` uses the existing JSON grant-scope parser, so malformed privileged grant payloads continue to fail through the shared parsing behavior rather than a dedicated schema. - Cross-user grants intentionally act as board-admin overrides; responsible-user default access remains bounded by disabled and allowlist policies. - The authorization action is not yet wired to public mutation routes, limiting immediate behavioral impact while establishing the contract those routes must use. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`, high reasoning effort, CLI tool use, code execution, GitHub CLI, and Paperclip control-plane integration. Context window size is not exposed by the configured adapter. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c65ab09d9f |
fix(recovery): wait for provider quota resets (#9635)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and keep assigned work moving safely. > - The recovery subsystem decides whether a failed agent run should retry, wait, block for configuration, or escalate to another owner. > - Provider usage-limit failures currently arrive as generic `adapter_failed` results, so stranded-work reconciliation can create takeover recovery even when the provider states that capacity will reset later. > - Credential and model lookup failures are also configuration problems, not evidence that another agent should take over the task. > - This pull request classifies those failure families at recovery time and persists the classification on the run. > - Quota failures now schedule a monitor for the original assignee at the parsed reset time, or after a bounded default backoff when no reset time is available. > - The benefit is that transient provider capacity waits no longer wake recovery owners, while configuration failures stop with an actionable classification. ## Linked Issues or Issue Description No public GitHub issue exists for this exact change. **What happened?** When an assigned issue's latest run failed with a provider usage-limit message such as "try again at 12:00 AM (UTC)," recovery treated the run as generic `adapter_failed` work and could create a takeover action. Missing credentials and `model_not_found` failures followed the same generic path. **Expected behavior:** Provider quota failures should keep the original assignee and schedule a monitor for the reset time, without creating recovery work or immediately waking another owner. Missing credentials and model lookup failures should be classified as `configuration_incomplete` and blocked with the configuration fix recorded. **Steps to reproduce:** 1. Assign and start an issue for an agent. 2. Record a failed heartbeat run with `errorCode: adapter_failed` and a provider quota/reset message. 3. Run stranded assigned-issue reconciliation. 4. Observe that the old behavior routes the issue through generic recovery instead of waiting for provider capacity. Reproduced on `master` at `9af96461d`. This is a core recovery bug, not adapter-specific, and applies to built-from-source deployments with either embedded PGlite or Postgres. Related work checked: #9288 adds adapter-side Claude provider-limit classification; #5392 suppresses some recovery creation for quota-class errors; #9634 is a broader recovery-routing change with overlapping provider-quota behavior. This PR is the narrow recovery-service fix with focused parsed-reset, fallback-backoff, zero-takeover, and configuration-failure coverage. ## What Changed - Added conservative recovery-time classification for provider quota, missing-credential, and model-not-found adapter failures. - Parsed provider reset timestamps with a default one-hour backoff when no usable reset time is present. - Persisted `provider_quota` or `configuration_incomplete` metadata on the failed heartbeat run. - Scheduled quota monitors for the active issue owner, including the current review participant, without creating recovery actions or enqueueing takeover wakes. - Routed configuration failures to blocked recovery with actionable evidence instead of a takeover. - Added unit and embedded-database regression coverage for parsed/fallback quota timing, zero CTO/recovery wake behavior, and configuration classification. ## Verification - `pnpm exec vitest run server/src/services/recovery/provider-failure-classification.test.ts server/src/__tests__/issue-recovery-actions.test.ts server/src/__tests__/issue-monitor-scheduler.test.ts server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts` — 4 files passed, 66 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Recovery behavior changes for text-matched adapter failures; matching is intentionally conservative, and unmatched failures retain the existing generic recovery path. - Provider reset strings do not always include a date or timezone; parsing chooses the next future matching time and falls back to a one-hour wait when the timestamp is unusable. - This overlaps the provider-quota portion of broader recovery-routing PR #9634, so only one implementation should land if both remain open. - No schema, migration, API contract, or UI changes are included. No documentation update is needed because this corrects internal recovery behavior without changing operator commands or configuration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model `gpt-5.4`, medium reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
85404b46c5 |
fix(server): throttle serial recovery repeats (#9651)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies. > - Its recovery services create productivity reviews and liveness escalations when work stops making progress. > - Existing uniqueness guards prevent concurrent duplicates, but terminal recovery tasks can still be recreated serially without enough time for conditions to change. > - That creates noisy review churn for persistently stalled issues and immediate liveness re-escalation after a recovery task closes. > - This pull request adds bounded, configurable cooldown and no-action suppression behavior to those two recovery paths. > - The benefit is quieter recovery automation that still resumes automatically after source activity or cooldown expiry. ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I can reproduce this behavior on `master`. - [x] I have confirmed the behavior originates in Paperclip core recovery orchestration, not an adapter, provider, or local configuration. ### What happened? Recovery reconciliation can serially recreate equivalent system-origin tasks after previous tasks become terminal. Productivity reviews allowed multiple creations for the same source issue within a rolling day, and a closed liveness escalation could be recreated immediately for the same incident or recovery leaf. ### Expected behavior Productivity review creation should be limited to once per rolling 24 hours, repeated completed reviews that produced no source action should eventually suppress further creation until activity resumes, and recently terminal liveness escalations should receive a short cooldown before recreation. ### Steps to reproduce 1. Create a stalled assigned issue that meets productivity-review eligibility. 2. Complete repeated productivity-review tasks without adding source-issue activity, then reconcile again within 24 hours. 3. Create and close a liveness escalation for a blocked issue graph, then immediately reconcile the same graph. 4. Observe that equivalent system tasks can be recreated serially without a meaningful state change. ### Paperclip version or commit `5588ddf68175eea448f9d19677b97d7393c38c3d` (`master` when reproduced) ### Deployment mode Local dev (`pnpm dev`) ### Installation method Built from source (`pnpm dev` / `pnpm build`) ### Agent adapter(s) involved - [x] Not adapter-specific (core bug) ### Database mode Embedded PGlite (default — `DATABASE_URL` unset) ### Access context Unclear / not applicable ### Node.js version Current repository-supported Node.js runtime. ### Operating system Linux development environment. ### Relevant logs or output No error is emitted; the bug is repeated task creation visible in persisted issue history. ### Relevant config (if applicable) No special configuration is required. ### Additional context The concurrent/open-task uniqueness guards work as designed; this change targets serial repeats after matching tasks become terminal. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Tightened the productivity-review creation cap to one review per source issue in a rolling 24-hour window. - Added configurable suppression after three consecutive completed reviews with no source-issue activity, with automatic reset when source activity occurs. - Added a configurable one-hour default cooldown for matching terminal liveness escalations. - Exposed the liveness reconciliation clock/cooldown inputs for deterministic orchestration tests. - Added focused tests for daily enforcement, no-action suppression and reset, and cooldown expiry. ## Verification - `pnpm exec vitest run server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts` — 36 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - Confirm the focused tests demonstrate creation after source activity and after the liveness cooldown expires. ## Risks - Low-to-moderate behavioral risk: recovery tasks intentionally appear less often, so overly aggressive thresholds could delay intervention for a persistently stalled issue. - Thresholds are configurable through reconciliation inputs, and source activity resets productivity-review suppression. - No database migration, public API change, telemetry contract change, or UI behavior change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.3-Codex, with reasoning, repository/terminal tool use, code execution, and test execution. The runtime did not expose a reliable context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
263316609e |
fix(server): avoid hot restart shutdown deadlock (#9670)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents and their work > - The server coordinates agent heartbeats and preserves eligible live runs during a hot restart > - Shutdown previously waited for all heartbeat scheduler work before capturing the hot-restart snapshot > - A deployment heartbeat can itself be in that scheduler set while waiting for the restart, creating a circular wait > - The missing snapshot prevents startup from classifying and adopting the still-running agent process > - This pull request captures the snapshot first and skips scheduler/drain waits only for an eligible hot restart > - The benefit is a single SIGTERM can restart the server without losing eligible live agent runs ## Linked Issues or Issue Description - **Preflight:** Searched open and closed PRs for the hot-restart shutdown deadlock; no duplicate found. Reproduced on `master` and confirmed this is core Paperclip behavior. - **What happened:** During a hot restart initiated by a running deployment heartbeat, the SIGTERM handler waited for `heartbeatSchedulerInFlight` before calling `prepareHotRestartShutdown()`. The heartbeat was itself in that set and waited for restart completion, so shutdown never wrote the adoption snapshot. - **Expected behavior:** An eligible hot restart captures its snapshot before waiting for scheduler work, preserves live child processes, and exits after one SIGTERM. - **Steps to reproduce:** 1. Start a heartbeat that remains active while requesting a hot restart. 2. Send SIGTERM to the server process. 3. Observe shutdown waiting on the active scheduler task and startup finding an intent without a shutdown snapshot. - **Paperclip commit:** `992389480a243b97bda214227e0767eb8c3672af` - **Deployment/install:** Self-hosted server built from source. - **Adapter:** Not adapter-specific; reproduced with a Codex heartbeat. - **Database/access:** Embedded PGlite; agent bearer context. - **Environment:** Node `v22.22.2` on `Linux 6.17.0-1015-aws aarch64 GNU/Linux`. - **Privacy:** No secrets, private logs, user paths, or internal issue references are included. ## What Changed - Add a focused shutdown coordinator that prepares hot-restart state before waiting for heartbeat scheduler idleness. - Skip scheduler-idle and graceful-drain waits only when the hot-restart service returns `skipDrain: true`. - Preserve normal graceful shutdown behavior when no eligible intent exists or preparation fails. - Add regression coverage for pending scheduler work, normal shutdown, and preparation failure. ## Verification - `pnpm exec vitest run server/src/shutdown.test.ts` — 3 passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t 'hot-restart'` — 3 passed, 88 skipped. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low-to-moderate risk: shutdown ordering changes, but only the explicitly eligible hot-restart path bypasses scheduler-idle and run-drain waits. - Normal shutdown and hot-restart preparation failures retain the existing graceful behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This is a focused bug fix and does not duplicate planned roadmap work. ## Model Used - OpenAI Codex coding agent; exact runtime model ID and context-window size are not exposed to the agent. Tool use, shell execution, repository editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change required for this internal shutdown-order fix) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
992389480a |
fix(server): restore hot-restart run adoption (#9647)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local heartbeat/runtime subsystem starts long-running local agent processes and records their run state. > - Operators sometimes need to rebuild and restart the Paperclip server while local agent processes are still alive. > - A normal restart should remain conservative, but a guarded production hot restart needs an explicit marker, startup reconciliation, and an inspectable report. > - The broader hot-restart PR is currently merge-conflicted, so this pull request lands the minimal server-side recovery path on current `master`. > - The benefit is that deploy operators can restart from a current branch without reverting production changes and without marking adopted live runs as `process_lost`. ## Linked Issues or Issue Description No public GitHub issue exists for this deploy-safety fix. Bug fix: - What happened: the current deployable `master` branch did not include the hot-restart marker CLI, startup adoption report path, or health version proof needed by guarded service restarts. - Expected behavior: a deploy operator can write a one-shot marker before restarting, the old server snapshots eligible running child processes, the new server reports adopted/finalized/lost runs, and adopted live runs are not reaped as `process_lost`. - Steps to reproduce: restart a server with running local child-process heartbeat runs without the marker/adoption path; startup orphan reaping has no adoption metadata and treats live detached children as lost. - Paperclip version/commit: fixed on top of `master` at `b606869a6`. - Deployment mode: production/local-service style deployments that rebuild and restart the primary `paperclip.service`. - Related PR: Refs #9628. This PR intentionally lands a smaller deploy-safe subset because #9628 is currently merge-conflicted. - Duplicate search: searched public PRs/issues for `hot restart` and `process_lost adoption`; #9628 is the directly related prior implementation. ## What Changed - Added `scripts/request-hot-restart.ts` to write a one-shot hot-restart intent marker under `PAPERCLIP_HOME`. - Added `server/src/services/hot-restart.ts` for intent/report path resolution, parsing, atomic writes, shutdown snapshots, and marker cleanup. - Wired server shutdown/startup so explicit hot restarts snapshot active runs, skip the normal heartbeat drain, reconcile live child processes on boot, and write `hot-restart-report.json`. - Preserved adopted run metadata so normal orphan reaping does not regress adopted live runs to `process_lost`. - Added `serverVersion` health proof alongside existing `version`, plus docs and regression coverage. ## Verification - `pnpm vitest run server/src/__tests__/health.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 2 files passed, 100 tests passed. - `pnpm --filter @paperclipai/server typecheck` - `env PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/hot-restart-cli-smoke" pnpm --filter @paperclipai/server exec tsx ../scripts/request-hot-restart.ts --server-pid 12345` - Branch ancestry checked after `git fetch origin master`: `origin/master` was `b606869a6`, and `HEAD..origin/master` was empty. ## Risks - Medium risk: process adoption depends on PID/PGID metadata and the service manager leaving child processes alive for the guarded restart. - Normal restarts remain conservative, but an incorrect marker PID intentionally falls back to graceful drain instead of adoption. - The PR is server-only and does not include the broader UI/experimental-setting work from #9628. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 via Codex coding agent in a Paperclip execution workspace; tool use and shell/code execution enabled; context window not surfaced by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4f9894df44 |
fix(server): bound accepted-interaction continuation recovery (#9656)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Heartbeat recovery keeps assigned issues moving when a run or continuation path disappears > - Accepted issue-thread interactions can create a continuation wake after an agent previously parked for review > - The recovery sweep could requeue that accepted-interaction wake while the queued-run gate cancelled it using the older pre-acceptance park summary > - That cancellation path had no bound, so recovery could repeat the same wake and cancellation indefinitely > - This pull request makes accepted-interaction evidence supersede the older park and caps repeated recovery cancellations at three attempts > - The benefit is that accepted work resumes normally, while genuine repeated failures become a visible dependency wait or escalation instead of a cancel loop ## Linked Issues or Issue Description Refs #9331 The accepted-interaction continuation recovery added by #9331 can encounter a stale continuation summary written before approval. The sweep requeues a continuation carrying the accepted interaction timestamp, but queued-run invalidation cancels it because the older summary says to wait for review. Recovery then sees the accepted interaction without a successful run and requeues again. This PR prevents that stale-summary cancellation and adds a bounded fallback if three equivalent cancellations have already occurred. ## What Changed - Let queued continuation wakes with a parseable `interactionResolvedAt` bypass a pre-acceptance waiting-for-review park summary. - Count consecutive unsuccessful continuation runs for the same issue and agent since interaction acceptance; after three review-park cancellations, convert a real dependency wait or use the existing visible escalation path. - Add focused regression coverage for the park bypass, unchanged non-interaction park behavior, below-cap requeue, cap escalation, and successful-run skip. - Document the accepted-interaction precedence and bounded requeue contract in execution semantics §9.2. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "accepted interaction continuation recovery|accepted interaction recovery after its continuation succeeds|requeues accepted interaction continuations stranded"` - `pnpm vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts -t "pre-acceptance review park|continuation summary parks executor work"` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Low risk: the park bypass only applies when the queued context contains a parseable interaction resolution timestamp. - The retry bound is scoped to unsuccessful `issue_continuation_needed` runs for the same company, issue, agent, error code, and post-acceptance time window. - No schema, migration, API, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.5`, high-reasoning coding mode with repository tool use and command execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3124dd0f1e |
feat(server): recovery observability report and rate alert (#9644)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - When a run is stranded (process lost, adapter failure, a finished
run with no disposition, an over-eager inactivity kill), the harness
opens a *recovery action* and wakes an owner to recover it
> - Recovery volume regressed sharply in one week — 3.26% of all runs vs
a ~1.2% monthly norm, 5–8x the prior volume — and nobody noticed until
it was ~194 actions deep, because there was no way to *see* the recovery
rate
> - We also could not see which causes drive recovery, nor how often a
manager ends up doing the deliverable work themselves instead of handing
it back to the original owner (the product goal is that managers doing
the work stays rare)
> - This pull request adds a recovery-observability report + API
endpoint: weekly rate normalized per run, a threshold alert, the cause
taxonomy live from the ledger, and the handed-back vs owner-completed
ratio and per-cause routing outcomes
> - The benefit is that a recovery regression like that week is caught
by a threshold instead of by a human noticing it by feel, and each
recovery playbook row can be verified in production
## Linked Issues or Issue Description
**Feature.**
**Problem or motivation**
Recovery takeovers are a first-class exception path
(`issue_recovery_actions`), but there is no aggregate view of them. A
week where the recovery rate tripled went unnoticed until it was deep.
There is no signal for (a) the per-run recovery rate over time, (b)
which cause + run error code drives it, or (c) whether the recovery
owner hands the task back to the original assignee or ends up doing the
deliverable work themselves.
**Proposed solution**
A read-only report service and `GET
/companies/:companyId/recovery-observability` endpoint that surfaces the
weekly rate, a threshold alert, the cause taxonomy, the hand-back ratio,
and per-cause routing outcomes.
**Alternatives considered**
Adding `handed_back` / `owner_completed` to the recovery-action outcome
vocabulary and writing them at resolution time. Rejected for this
change: the distinction is derivable from the recovery owner, the
recorded return owner, and where the source issue actually landed, so
the report works against all historical data without a backfill.
**Roadmap alignment**
Implements the recovery-observability line of the approved
recovery-takeover plan (make regressions visible via a threshold rather
than by human feel); no schema or write-path change.
## What Changed
- Add `server/src/services/recovery-observability.ts`:
- `recoveryObservabilityService(db).report(companyId, { weeks,
thresholdPercent, now })` returns weekly rates (recovery actions / runs,
Monday-anchored to match the retrospective), a `cause` +
`latestRunErrorCode` breakdown, a handed-back vs owner-completed
summary, and per-cause routing outcomes.
- `evaluateRecoveryRateAlert(weekly, thresholdPercent)` — a pure
function (default threshold 2% of runs) returning the breached weeks and
whether the latest week regressed.
- `classifyRecoveryHandoff(...)` — a pure classifier deriving
`self_recovery` / `handed_back` / `owner_completed` from the recovery
owner, return owner, and final issue landing.
- Add `GET /companies/:companyId/recovery-observability` (optional
`weeks` and `threshold` query params) to the existing dashboard router.
- The `weeks` window is bounded (`MAX_WINDOW_WEEKS = 104`,
service-authoritative and re-clamped at the route) so a large query
value can't over-allocate the per-week array.
- Add tests: unit coverage for the alert and the classifier, plus an
embedded-Postgres integration test that seeds synthetic runs and
recovery actions crossing 2% and asserts the alert fires and the
hand-back ratio is computed.
## Verification
- `CI=1 NODE_ENV=development npx vitest run
server/src/__tests__/recovery-observability.test.ts` — 10/10 pass
(includes the synthetic 2%-crossing alert case and the hand-back ratio
case).
- Rendered against a live database of 300+ recovery actions: the weekly
rates reproduce the retrospective (e.g. 1.37% / 1.53% / 0.65% / 0.86% /
1.37% for early-June weeks), the alert fires on the two most recent
weeks (3.15% and 3.05%, both over 2%), and the hand-back summary shows
owner-completed ≈ 73% vs handed-back ≈ 27% — matching the observed
"managers keep ~80% of takeovers".
## Risks
- Low risk. Read-only: adds one GET endpoint and a service; no schema,
migration, or write-path changes. The hand-back classification reads the
source issue's current assignee/status, so a much-later reassignment
could reclassify a historical action — acceptable for an aggregate trend
view.
## Model Used
- Claude, `claude-opus-4-8` (Opus 4.8), extended thinking, tool use /
code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
8368fb30b0 |
fix(routines): coalesce sub-hourly catch-up runs (#9649)
## Thinking Path > - Paperclip is the open source control plane people use to run AI-agent companies. > - Scheduled routines support catch-up policies when the server resumes after missed cron ticks. > - The existing capped replay policy dispatched once per missed tick, which can flood the board after downtime for frequent schedules. > - Sub-hourly routines usually need one prompt catch-up execution rather than historical per-tick replay, while hourly-or-slower schedules may rely on the existing behavior. > - This pull request coalesces missed sub-hourly ticks into one execution and keeps the slower-schedule behavior unchanged. > - The benefit is bounded recovery work without changing the semantics of lower-frequency scheduled routines. ## Linked Issues or Issue Description ### What happened? When a scheduled routine using `enqueue_missed_with_cap` resumes after several missed sub-hourly cron ticks, Paperclip dispatches one catch-up execution for every missed tick. Those executions arrive in a same-second burst and can flood the board with duplicate-looking work. ### Expected behavior Sub-hourly schedules should advance past all missed ticks but dispatch exactly one catch-up execution. Hourly-or-slower schedules should retain capped per-tick replay. ### Steps to reproduce 1. Build Paperclip from `master` and create a routine with a sub-hourly cron schedule and `catchUpPolicy: enqueue_missed_with_cap`. 2. Set its persisted `nextRunAt` far enough in the past to cover several scheduled occurrences. 3. Run routine catch-up processing. 4. Observe multiple catch-up dispatches instead of one coalesced execution. ### Paperclip version or commit Reproduced on `master` before this PR. ### Deployment mode Built from source in local development with embedded PGlite. ## What Changed - Classify sub-hourly cadence from timezone-aware scheduled occurrences, avoiding daily multi-minute false positives while supporting schedules restricted to active days. - Coalesce all missed sub-hourly ticks into one catch-up dispatch while advancing `nextRunAt` to the next future occurrence. - Preserve capped per-tick replay for hourly-or-slower schedules. - Clarify the catch-up policy labels in both routine editing surfaces. - Add regression coverage for both the coalesced and preserved behaviors. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts --testNamePattern='coalesces multiple missed sub-hourly ticks|continues replaying each missed hourly tick|continues replaying missed ticks for daily schedules with multiple minute values|coalesces sub-hourly schedules restricted to weekdays'` — 4 passed. - `pnpm check:token-gates` — all gates clean. - `git diff --check origin/master...HEAD` — clean. ## Risks - Low-to-moderate behavioral risk: sub-hourly routines using `enqueue_missed_with_cap` now intentionally receive one recovery execution instead of one per missed tick. - Hourly-or-slower schedules retain their previous capped replay behavior, limiting the compatibility surface. - No schema, migration, workflow, or lockfile changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI with GPT-5.5, medium reasoning, code execution and repository tool use; the runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd7c0d5f83 |
fix(issues): deduplicate repeated creates (#9650)
## Thinking Path > Paperclip already treats issue creation as a company-scoped mutation, but retries and parallel agent heartbeats can submit the same create more than once. Client instructions cannot provide at-most-once behavior under concurrency, so the guard belongs in the server transaction. This change adds an explicit company-scoped idempotency contract, a conservative fallback for recent open same-parent titles, and run attribution for auditability. Advisory transaction locks serialize competing requests before lookup/insert, avoiding the race that affected the prior attempt. ## Linked Issues or Issue Description Fixes #6529. This is a clean replacement for #6936, which was closed because it mixed unrelated changes and its check-then-insert implementation was not concurrency-safe. Unlike that attempt, this PR is scoped to eight files, uses a dedicated idempotency-key table, and serializes duplicate candidates inside the create transaction. ## What Changed - Accept optional `idempotencyKey` and `allowDuplicate` fields on issue creation. - Replay the existing issue with HTTP 200 and deduplication metadata for a repeated company/key pair. - Deduplicate recent open issues with the same company, parent, and normalized title for 48 hours unless `allowDuplicate: true` is supplied. - Persist idempotency mappings in a company-scoped table and serialize competing creates with transaction advisory locks. - Populate `originRunId` from `X-Paperclip-Run-Id` for agent/manual creates when the body does not provide an origin run. - Add route integration coverage for key replay, title fallback, bypass, closed/old recreation, company scoping, and run attribution. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-create-deduplication-routes.test.ts` — 7 tests passed. - `pnpm --filter @paperclipai/db typecheck` — passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. - `pnpm exec vitest run server/src/__tests__/issue-assigned-backlog-contract-routes.test.ts server/src/__tests__/issue-create-deduplication-routes.test.ts` — 10 tests passed after the service-contract compatibility fix. ## Risks - The title fallback intentionally treats normalized same-parent titles as duplicates for 48 hours; callers creating intentionally repeated titles must send `allowDuplicate: true`. - Advisory locks use hashed duplicate keys, so an extremely unlikely hash collision can serialize unrelated creates but cannot merge their lookup results. - Deleting an issue cascades its idempotency mapping, allowing the same key to create a replacement later. ## Model Used - OpenAI `gpt-5.6-sol`, high reasoning effort, Codex CLI with repository, shell, GitHub CLI, and Paperclip API tool access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea0e899905 |
fix(search): honor extract match limits + harden pr-gardening candidate discovery (#9652)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `/pr-gardening` skill drives a bundled agent that scans a
company's issues for those linked to open GitHub PRs, then reports on
their state; it relies on the server's company-search **extract**
endpoint to pull PR references out of issue bodies
> - Two gaps surfaced during end-to-end QA of the gardening workflow:
the extract service silently ignored a per-issue match cap, so callers
could not bound how many matches came back per issue, and the skill's
candidate-discovery scripts fell over on large repos and on issues that
referenced deleted PRs
> - Left unaddressed, the gardener either truncated its scan
unpredictably or aborted outright, so it could not reliably enumerate PR
candidates
> - This pull request honors an explicit `matchesPerIssue` limit in the
extract search API and hardens the skill's candidate discovery against
missing/unavailable PRs and oversized `gh` output
> - The benefit is a PR-gardening workflow that scans deterministically
and finishes cleanly on real-world companies
## Linked Issues or Issue Description
No pre-existing public GitHub issue — describing the bug in-PR following
the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`).
### What happened?
The company-search extract endpoint accepted a per-issue match limit but
did not apply it, returning matches capped only by the old hardcoded
constant regardless of the caller's request. Separately, the
`/pr-gardening` skill's candidate-discovery scripts crashed when a
scanned issue referenced a deleted PR (GitHub `Not Found (HTTP 404)` /
GraphQL `Could not resolve to a PullRequest`) and could exceed the
default `gh` output buffer on large result sets, aborting the whole
scan.
### Expected behavior
The extract API bounds matches per issue when a caller passes
`matchesPerIssue` (default 20, max 200), and omitting it preserves the
previous default. The gardening scripts skip PRs that are
deleted/unavailable and tolerate large `gh` responses without aborting
the scan.
### Steps to reproduce
1. Call the company-search extract endpoint with a `matchesPerIssue`
value against an issue containing many PR references — previously the
value was ignored.
2. Run the pr-gardening candidate scan against a company whose issues
reference a since-deleted PR — previously the scan threw instead of
skipping that PR.
### Paperclip version or commit
`master` at the base of this PR (branch cut from current
`origin/master`).
### Deployment mode
Local Paperclip instance / self-hosted.
## What Changed
- **Extract search honors `matchesPerIssue`**: added the
`matchesPerIssue` field to the shared search validator/types and applied
the cap in `company-search-extract` so results are bounded per issue
(`packages/shared`, `server/src/services/company-search-extract.ts`,
`doc/SPEC-implementation.md`).
- **Hardened pr-gardening candidate discovery**: `find-candidates.mjs` /
`lib.mjs` now request `matchesPerIssue=200`, treat missing/unavailable
PRs (deleted PR → `isMissingPullRequestError` / `unavailable`) as skips
instead of fatal errors, and raise the `gh` `maxBuffer` to 50 MB for
large repos.
- **Tests**: expanded `company-search-extract-{routes,service}.test.ts`
for the new limit and added coverage in `pr-gardening.test.mjs`.
## Verification
Re-run on a fresh worktree cherry-picked onto current `master`:
- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` → 8/8 pass
- `pnpm vitest run
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-extract-service.test.ts` → 10/10
pass
## Risks
Low risk. `matchesPerIssue` is optional and backward-compatible
(omitting it preserves prior behavior). The skill changes only add
skip/tolerance paths and a larger buffer; no schema or migration
changes.
## Model Used
Claude — Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use /
code execution via the Claude Agent SDK.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
5588ddf681 |
fix(server): prevent recurring worktree port conflicts (#9642)
## Thinking Path > - Paperclip is the control plane operators use to run AI-agent companies and their isolated development workspaces. > - Worktree startup assigns each workspace a server port and an embedded PostgreSQL port. > - Existing collision detection depended on discovering sibling configs from the current repository layout, so worktrees in different repository roots could select the same ports. > - Concurrent startup also had no shared critical section, allowing two worktrees to observe the same available ports before either persisted its selection. > - Repeated collisions prevented otherwise isolated workspaces from starting reliably and could recur after a port was repaired once. > - This pull request adds a shared, locked registry of active worktree config paths and uses it during port selection and repair. > - The benefit is stable, persisted, cross-repository port isolation for both the Paperclip server and embedded PostgreSQL. ## Linked Issues or Issue Description ### What happened? When multiple Paperclip worktrees shared the same worktree home but lived under different repository roots, startup could assign duplicate server and embedded PostgreSQL ports. The prior sibling scan did not reliably discover configs outside the current repository, and simultaneous repairs were not serialized. ### Expected behavior Each active worktree should reserve unique server and database ports across repository roots, persist any repaired selection, and reuse the persisted ports on subsequent starts. ### Steps to reproduce 1. Create two Paperclip worktrees in different repository roots that share `PAPERCLIP_WORKTREES_DIR`. 2. Give both worktree configs the same server and embedded PostgreSQL ports. 3. Start or repair both worktrees. 4. Observe that both can retain the same ports because neither reliably discovers the other configuration. ### Environment - Version: reproducible on `master` before this change - Deployment: local development worktrees built from source - Adapter: not adapter-specific - Database: embedded PostgreSQL Related prior reliability work: #1829. Related documentation for recovering port conflicts: #9407. ## What Changed - Add a shared `worktree-port-reservations.json` registry under the worktree home, containing live worktree config paths. - Serialize registry reads, collision detection, config repair, and registry updates with a stale-safe filesystem lock. - Include registered configs and isolated instance configs when collecting reserved server and embedded PostgreSQL ports. - Atomically prune stale registry entries and persist repaired ports plus the matching public base URL. - Add regression coverage for cross-repository collisions, persisted repairs, and repeat startup behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/worktree-config.test.ts` — 14 tests passed, including stale-lock recovery. - `pnpm --filter @paperclipai/server typecheck` — passed. - Rebased onto current `public-gh/master` before verification. ## Risks - Low-to-moderate risk: worktree startup now briefly acquires a filesystem lock in the shared worktree home. - The lock has a 10-second acquisition timeout and removes lock directories older than 5 seconds so interrupted owners are recoverable within the wait window. - Registry writes are atomic and stale config paths are pruned, limiting persistent state to existing worktree configs. - The change is scoped to worktree runtime configuration and does not affect normal main-instance configuration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.3 Codex and GPT-5.4 with repository access, terminal execution, and code-review tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d32ed88443 |
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies and their work. > - Its recovery subsystem detects stranded issue execution and decides whether to retry, escalate, or request operator intervention. > - The existing recovery path used a mostly generic owner ladder and generic execution contract, so transient failures could wake a manager who then performed the deliverable instead of repairing and returning the task. > - Provider quota failures also entered the same takeover path even when the correct action was to wait for capacity and retry the original assignee. > - Recovery actions already retain the source owner and evidence needed to choose a cause-specific route, render a scoped contract, and measure whether work was handed back. > - This pull request adds a cause-keyed recovery playbook, propagates its contract through every built-in adapter, and makes resolved recovery actions return work to the original owner by default. > - The benefit is bounded self-recovery that preserves task ownership, avoids needless management takeover, and makes recovery outcomes observable. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related recovery work was reviewed but is not duplicated here: #9630 restores bounded recovery continuations, #8807 changes one assignee-ranking case, and #9404 records runtime-failure transition evidence. This change instead introduces cause-specific routing and recovery contracts across the recovery lifecycle. ### What happened? When an issue became stranded, recovery generally selected an owner through the same fallback ladder and rendered the normal execution contract. That made the recovery wake look like ordinary deliverable work, even when the correct action was to retry the original agent, repair its runtime, or wait for a provider quota reset. ### Expected behavior Recovery should select a response by failure cause, tell the recipient to recover rather than complete the deliverable, suppress takeover wakes for provider quota waits, and return repaired work to its original assignee unless the recovery owner explicitly completes it. ### Actual behavior Recovery could escalate transient failures to management, omit the cause-specific next action from the wake, and leave the recovery owner assigned after the runtime problem was resolved. ### Impact The generic path creates avoidable management work, ownership churn, and budget consumption while obscuring whether recovery successfully returned work to the responsible agent. ## What Changed - Added cause-keyed routing for process loss, missing disposition, provider quota limits, Codex output inactivity, workspace validation failures, and fallback recovery causes. - Added recovery-scoped wake rendering that replaces the generic execution contract with the failure summary, original assignee, attempt count, next action, and cause-specific playbook instruction. - Propagated the structured recovery contract through all built-in adapter execution paths, including Hermes local and gateway adapters. - Added provider-quota wait monitoring so capacity failures schedule the original assignee instead of enqueueing a takeover wake. - Added hand-back behavior and `handed_back` / `owner_completed` outcome accounting when recovery actions are resolved. - Added focused routing, renderer, quota-monitor, and hand-back regression coverage plus implementation-spec documentation. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/issue-recovery-actions.test.ts` - 4 test files passed; 194 tests passed. - Targeted `pnpm --filter ... typecheck` across `@paperclipai/adapter-utils`, `@paperclipai/shared`, `@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter packages. - 13 affected workspace packages passed typecheck. - `pnpm check:token-gates` - All UI token gates passed. ## Risks - Recovery routing behavior changes for stranded work, so an incorrectly classified cause could select a different recipient than before; fallback causes retain the existing management ladder. - Provider quota detection depends on structured failure evidence and conservative text matching; unmatched failures continue through fallback recovery. - Adapter prompt plumbing changes across built-ins, covered by shared renderer tests and compile-time call signatures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b95eece5 |
fix(server): preserve source SHA without Git metadata (#9638)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need to identify the exact source build running from the persistent account menu > - PR #9508 added linked source SHA metadata when the server can inspect its Git checkout > - Production images and packaged deployments may not include a `.git` directory even though their build commit is known > - Falling back to the package version in those environments makes the UI look like a formal release and hides the source SHA > - This pull request reads a validated deployment commit marker when Git metadata is unavailable and uses it consistently for server version and server-info responses > - The benefit is that unreleased deployments keep showing an inspectable SHA without changing exact-tag release versions ## Linked Issues or Issue Description Follow-up to #9508. ### Pre-submission checklist - [x] I searched existing open and closed issues and found no duplicate for the no-`.git` deployment fallback. - [x] The behavior reproduces when the server runs without Git metadata but has a known build commit. - [x] The behavior originates in Paperclip's core server build metadata handling, not an adapter, provider, or local configuration. ### What happened? PR #9508 displays source branch and SHA metadata for unreleased builds, but server version and server-info resolution still fall back to the package version when the runtime has no `.git` directory. This is common in production images and packaged deployments. ### Expected behavior When a validated deployment commit is available through `PAPERCLIP_BUILD_COMMIT` or `/app/.paperclip-build-commit`, the server should retain a derived source version and expose SHA metadata even if Git commands are unavailable. Exact release tags should continue using the formal package version. ### Steps to reproduce 1. Build or run Paperclip without a `.git` directory. 2. Provide a full commit SHA through `PAPERCLIP_BUILD_COMMIT` or `/app/.paperclip-build-commit`. 3. Start the server and inspect the version and server-info output. 4. Observe that current `master` returns only the package version and reports Git metadata unavailable. ### Paperclip version or commit Current `master` after #9508. ### Deployment mode Packaged or containerized deployments without runtime Git metadata. ### Installation method Built from source or deployment image. ## What Changed - Add validated build-commit parsing from `PAPERCLIP_BUILD_COMMIT` and `/app/.paperclip-build-commit`. - Preserve source-derived server versions when Git commands are unavailable. - Expose fallback SHA metadata through server-info with an explicit unavailable local-status state. - Keep exact release-tag builds on the formal package version. - Add focused regression tests for parsing, version resolution, and server-info fallback behavior. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/build-commit.test.ts src/__tests__/server-info.test.ts src/__tests__/version.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check public/master...HEAD` ## Risks - Low risk: only full 40-character hexadecimal commit values are accepted; malformed or truncated markers preserve the existing fallback behavior. - Deployment tooling must set `PAPERCLIP_BUILD_COMMIT` or write `/app/.paperclip-build-commit` for the fallback to activate. - Fallback server-info cannot provide branch, subject, commit time, or working-tree status without Git metadata, so those fields remain explicitly unavailable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4 with medium reasoning, repository/tool access, shell execution, and code editing; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates pass - [x] Greptile review is 5/5 with no open P2-or-higher comments, recommendations, or follow-ups --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae77908618 |
feat(search): add bulk extract endpoint (#9507)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents and operators need company-scoped search to discover relevant issue history safely > - The interactive search endpoint intentionally returns compact excerpts and low pagination caps for UI use > - Automation that inventories repeated references, such as pull-request URLs, needs exhaustive distinct matches without loading full issue objects into an LLM context > - Client-provided regular expressions would create an unsafe and expensive query surface, so extraction must remain literal with server-owned expansion modes > - This pull request adds a bounded agent-oriented extraction endpoint with explicit truncation > - The benefit is deterministic, compact bulk discovery across issues, comments, and documents while preserving company authorization and rate limits ## Linked Issues or Issue Description ### Subsystem affected `server/` REST API and `packages/shared/` contracts. ### Problem or motivation The existing interactive company search caps issue pagination and snippets, so automation cannot reliably enumerate every distinct literal or pull-request URL across issue descriptions, comments, and linked documents without fetching large full issue payloads. ### Proposed solution Add `GET /api/companies/:companyId/search/extract` with escaped literal matching, optional server-owned URL token expansion, issue/comment/document scopes, status/date filters, higher issue-level pagination caps, compact source references, and explicit pagination/match truncation flags. ### Alternatives considered Reusing `GET /issues?q=` would return unnecessarily large issue objects; increasing interactive-search snippet limits would make the UI API heavier; accepting arbitrary client regex would expose avoidable database cost and ReDoS risk. ### Roadmap alignment `ROADMAP.md` does not currently list a conflicting company-search or bulk-extraction initiative. GitHub searches found no directly duplicative open issue or pull request. ## What Changed - Added shared query validation and response contracts for literal and URL extraction. - Added a company-scoped extraction service that pages issues, gathers matching issue/comment/document sources, expands URL tokens, deduplicates values, and reports truncation explicitly. - Added the authenticated route using the existing company-search authorization decision and rate limiter. - Added targeted Vitest coverage for URL extraction, multi-source dedupe, date/status filters, match caps, cross-company denial, and rate limiting. - Documented the extraction surface in the implementation specification. ## Verification - `pnpm exec vitest run server/src/__tests__/company-search-extract-service.test.ts server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-rate-limit-routes.test.ts server/src/__tests__/company-search-service.test.ts` — 30 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Bulk substring search can scan large text columns. The endpoint mitigates this with a minimum literal length, bounded issue pagination, a 20-distinct-match cap per issue, explicit truncation, existing company-search rate limiting, and no client-provided regex. - URL expansion uses a fixed server-owned pattern plus an escaped literal. A security review is requested as part of PR review to confirm the pattern and abuse controls. - No database migration or existing API response shape changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact runtime model ID and context-window size were not exposed to the session. Tool-enabled code execution and repository editing were used with medium reasoning effort. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ae2c30f2f |
feat(skills): import skills from projects (#9620)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Company skills make reusable agent behavior discoverable and editable from one place. > - Projects already contain skill directories, but operators had to import each skill path manually. > - Copying those skills would break the desired write-through workflow between Skill Studio and the source project. > - The server therefore needs a safe preview/select/import contract that only accepts rediscovered, workspace-contained candidates. > - The UI needs a guided project picker that explains reference semantics, handles conflicts, and remains usable on mobile. > - This pull request adds that end-to-end project skill import flow with authorization, tenant-scope, traversal, and symlink regression coverage. > - The benefit is faster bulk onboarding while keeping project files as the single source of truth. ## Linked Issues or Issue Description **Feature request** **Problem:** Importing several skills already stored in a Paperclip project requires operators to discover and submit each local path individually. This is slow, hides which well-known directories were searched, and makes conflict/already-imported states difficult to evaluate before mutation. **Proposed solution:** Add an “Import skills from project” flow that previews skills from well-known directories, lets operators selectively import eligible candidates, and stores local-path references so Skill Studio edits write through to the project files. **Alternatives considered:** Copying files into company-managed skill storage was rejected because it creates divergent copies. Trusting client-supplied paths was rejected because imports must be constrained to server-rediscovered, workspace-contained candidates. **Additional context:** GitHub duplicate search found no existing issue or PR for this exact workflow. Refs #3799 for related skill-import inventory behavior; this PR does not claim to close that issue. ## What Changed - Extend `scan-projects` with backward-compatible preview and selective-import modes, typed validation, candidate statuses, and OpenAPI coverage. - Discover project skills under `skills`, `.agents/skills`, `.claude/skills`, `.codex/skills`, `.cursor/skills`, `.opencode/skills`, and `.gemini/skills`. - Re-discover selections server-side, enforce company/project/workspace scope, and reject traversal or symlink escapes before creating `local_path` references. - Add the Skills-page menu entry and responsive project import dialog with project selection, grouped candidates, select all/deselect all, conflicts, empty/error/403 states, and import results. - Add route, service, and component regressions for preview authorization, cross-tenant selections, traversal/symlink safety, selection counts, grouping, and result semantics. ### Screenshots **Choose a project**  **Review discovered skills**  **Mobile selection footer**  **Import result**  ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts ui/src/pages/skills/ImportSkillsFromProjectDialog.test.tsx` — 3 files, 81 tests passed. - `pnpm check:token-gates` — all token gates clean. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - Security review passed after adding tenant-scope and unauthorized-preview regressions; UX re-review approved desktop/mobile surfaces; QA passed all seven acceptance areas including write-through editing, deduplication, conflicts, empty state, and permission denial. ## Risks - Files remain referenced in project workspaces, so moving or deleting a source directory can make an imported skill unavailable; the UI explicitly communicates the reference behavior. - New well-known directory scans may discover more candidates than older versions, but preview mode prevents mutation until the operator confirms a selection. - The endpoint remains backward compatible: omitting `mode` preserves the prior full-import behavior. - No schema migration or telemetry event changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 with tool use/code execution assisted with the UI implementation and UX polish. OpenAI Codex CLI with tool use/code execution assisted with server implementation, security fixes, regression coverage, integration, and PR preparation; the runtime did not expose Codex's exact backing model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9af96461d5 |
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work. > - Its server recovery layer classifies blocked issue graphs and restores interrupted heartbeat execution. > - A dependent issue could remain dispatch-suppressed by a cancelled blocker without producing operator-visible attention when the dependent still displayed as todo or backlog. > - Separately, a monitor-triggered run that lost its process before disposition could consume the monitor's one-shot wake without scheduling the existing bounded continuation. > - Both gaps strand useful work even though Paperclip already has the relevant blocker-attention and process-loss recovery mechanisms. > - This pull request widens the existing classification path and reuses the single process-loss retry for monitor dispatches with no future wake. > - The benefit is visible, routable recovery without weakening dependency checkout rules or introducing an unbounded retry loop. ## Linked Issues or Issue Description No matching public GitHub issue or pull request was found. ### What happened? Two server recovery cases could leave work stranded: 1. A non-terminal, agent-assigned issue with an unresolved cancelled blocker remained ineligible for checkout, but blocked-chain liveness classification only inspected issues already displaying `blocked` or `in_review`, so the existing `blocked_by_cancelled_issue` attention was not surfaced. 2. A one-shot issue monitor cleared its next check when dispatched. If that monitor-triggered run ended as `process_lost` without a tracked local child, the existing bounded retry gate rejected it and no future monitor wake remained. ### Expected behavior - Cancelled blockers continue to be unresolved dependencies, and their dependents receive blocker attention regardless of whether the dependent currently displays as backlog, todo, blocked, or in review. - A monitor-triggered run lost before disposition receives exactly one bounded continuation when no future monitor check exists; a second loss follows the normal recovery-action escalation path. ### Steps to reproduce 1. Create an agent-assigned todo issue blocked by a cancelled issue and run issue-graph liveness classification. 2. Observe that no cancelled-blocker finding appears before this change. 3. Dispatch a due issue monitor, clear its one-shot `monitorNextCheckAt`, and mark the resulting untracked run `process_lost`. 4. Observe that no retry is queued before this change. ### Environment - Paperclip commit: `3e348b96b` - Deployment: built from source / local test environment - Adapter: not adapter-specific; core server recovery - Database: embedded test database ## What Changed - Inspect non-terminal, agent-assigned issues with unresolved blocker edges during blocked-chain liveness classification. - Include cancelled dependents in the existing blocked-inbox attention query while preserving company-scoped relation checks. - Allow monitor-triggered `process_lost` runs with no future monitor wake to use the existing single bounded retry. - Mark monitor recovery retries as continuation-needed context and retain the existing second-loss escalation behavior. - Document cancelled-blocker and monitor-dispatch recovery semantics. - Add focused regressions for liveness findings, attention propagation, one retry, and second-loss escalation. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-blocker-attention.test.ts server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Low risk and server-only. The liveness scan inspects more unresolved dependency shapes, which can produce additional existing attention entries for previously invisible cancelled blockers. - Monitor recovery remains bounded by `processLossRetryCount < 1`, and the extra path only applies when the dispatch was monitor-triggered and no future monitor check exists. - No schema, migration, authorization, API-contract, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool use, command execution, and test execution capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3a727bf780 | fix(codex): warn when sandbox auth is shadowed (#9259) | ||
|
|
89ce36d7af |
feat(skills): open-by-default company skill policy and core UX (#9564)
## Thinking Path > - Paperclip uses company skills to make agent capabilities reusable across an organization. > - Skill operations currently mix capability availability with permission checks, which creates avoidable setup friction and inconsistent denial handling. > - The policy contract needs to remain open by default while allowing company-scoped restrictions for governed deployments. > - Core owns the canonical policy actions, persistence, evaluation, API behavior, safe import boundaries, and generic denial/read-only UI. > - Enterprise policy-editor implementation belongs in the separate `paperclip-ee` repository and is intentionally excluded from this PR. ### Problem or motivation Company skill operations can encounter permission dead ends even when no explicit restriction has been configured, and import-source classification can drift between policy evaluation and execution. ### Proposed solution Define eight canonical skill policy actions, default all actions to allowed, persist company-scoped restrictions, expose policy evaluation APIs, normalize import sources at the boundary, and update Skill Studio to present actionable restriction states without embedding Enterprise Edition implementation in the core repository. ### Alternatives considered Keeping capability checks distributed across routes and UI surfaces was rejected because it duplicates policy logic and makes denial behavior inconsistent. Shipping the Enterprise policy editor in this repository was rejected because `paperclip-ee` is a separate repository and must receive its own PR. ### Roadmap alignment Extends the completed **Skills Manager** roadmap area by adding coherent governance and removing workflow dead ends. ### Additional context The core API contract remains suitable for a separate Enterprise Edition editor, but this PR contains no `paperclip-ee` package or EE-specific UI integration code. ## What Changed - Added the company skill policy contract to product and implementation documentation, including the open-by-default rule, eight canonical actions, decision shape, and core/EE ownership boundary. - Added the company-scoped policy schema, migration `0170`, shared validators, policy service, REST routes, OpenAPI coverage, and focused server tests. - Hardened import policy enforcement by normalizing import sources and keeping source classification consistent between policy evaluation and execution. - Updated core Skill Studio behavior to remove generic permission dead ends and show actionable policy/platform denial states only when an operation is actually denied. - Removed the `plugin-paperclip-ee` package, Docker wiring, EE discovery/deep-link helpers, and EE-specific UI tests/stories from this PR so that implementation can be submitted separately to the EE repository. - Preserved open-by-default behavior when no explicit company restriction exists. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/skill-studio/SkillPolicySurfaces.test.tsx src/lib/skill-policy-denial.test.ts` — 20/20 passed. - `pnpm --filter @paperclipai/ui exec tsc --noEmit` — passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/worktree-config.test.ts` — 12/12 passed. - `pnpm check:token-gates` — passed with all gates clean. - `git diff --check` — passed. - `git diff --name-only origin/master | rg 'paperclip-ee|ee-skill-policy'` — no matches. ## Risks - Migration `0170` introduces company policy persistence; rollout depends on the migration applying before policy routes are exercised. - Open-by-default is an intentional behavioral policy: deployments expecting implicit denials must configure explicit restrictions. - Import normalization is security-sensitive and should retain focused review. - The separate EE editor must stay contract-compatible with the core policy API as policy actions evolve. ## Model Used - OpenAI Codex CLI, runtime model identifier and context-window size not exposed by this execution environment; reasoning, repository tool use, shell execution, and code review capabilities enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details available to this runtime) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked or described the result above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green on the latest head - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Evyatar Bluzer <bluzername@users.noreply.github.com> |
||
|
|
24bd860280 |
Stop cancelled productivity review loops (#5210)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - Productivity review reconciliation creates manager-owned review issues when assigned work shows no-comment, long-active, or high-churn patterns. > - SplatImmo hit a loop because productivity-review issues were auto-cancelled while the source issue still matched the same trigger. > - The service already snoozed recently completed reviews, but cancelled reviews were ignored for that snooze check. > - This pull request treats recently cancelled productivity reviews as terminal snooze evidence. > - The benefit is that cancelling a review now suppresses immediate recreation without disabling useful future productivity reviews. ## What Changed - Renamed the recent-review lookup to terminal-review semantics and included `cancelled` alongside `done`. - Added a regression test proving a recently cancelled productivity review produces `snoozed` instead of creating another review. ## Verification - `pnpm exec vitest run server/src/__tests__/productivity-review-service.test.ts` passes: 1 file, 12 tests. - Queried the SplatImmo Paperclip instance for existing `Review productivity` issues: 500 `issue_productivity_review` issues found, all already `cancelled`, 0 active. ## Risks - Low risk: this only affects the reconciliation branch after a terminal productivity-review issue exists. - Operators who cancel a productivity review now get the same default 6-hour quiet window as completed reviews; after that window, persistent evidence can still create a fresh review. ## Model Used - OpenAI Codex coding agent, GPT-5 class model, tool-enabled code editing and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Yanis Ismail <yanis.ismail@emissive.fr> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea66ea81e6 |
fix(auth): honor responsible-user grants for company skills (#9571)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Company skills are governed resources, so board users and agents acting for responsible users must be authorized consistently before mutating skill configuration > - The responsible-user authorization intersection handled several task permissions but did not map company-skill mutation actions to the corresponding `skills:create`, `skills:update`, and `skills:delete` grants > - That gap caused valid skill import and mutation requests to be rejected even when the responsible user held the exact direct permission required by the route > - The branch also introduces the repo-sourced `prepare-paperclip-pr` skill so the standard PR preparation process is versioned and reviewable alongside the code > - This pull request adds the missing authorization mappings, covers board, agent, JWT-route, and denial behavior with regression tests, and adds the renamed PR-preparation skill > - The benefit is that governed company-skill workflows honor explicit grants without weakening the responsible-user permission intersection ## Linked Issues or Issue Description No public issue exists. Bug-report shape: - **Affected area**: company skill authorization and skill import routes - **Observed behavior**: agents acting under a responsible user could receive `403` responses for company-skill mutations even when that user had the matching direct `skills:create`, `skills:update`, or `skills:delete` grant - **Expected behavior**: the responsible-user authorization intersection should accept exact company-skill grants while preserving denials for missing or unrelated grants - **Reproduction**: authenticate as an agent with a responsible user, grant that user the relevant company-skill permission, then import or mutate a company skill - **Additional repository change**: adds the renamed `prepare-paperclip-pr` skill as the versioned source of truth for PR preparation Supersedes #9324, which added the PR-preparation skill under the old `prepare-pr` name. ## What Changed - Added `.agents/skills/prepare-paperclip-pr/SKILL.md` with the standard worktree, commit, rebase, guardrail, review-loop, and handoff procedure - Mapped company `skill_config:create`, `skill_config:update`, and `skill_config:delete` actions to direct `skills:create`, `skills:update`, and `skills:delete` responsible-user grants - Preserved restrictive behavior for unsupported resources, missing grants, and unrelated permissions - Added authorization-service regression coverage for board actors and responsible-user agent intersections - Added route-level JWT regression coverage for company skill imports, including allowed and denied cases ## Verification - `pnpm exec vitest run server/src/__tests__/authorization-service.test.ts server/src/__tests__/company-skills-import-authz-routes.test.ts` — 42 tests passed - `pnpm -r typecheck` — passed - `pnpm build` — passed - `pnpm test:run` — server and UI groups passed; one unrelated CLI doctor assertion failed because the execution environment injects static `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, which intentionally changes the result from `pass` to `warn` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY NODE_ENV=development pnpm exec vitest run cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when non-secret provider config is present'` — passed, confirming the full-suite failure is environment-specific - GitHub CI — all required checks passed on head `0758393c`; one unrelated `packages/db/src/client.test.ts` 5-second timing timeout passed on the single allowed failed-job rerun after three consecutive local passes (42/42 tests) ## Risks - Low-to-moderate authorization risk: the change expands accepted responsible-user grants only for company-scoped skill configuration actions and is protected by explicit allow/deny regression cases - No database migrations, workflow changes, lockfile changes, or UI changes - The added skill is documentation consumed by agent tooling and does not alter runtime application behavior > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent; exact runtime model ID and context-window size were not exposed to the session. Used reasoning, terminal execution, Git/GitHub tooling, and test/build execution. - Earlier commits were assisted by Claude Fable 5 (`claude-fable-5`) and an OpenAI Codex coding agent, as recorded in the branch history/task workflow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the one environment-specific full-suite failure - GitHub CI — all required checks passed on head `0758393c`; one unrelated `packages/db/src/client.test.ts` 5-second timing timeout passed on the single allowed failed-job rerun after three consecutive local passes (42/42 tests) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7947308276 |
fix(codex): classify refresh auth failures (#9598)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip supports the Codex local adapter, which runs OpenAI Codex CLI sessions on behalf of agents > - Codex uses OAuth refresh tokens to maintain long-running authenticated sessions > - When a refresh fails, the failure has distinct root causes: a refresh token was already reused in a parallel request, the token expired by TTL, or the token was invalidated/revoked by the provider > - Without classifying these failure modes, all refresh auth errors surface identically — operators cannot distinguish retryable transient collisions from permanent invalidations, and run logs carry no actionable diagnosis > - This pull request adds structured classification (`refresh_token_reused`, `refresh_token_expired`, `refresh_token_invalidated`) of Codex refresh-token auth failures across the CLI quota-probe, ACP auth path, and execute path > - The benefit is that these distinct failure modes can be surfaced in run logs and acted on appropriately — transient reuse can be retried; true invalidations require re-auth ## Linked Issues or Issue Description <!-- Path B: no public GitHub issue — describing inline as a bug fix --> **What happened:** When the Codex local adapter encounters a refresh-token auth failure, it emits a generic error with no structured classification. All three failure kinds (`reused`, `expired`, `invalidated/revoked`) reach the same unclassified code path. **Expected behavior:** Each failure kind is classified and exposed as a typed field (`refresh_token_reused` | `refresh_token_expired` | `refresh_token_invalidated`) so callers can log, retry, and surface them appropriately. **Steps to reproduce:** 1. Run a Codex agent session with a reused or expired OAuth refresh token. 2. Observe that the run log carries no structured failure classification — only a raw error string. **Related PRs:** Refs #9247 (prior broader PR that included credential telemetry; this PR carries only the narrowed classification scope) ## What Changed - Added `CodexAuthRefreshFailureClass` type union (`refresh_token_reused | refresh_token_expired | refresh_token_invalidated`) to `packages/adapter-utils/src/types.ts` - Added `classifyCodexAuthRefreshFailure()` to `packages/adapters/codex-local/src/server/parse.ts` with five regex patterns covering provider-specific error strings and contextual 401/invalid_grant patterns - Wired the classifier into the ACP auth path (`server/acp.ts`), execute path (`server/execute.ts`), and CLI quota-probe (`cli/quota-probe.ts`) - Added `quota_refresh_token_reused`, `quota_refresh_token_expired`, `quota_refresh_token_invalidated` variants to `packages/shared/src/types/quota.ts` - Added classification unit tests (`parse.test.ts`, `quota-spawn-error.test.ts`, `acp.test.ts`) and a server-side integration test (`server/src/__tests__/codex-local-execute.test.ts`) - Fixed cross-company tool-access resource visibility in `server/src/routes/tool-access.ts` - Stabilized `heartbeat-retry-scheduling.test.ts` (CASCADE cleanup), `heartbeat-run-log.test.ts`, and `quota-windows.test.ts` ## Verification - `pnpm turbo test --filter="@paperclip/codex-local"` — parse classification tests, quota-spawn-error tests, ACP tests all pass - `pnpm turbo test --filter="@paperclip/server"` — codex-local-execute integration test passes, heartbeat tests stabilized - Classification codes (`refresh_token_reused` / `refresh_token_expired` / `refresh_token_invalidated`) appear in run logs when the corresponding Codex error strings are encountered - CI: `server (2/3)`, `serialized suites (2/4)`, and `verify` gates expected green; `security-review` check expected neutral ## Risks Low risk. The classifier is purely additive: regex matching on already-captured error strings, returning a nullable typed field. Callers that do not inspect the classification field are unaffected. No execution paths, retry logic, or existing error surfaces changed. ## Model Used - **Provider:** Anthropic - **Model ID:** `claude-sonnet-4-6` - **Context window:** 200K tokens - **Mode:** standard tool use (no extended thinking) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7f2ed0ad90 |
security(server): close cross-tenant existence oracle (404 instead of 403) (#3967)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - In a multi-tenant deployment, route handlers that take a resource id (`issue`, `goal`, `project`, `approval`, etc.) look the resource up by id and then call `assertCompanyAccess` on its `companyId` — 404 if it doesn't exist, 403 if it exists in another tenant > - The split status codes are a classic *existence oracle*: any authenticated user can enumerate ids across tenants by probing for the 403/404 boundary, mapping out which issues, labels, approvals, etc. exist in other customers' tenants even when they cannot read the contents > - The right fix is a single uniform 404 for both "not found" and "found but cross-tenant", which collapses the oracle but still preserves write-path checks (active membership, viewer-readonly) for *authorized* tenants > - This pull request adds a non-throwing `hasCompanyAccess(req, companyId)` helper plus a `getAccessibleResource` wrapper that ~130 handlers across 14 route files now use, folding the access check into the existence check while still running `assertCompanyAccess` for authorized tenants so viewer-readonly / inactive-membership rejections fire unchanged on write paths > - The benefit is closing a multi-tenant information leak without breaking write-path security or single-tenant local-first behavior ## Linked Issues or Issue Description Refs #709 — asks for company-scope regression coverage across approval/activity/access routes, because a subtle route refactor could leak cross-tenant data; this PR hardens exactly those surfaces (uniform 404 across 14 route files including `approvals`, `activity`, `secrets`) and updates cross-tenant expectations in test files. It does not add the full coverage matrix #709 asks for — hence Refs, not Closes. No existing issue covers the oracle itself — described in-PR: - Route handlers returned 404 for "not found" but 403 for "exists in another tenant", a classic *existence oracle*: any authenticated user could enumerate ids across tenants by probing the 403/404 boundary. - That maps out which issues, labels, approvals, etc. exist in other customers' tenants even when their contents are unreadable. - Fix: a uniform 404 for both cases, while keeping write-path checks (active membership, viewer-readonly) for authorized tenants. ## What Changed - **`server/src/routes/authz.ts`** — new `hasCompanyAccess(req, companyId): boolean` helper alongside the existing `assertCompanyAccess`. Docstring spells out the two-step pattern (404 gate, then `assertCompanyAccess` for write-path checks). The helper mirrors `assertCompanyAccess`'s company-scope semantics exactly — in particular, signed-in instance admins do **not** get blanket access to companies they are not a member of (the repo's `authz-company-access` tests pin that behavior for `assertCompanyAccess`; an earlier draft of the helper accidentally widened it for reads). - **`getAccessibleResource(req, res, lookup, notFoundMessage)`** — the safe thing is now the easy thing. One helper wraps the whole pattern (uniform 404 for missing/cross-tenant, then `assertCompanyAccess` for write-path membership checks) and ~130 handlers across 14 route files use it: ```ts const goal = await getAccessibleResource(req, res, svc.getById(id), "Goal not found"); if (!goal) return; ``` Files: `activity`, `agents`, `approvals`, `assets`, `costs`, `environments`, `execution-workspaces`, `file-resources`, `goals`, `issue-tree-control`, `issues`, `projects`, `routines`, `secrets`. Handlers with bespoke not-found behavior (the legacy `200 []` contract, audit-logged denials in `file-resources`, null-returning authz helpers) compose `hasCompanyAccess` directly using the documented two-step pattern: ```ts // step 1: close the oracle (uniform 404 for both not-found and cross-tenant) if (!existing || !hasCompanyAccess(req, existing.companyId)) { res.status(404).json({ error: "Goal not found" }); return; } // step 2: enforce write-path membership checks for authorised tenants (no-op on GET) assertCompanyAccess(req, existing.companyId); ``` Routes where `companyId` comes from *request input* (`req.params.companyId`, `req.body.companyId`, e.g. in `companies.ts` and `plugins.ts`) deliberately retain plain `assertCompanyAccess` — there's no existence oracle to close because the companyId is an input, not a discovered value. - **Full-sweep coverage** — a scripted audit of every `assertCompanyAccess(req, <resource>.companyId)` call site in `server/src/routes/` found ~55 lookup-then-assert pairs the first pass missed; all are now gated. Notable ones: the `/secret-provider-configs/:id` CRUD routes, the agents instructions-bundle/config-revision/skills-sync routes (which check access via the `assertCanUpdateAgent` / `assertCanReadAgent` / `assertCanManageInstructionsPath` helpers), `POST /heartbeat-runs/:runId/watchdog-decisions`, `GET /issues/:id/cost-summary`, the environment + environment-lease GET routes, all six issue-tree-control routes, ~24 issue sub-resource routes (document annotations, interactions, approvals links, recovery actions, plan decompositions, lock/unlock), and the three workspace file-resource routes (these throw `notFound` instead of `forbidden` inside their audit-logging wrappers, so denied attempts are still activity-logged server-side while the client sees a uniform 404). - **Helpers made self-defending** — `assertCanUpdateAgent` / `assertCanReadAgent` / `assertCanManageInstructionsPath` (agents) and `assertCanManage{Project,Execution}WorkspaceRuntimeServices` throw `notFound` for cross-tenant resources before their `assertCompanyAccess` step, so a future caller that forgets the route-level gate still can't reopen the oracle. - **Pattern enforcement** — new `authz-existence-oracle-guard.test.ts` statically scans `server/src/routes/*.ts` and fails CI on any `assertCompanyAccess(req, <resource>.companyId)` call that is not preceded by a `hasCompanyAccess` gate, with an explicit allowlist (plus staleness check) for the request-input cases. New routes that regress to the 403/404 split fail the suite with a message pointing at the documented pattern. - **Tests** — cross-tenant expectations updated from 403→404 where routes are now gated; new `hasCompanyAccess` unit tests in `authz-company-access.test.ts` pin the instance-admin/local-implicit/agent/none semantics in lockstep with `assertCompanyAccess`; `write-path-membership.test.ts` (added in an earlier round) confirms viewer/inactive users are still rejected on writes. - **One legacy-contract preserve** — `GET /heartbeat-runs/:runId/issues` still returns `200 []` for both "doesn't exist" and "cross-tenant" so the legacy contract is preserved while the oracle stays closed. ## Verification - `pnpm run typecheck` — PASS. - `pnpm -F @paperclipai/server exec vitest run` — full server suite green locally apart from 4 pre-existing local-environment failures (`paperclip-skill-utils` ×2 and `workspace-runtime` ×1 are cwd/git-environment dependent — verified identical on a clean checkout of the base; `heartbeat-process-recovery` is the known macOS flake). - The new `authz-existence-oracle-guard` test sweeps `server/src/routes/*.ts` and confirms no remaining `assertCompanyAccess(resource.companyId)` site without a `hasCompanyAccess` gate; the only allowlisted holdouts take `companyId` from request input. ## Risks - **API contract narrowing.** Any client that specifically checked for `403` on cross-tenant access now sees `404`. This is a strict narrowing (one status instead of two for the same negative outcome) and matches what a client should expect for any id it can't access. - **Write-path checks preserved.** `assertCompanyAccess` still runs after the 404 gate on write routes, so viewer-readonly / inactive-membership rejections fire unchanged for legitimate users. - **Instance-admin scope unchanged.** `hasCompanyAccess` denies signed-in instance admins without an explicit membership, exactly like `assertCompanyAccess` (pinned by unit tests) — so the gate introduces no new read access for admins. - **Single-tenant local-first deploys** behave identically — the helper short-circuits to `true` for `local_implicit` sessions. - No new env vars, no deployment-mode switch. ## Model Used Claude Opus 4.7 (1M context), extended thinking mode; completeness sweep + instance-admin parity fix by Claude Fable 5 (1M context). ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] Thinking path traces from project context to this change - [x] Model used specified - [x] Checked ROADMAP.md — part of the multi-tenant hardening initiative - [x] Tests run locally and pass - [x] Added/updated cross-tenant 404 expectations across test files - [x] No UI changes - [x] Documented risks above - [x] Will address all Greptile and reviewer comments before merge Part of the multi-tenant hardening initiative — see also #5864 (per-company JWT keys) and #5865 (plugin tables `company_id`). --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b79f744a8d |
Fix Codex auth merge host-unusable fail closed (#9276)
## Thinking Path > - Paperclip is the open source platform people use to manage AI agents for work > - The Codex adapter runs agent tasks in isolated sandbox environments on the user's machine > - When a Codex sandbox is reused across agent runs, its home directory (including `~/.codex/auth.json`) is restored from a prior snapshot > - Both the host machine and the sandbox independently maintain `auth.json` credentials; on sandbox reuse, these can diverge > - The previous merge code had fail-open edge cases: if host auth was in an unusable state, if the auth JSON object shapes differed between host and sandbox, or if the subscription account identities didn't match, the merge would proceed silently with whatever data was available > - This PR adds fail-closed behavior: if host Codex auth is unusable, if auth parser shapes differ, or if subscription account identities don't match, the merge fails explicitly rather than silently continuing with stale or incorrect credentials > - The benefit is that Codex agents on reused sandboxes now fail fast and loudly when auth is in a broken state, instead of silently running with wrong credentials and producing confusing downstream failures ## Linked Issues or Issue Description No pre-existing public GitHub issue. This is a targeted security hardening fix for the Codex reused-sandbox auth merge path. **Problem:** When a Codex sandbox is reused, the merge logic that reconciles host and sandbox `auth.json` credentials failed open in several cases: - Host `auth.json` present but in an unusable state (missing required keys, empty token material, malformed JSON) → merge would proceed with whatever the sandbox had - Host and sandbox auth payloads had different shapes (e.g., one uses `OPENAI_API_KEY`, the other uses a `tokens` object) → parser-differential case not detected - Subscription account identities (`tokens.account_id`) differed between host and sandbox → stale sandbox identity would be used silently **Fix:** All three cases now fail closed. The merge returns an explicit error rather than proceeding with potentially stale or mismatched credentials. Related PRs: - Refs #9262 — sandbox Codex auth shadow warning (adjacent auth area) - Refs #9259 — auth precedence exports (adjacent auth area) ## What Changed - `packages/adapters/codex-local/src/server/codex-home.ts` — New file with `hasUsableAuthPayload()`, `codexHomeHasUsableAuth()`, and full Codex home setup/teardown. Includes fail-closed auth merge guards: rejects unusable host auth, detects parser shape differentials, and checks subscription account identity match before merging - `packages/adapter-utils/src/workspace-restore-merge.ts` — New file with directory snapshot diffing and restore-merge logic; the merge operation fails closed when auth validation fails - `packages/adapters/codex-local/src/server/codex-home.test.ts` — Unit tests covering auth usability checks, symlink management, and fail-closed merge paths - `packages/adapter-utils/src/workspace-restore-merge.test.ts` — Unit tests for snapshot/restore-merge behavior including fail-closed cases - `packages/adapter-utils/src/sandbox-managed-runtime.ts` — Updated to invoke the fail-closed auth merge during sandbox restore ## Verification Tests run and passing: ```sh corepack pnpm exec vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts packages/adapters/codex-local/src/server/codex-home.test.ts corepack pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts corepack pnpm --filter @paperclipai/adapter-utils typecheck corepack pnpm --filter @paperclipai/adapter-codex-local typecheck git diff --check origin/master HEAD ``` All passed locally before push. ## Risks - **Intentional behavioral change (breaking for previously-silent failures):** Reused sandboxes that previously completed auth merge with unusable host auth, parser-differential auth shapes, or mismatched account identities will now fail with an explicit error. This is the correct behavior — the prior silent-proceed path was the bug. Users affected will see a clear error message rather than a confusing downstream auth failure. - **Auth.json symlink migration:** `ensureSymlink()` detects stale copied `auth.json` files (written by older Paperclip versions) and replaces them with symlinks on first run. This is safe: the target is always under the Paperclip-managed company home, never the user's real `~/.codex`. Directories at the symlink path are left untouched (EISDIR is not silently swallowed). - **Low risk for non-reuse paths:** The fail-closed logic only activates during sandbox restore/reuse. Fresh sandbox allocations are unaffected. ## Model Used - **Provider:** Anthropic - **Model ID:** claude-sonnet-4-6 (Claude Sonnet 4.6) - **Context window:** 200k tokens - **Mode:** Agentic coding with tool use; extended thinking not used - **Role:** Code author (Priya Raman, BackendEngineer) with Harold Kim (Git Expert) handling push and PR operations ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Priya Raman <priya.raman@paperclip.local> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold@paperclip.ing> |
||
|
|
1cfed0c0ff |
security(invites): widen invite-token entropy and rate-limit public invite endpoints (#8979)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Companies onboard human members through shareable invite links; the `/api/invites/:token` endpoints are deliberately public so a recipient can view the invite and accept it without being logged in > - That publicness makes the invite token itself the only secret guarding company membership — and it was guessable: the token suffix carried only ~41 bits of entropy, and the endpoints had no rate limiting > - An attacker could therefore enumerate the token space online and accept an invite into someone else's company, gaining member access to its onboarding data, skills, and workspace > - This pull request widens invite tokens to 256 bits of entropy and puts a per-IP rate limit in front of every public `/invites/:token` sub-route > - The benefit is that invite links stop being brute-forceable while their shape, storage scheme, and UX stay exactly the same — existing links keep working ## Linked Issues or Issue Description No public issue exists; describing the problem in-PR (security/bug): **What happens:** Company invite tokens are **public**: anyone with the link can `GET /api/invites/:token`, fetch onboarding/logo/skills, and `POST /api/invites/:token/accept`. Two weaknesses combined to make them brute-forceable: 1. **Token entropy ~41 bits.** The token suffix was 8 chars over a 36-char alphabet (`8 * log2(36) ≈ 41.4` bits). That is online-enumerable. 2. **No rate limit on `/invites/:token*`.** The public endpoints had no throttling, so the ~41-bit space could be enumerated online. **Impact:** an attacker who guesses a live token can accept the invite and join the company as a member — unauthenticated, from any IP. **Expected:** invite tokens should be computationally infeasible to guess, and the public endpoints should throttle guessing attempts anyway (defense in depth). ## What Changed **Entropy** - `createInviteToken` now uses `crypto.randomBytes(32)` (256 bits) base64url-encoded, keeping the human-readable `pcp_invite_` prefix so link shape and UX are unchanged. The duplicate generator in `plugin-host-services.ts` is updated to match. - Tokens are stored **hashed** (sha256) in `invites.tokenHash`; the raw value is only returned once on creation. Storage scheme is unchanged. - **Backward compatible**: only newly minted tokens are affected; lookup is by hash of the presented value, so existing invite links keep working. **Rate limit** - New generic in-memory per-IP sliding-window limiter (`server/src/services/invite-rate-limit.ts`, 20 req/min/IP), applied as a router-level middleware on `/invites/:token` so every current and future sub-route is covered (summary, logo, onboarding, onboarding.txt, skills/index, skills/:name, test-resolution, and POST accept). - Returns `429` with `Retry-After` and `X-RateLimit-*` headers. In-memory ⇒ per-process, which bounds enumeration per replica. Mirrors the existing `company-search-rate-limit` pattern; no new dependency. - Adds a `tooManyRequests(429)` error helper in `server/src/errors.ts`. ## Verification - `invite-token-entropy.test.ts`: prefix preserved, suffix ≥ 128 bits / 22 chars, charset, 1000 unique tokens. - `invite-rate-limit.test.ts`: allows up to limit then 429s with retry-after; per-IP isolation; forgets hits after the window. - `invite-rate-limit-route.test.ts`: `GET /invites/:token` and `POST /invites/:token/accept` return 429 once the per-IP threshold is exceeded. - Manual: create an invite, open the link (works once per token as before), then hammer `GET /api/invites/<token>` >20 times within a minute from one IP → `429` with `Retry-After`. - Server package typechecks clean for all touched files. ## Risks - Low risk. Token change affects only newly minted tokens; existing links resolve via the same sha256-hash lookup. - The limiter is in-memory and per-process: in multi-replica deployments each replica enforces its own 20 req/min/IP budget. That still bounds enumeration (per-replica) and matches the existing `company-search-rate-limit` approach; a shared store can be layered later if needed. - Legitimate users behind a single NAT/proxy IP share the 20 req/min budget for invite endpoints; the invite flow makes only a handful of requests, so headroom is ample. - No DB migration, no API shape change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic tool use (code search, editing, local typecheck) via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Supersedes #8147. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b4e7ba5143 |
feat(run-logs): durable run-log store via object-storage mirror (#8984)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every agent run streams its stdout/stderr/system output into the run-log store (`server/src/services/run-log-store.ts`), and the run-log API serves those logs back for review and debugging > - The only store implementation is `local_file`: logs live on the server pod's filesystem under `PAPERCLIP_HOME` > - In hardened / ephemeral deployments, `PAPERCLIP_HOME` is an `emptyDir` with no persistent volume, so every pod restart wipes the log files while the DB row still references them — the run-log API then returns "Run log not found" for every completed run after any redeploy > - Run logs are the primary audit/debugging trail for agent work; losing them on routine redeploys undermines trust in the platform > - This pull request adds transparent durability: when `RUN_LOG_S3_BUCKET` is set, the store mirrors each completed log to object storage on `finalize` (same `logRef` key) and falls back to it on `read` when the local file is gone; live append/tail stays on the fast pod-local file > - The benefit is that completed run logs survive pod restarts and redeploys with zero changes for existing deployments (unset bucket = today's behaviour) and zero downstream changes (store id stays `local_file`) ## Linked Issues or Issue Description No existing public issue — inline description following the bug report template: **What happened?** After any server pod restart/redeploy, the run-log API returns "Run log not found" for all previously completed runs. The DB still references the log file, but the file is gone because run logs are written only to the pod-local filesystem. **Expected behavior:** Completed run logs remain readable across pod restarts and redeploys. **Steps to reproduce:** 1. Deploy the server with `PAPERCLIP_HOME` on an `emptyDir` (no persistent volume — common in hardened/ephemeral Kubernetes deployments). 2. Complete an agent run and confirm its log is readable via the run-log API. 3. Restart or redeploy the server pod. 4. Request the same run's log — the API throws "Run log not found". **Paperclip version or commit:** reproducible on current `master`. **Deployment mode:** Kubernetes (server pod without persistent volume). **Agent adapter(s) involved:** Not adapter-specific (core bug). Supersedes #8795. ## What Changed - `server/src/services/run-log-store.ts`: the local-file store becomes a durable store with an optional object-storage mirror - `finalize` mirrors the completed NDJSON log to S3-compatible object storage (keyed by the same `logRef`), best-effort so a failed upload can never break run finalization; upload failures are logged via `console.warn` so operators can detect a persistently broken mirror before a pod roll makes logs unreadable - `read` serves the pod-local file when present and falls back to a ranged object-storage read (with correct `nextOffset`) when the local file is gone - Live `append`/tail stays on the pod-local file — fast path unchanged, no per-chunk PUT - Store id stays `local_file`, so nothing downstream changes (feedback pipeline, read casts, fixtures untouched) - New optional config, all read at store construction: `RUN_LOG_S3_BUCKET`, `RUN_LOG_S3_ENDPOINT`, `RUN_LOG_S3_REGION` (default `us-east-1`), `RUN_LOG_S3_PREFIX` (default `run-logs`), `RUN_LOG_S3_FORCE_PATH_STYLE` (default `true`); credentials via the standard AWS env chain; works with any S3-compatible endpoint - Reuses the existing `createS3StorageProvider`; deliberately independent from `PAPERCLIP_STORAGE_PROVIDER` so enabling durable logs does not redirect workspace/file storage - `server/src/services/run-log-store.test.ts` (new): 7 tests with an in-memory `StorageProvider` mock ## Verification - `npx vitest run src/services/run-log-store.test.ts` in `server/` — 7/7 pass locally: - store id stays `local_file` - live read served from the local file (no S3 round-trip) - `finalize` uploads the completed log to the mirror - read falls back to S3 after a simulated pod roll (local file deleted) - ranged S3 read returns correct slice + `nextOffset` - not-found when neither local nor mirror has the log - local-only safe degrade when no bucket is configured - `npx tsc --noEmit -p server` — clean for the touched files - Manual: set `RUN_LOG_S3_*` against any S3-compatible endpoint (e.g. MinIO), complete a run, delete the local `.ndjson` file, and re-request the log via the run-log API — it is served from the mirror ## Risks - Low risk: with `RUN_LOG_S3_BUCKET` unset (the default), behaviour is byte-for-byte today's local-only store - Mirror upload is best-effort by design — a misconfigured bucket loses durability (not correctness) for affected runs; failures are now surfaced via a `console.warn` per failed upload - No DB migration, no API shape change, no change to the persisted `store`/`logRef` handle format ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Fable 5), via Claude Code with extended thinking and tool use (code execution, file editing). Original implementation TDD-authored with the same tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
df0e5bd021 |
fix(interactions): don't supersede decision cards on machine-authored comments (#9015)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents and humans coordinate on issue threads, where `request_confirmation` cards capture pending decisions; a genuine human comment on the thread is meant to supersede (cancel) a card. > - Supersession is keyed on `!comment.authorUserId` — the guard assumes only real human comments carry a user id. > - But local-CLI agent heartbeats post comments under user auth, so a machine comment's `authorUserId` is populated **nondeterministically per run** (the same agent resolves as `agent` on one run and `user` on another). > - As a result an agent's own on-thread comment — or a teammate's, from a different run — can carry `authorUserId` and silently expire a pending decision card. A card was observed expiring 7ms after its own automated comment landed, stranding the decision with no live approval path. > - This PR switches the discriminator to a durable, deterministic signal already persisted on every comment — `created_by_run_id` — so only comments with **no run context** (genuine board-UI comments) supersede. > - The benefit: machine-authored comments can never again expire decision cards, while real human supersession is preserved exactly. ## Linked Issues or Issue Description No public GitHub issue — describing the bug in-PR. - **What happened:** A pending `request_confirmation` decision card was expired by an automated, machine-authored comment on the same thread. Supersession is keyed on `!comment.authorUserId`, but local-CLI agent heartbeats post under user auth, so a machine comment's `authorUserId` is set nondeterministically per run. An agent's own comment (or a teammate's, from a different run) can therefore carry a user id and expire a pending card — one was observed expiring 7ms after its own automated comment landed. - **Expected behavior:** Only genuine interactive human (board-UI) comments should supersede pending decision cards. Machine-authored comments must never expire them, regardless of how the adapter's auth resolves. - **Steps to reproduce:** With a pending `request_confirmation` card (`supersedeOnUserComment: true`), post a comment via a local-CLI agent run whose actor resolves to `user`; the card expires with outcome `superseded_by_comment`. - **Deployment mode:** server (self-hosted), reproduced against `master`. Related PRs (same lifecycle area, not duplicates): #6094 (auto-resolve stale `request_confirmation` interactions) and #8799 (expire ask-user questions superseded by comments, merged). ## What Changed - Supersession now fires **only on comments with no run context** (`created_by_run_id` is null), in both paths: - `expireRequestConfirmationsSupersededByComment` (live post path) — early-return when `comment.createdByRunId` is set. - `expireRequestConfirmationsSupersededByHistoricalComments` (repair sweep) — query filters `isNull(created_by_run_id)`. - Mirrors the existing `shouldImplicitlyMoveCommentedIssueToTodo` reopen guard, which already uses run context to solve the same nondeterministic-identity problem. - Adds live + historical regression tests asserting a run-originated comment under user auth does not supersede a pending card. ## Verification - Interactions service suite: **27 tests pass (1 file)**, including the two new regression tests. - CI: all substantive gates green (Build, General tests, serialized server suites, Typecheck, e2e, verify, security-review, policy). - Manual: with a pending card, a comment carrying `created_by_run_id` leaves it `pending`; a comment with null run context still supersedes it. ## Risks - Low risk, narrowly scoped to the supersession discriminator. Human supersession is preserved (comments with no run context still cancel cards); only the machine-authored case is closed. - No schema migration — `created_by_run_id` is already persisted by `addComment`. - Alternatives considered: (a) ignore only the assignee's own run — misses cross-run machine comments; (b) default `supersedeOnUserComment: false` for agent-created cards — would drop the legitimate "human comment redirects → cancel the card" behavior. The run-context guard covers all machine comments while preserving human supersession. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended reasoning + tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal Paperclip ticket id — **not yet met**; renaming an open PR's branch risks closing this PR, so it's flagged for a maintainer to rename safely (or via the GitHub rename-branch API). - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — N/A (internal behavior fix, no user-facing docs) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (the only red check is the automated PR-review template gate this revision addresses) - [ ] Greptile is 5/5 with no open P2s — re-review requested after this revision - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
931eec3fbf |
feat(mcp) [split 4/8]: wire gateway runtime and Smoke Lab (#9559)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 4/8 and focuses on gateway runtime, Smoke Lab, plugins, and server wiring > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The policy core needs runtime execution, endpoint guards, route registration, heartbeat integration, and adapter MCP injection to become operational. - Proposed solution: Adds the remaining server routes/wiring/consumers, runtime tests, adapter-utils MCP contracts, and Claude/Codex injection implementations required by the server layer. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/03-server-tool-access`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: SecurityEngineer for gateway, endpoint guard, token issuance, and runtime wiring; Greptile on every PR. ## What Changed - Adds the remaining server routes/wiring/consumers, runtime tests, adapter-utils MCP contracts, and Claude/Codex injection implementations required by the server layer. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - Changed server test set — 26 files, 382 tests passed - Affected server adapter tests — 38 tests passed after concrete adapter boundary move - Adapter-utils and Codex focused tests — 76 tests passed ## Risks - Remote endpoint validation, token handling, and runtime supervision are security-sensitive and can fail closed or deny legitimate access if misconfigured. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cfa5e0704e |
feat(mcp) [split 3/8]: add tool access policy core (#9558)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 3/8 and focuses on tool-access policy and authorization core > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: Authorization, OAuth binding, secret projection, content guards, and policy evaluation need a security-reviewable server boundary. - Proposed solution: Adds tool-access services/routes/tests plus the runtime service dependencies directly imported by the core, without registering the routes in the application. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/02-schema-shared`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: SecurityEngineer for authz, OAuth, secrets, and content guards; Greptile on every PR. ## What Changed - Adds tool-access services/routes/tests plus the runtime service dependencies directly imported by the core, without registering the routes in the application. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - Focused server Vitest run — 4 files, 143 tests passed ## Risks - Authorization bugs could permit cross-company or over-broad tool access; the PR remains inert until PR 4 wiring and requires dedicated security review. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c6d4ee10f7 |
test: align current-master regression expectations (#9577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its server and UI test suites protect company-scoped plugin access and instance settings behavior > - Recent governed-access contracts intentionally added company invocation scope and new experimental-setting defaults > - Four existing tests were not updated consistently with those contracts, causing current-master CI failures unrelated to the changes under review > - The runtime behavior is intentional, so changing production code would weaken the new authorization and settings contracts > - This pull request aligns the stale tests with current behavior and removes one UI assertion accidentally pulled forward from a later stacked feature > - The benefit is a focused, low-risk repair that restores master CI without changing application behavior ## Linked Issues or Issue Description - **Bug:** Current master has four regression failures in plugin authorization, plugin execution-workspace bridging, instance settings normalization, and experimental settings UI tests. - **Expected behavior:** Tests provide required company/invocation scope, use the governed object-shaped secret reference contract, include all current defaults, and only assert UI controls implemented at this stack level. - **Actual behavior:** Tests exercised obsolete request shapes or expected a later-stack Apps toggle that is not present on current master. - **Reproduction:** Run the four test files listed in the Verification section on master before this commit. ## What Changed - Updates plugin config authorization coverage to include company scope and an object-shaped `secret_ref` binding. - Supplies invocation company scope to execution-workspace host-client tests. - Adds `enableApps` and `enableSmokeLab` to normalized settings expectations. - Removes the premature Apps toggle UI test introduced without its later-stack implementation. ## Verification - `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts server/src/__tests__/plugin-execution-workspace-bridge.test.ts server/src/__tests__/instance-settings-service.test.ts ui/src/pages/InstanceExperimentalSettings.test.tsx` — 73 tests passed. - `pnpm exec vitest run packages/plugins/sdk/tests/host-client-factory.test.ts server/src/__tests__/plugin-secrets-handler.test.ts server/src/__tests__/instance-settings-routes.test.ts ui/src/lib/instance-settings.test.ts` — 39 tests passed. - `git diff --check` — passed. ## Risks - Low risk: test-only changes with no production runtime, schema, API, or UI behavior changes. - The removed Apps toggle assertion should return in the later stacked change that introduces the actual control. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1de0a3bb1e |
feat(mcp) [split 2/8]: add governed access contracts (#9557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 2/8 and focuses on database schema and shared governance contracts > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The governed access model needs additive persistence and synchronized shared types before server enforcement can compile. - Proposed solution: Adds migrations 0148–0169, tool-access and Smoke Lab schema, shared types/validators/gallery helpers, and the minimal compile-required contract consumers identified by boundary testing. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/01-demo-servers`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for migrations/validators; Greptile on every PR. ## What Changed - Adds migrations 0148–0169, tool-access and Smoke Lab schema, shared types/validators/gallery helpers, and the minimal compile-required contract consumers identified by boundary testing. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` — passed, including migration numbering and safety checks - `pnpm --filter @paperclipai/db test` — passed - `pnpm --filter @paperclipai/shared test` — passed ## Risks - Migration or contract mistakes could affect every upper layer; all migrations are additive/idempotent and compile consumers are included in this boundary. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
072f1e9b8e |
build(deps): bump dompurify from 3.4.8 to 3.4.12 (#9478)
[//]: # (dependabot-start) ⚠️ **Dependabot is rebasing this PR** ⚠️ Rebasing might not happen immediately, so don't worry if this takes some time. Note: if you make any changes to this PR yourself, they will take precedence over the rebase. --- [//]: # (dependabot-end) Bumps [dompurify](https://github.com/cure53/DOMPurify) from 3.4.8 to 3.4.12. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/cure53/DOMPurify/releases">dompurify's releases</a>.</em></p> <blockquote> <h2>DOMPurify 3.4.12</h2> <ul> <li>Fixed an issue where a hook would not get called for custom elements, thanks <a href="https://github.com/Rikuxx0"><code>@Rikuxx0</code></a></li> <li>Hardened the handling of hooks removing elements, <a href="https://github.com/mkrause-bee360"><code>@mkrause-bee360</code></a></li> <li>Added support for a few new SVG attributes, thanks <a href="https://github.com/cbn-falias"><code>@cbn-falias</code></a> & <a href="https://github.com/Develop-KIM"><code>@Develop-KIM</code></a></li> <li>Hardened the handling of declarative partial updates</li> <li>Updated the documentation is several spots, README, wiki, etc.</li> <li>Bumped several dependencies where possible</li> </ul> <h2>DOMPurify 3.4.11</h2> <ul> <li>Fixed an issue with a leaky config for hooks via <code>setConfig</code>, thanks <a href="https://github.com/trace37labs"><code>@trace37labs</code></a></li> <li>Bumped vulnerable development dependencies to arrive at plain 0 with <code>npm audit</code></li> <li>Updated the <code>osv-scanner</code> suppression list as no vulnerable dependencies are left for now</li> <li>Updated up the linting tool-chain and removed now-redundant lint directives</li> <li>Updated the documentation is several spots, README, wiki, etc.</li> <li>Bumped several dependencies where possible</li> </ul> <h2>DOMPurify 3.4.10</h2> <ul> <li>Refactored codebase for clarity: extracted the public type declarations into <code>types.ts</code></li> <li>Decomposed the three largest sanitizer functions into focused helpers</li> <li>Removed duplicated defaults and dead branches, consolidated <code>SAFE_FOR_TEMPLATES</code> scrubbing into single shared path</li> <li>Improved per-node performance by hoisting the mXSS probe regexes and testing <code>textContent</code> before <code>innerHTML</code></li> <li>Added a deterministic micro-benchmark harness (<code>npm run bench</code>) with a <code>--compare</code> mode</li> <li>Reduced CI cost by running the full three-engine browser suite once per PR</li> <li>Refreshed the <code>demos/</code> folder so every demo runs again, and added a SVG-via-<code><img></code> demo</li> <li>Documented the bench and <code>test:happydom</code> scripts in the README</li> <li>Completed the Attack Classes & Bypass History wiki page</li> <li>Bumped several dependencies where possible</li> </ul> <h2>DOMPurify 3.4.9</h2> <ul> <li>Further improved the handling of Trusted Types config options, thanks <a href="https://github.com/offset"><code>@offset</code></a></li> <li>Further improved the handling of <code>IN_PLACE</code> sanitization, thanks <a href="https://github.com/mozfreddyb"><code>@mozfreddyb</code></a></li> <li>Added more test coverage for <code>IN_PLACE</code> and Trusted Types related usage</li> <li>Bumped several dependencies where possible</li> <li>Updated README and wiki with more accurate documentation & attack samples</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/cure53/DOMPurify/commit/a9ca1e537422319a557a9a2aa61f003b23b4a197"><code>a9ca1e5</code></a> release: 3.4.12 (<a href="https://redirect.github.com/cure53/DOMPurify/issues/1537">#1537</a>)</li> <li><a href="https://github.com/cure53/DOMPurify/commit/0cae5187403132f96a6d357649e4b15633fc210a"><code>0cae518</code></a> release: 3.4.11 (<a href="https://redirect.github.com/cure53/DOMPurify/issues/1494">#1494</a>)</li> <li><a href="https://github.com/cure53/DOMPurify/commit/6ee5716f8336989753611beeca364957c0eb0c3e"><code>6ee5716</code></a> release: 3.4.10 (<a href="https://redirect.github.com/cure53/DOMPurify/issues/1478">#1478</a>)</li> <li><a href="https://github.com/cure53/DOMPurify/commit/52102472d46035857c52df19e44285f8a1e102fc"><code>5210247</code></a> release: 3.4.9 (<a href="https://redirect.github.com/cure53/DOMPurify/issues/1459">#1459</a>)</li> <li>See full diff in <a href="https://github.com/cure53/DOMPurify/compare/3.4.8...3.4.12">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f92a6648d5 |
build(deps): bump better-auth from 1.6.20 to 1.6.23 (#9479)
Bumps [better-auth](https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth) from 1.6.20 to 1.6.23. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/better-auth/better-auth/releases">better-auth's releases</a>.</em></p> <blockquote> <h2>v1.6.23</h2> <h2><code>better-auth</code></h2> <h3>Features</h3> <ul> <li>Added Yandex as a social OAuth provider (<a href="https://redirect.github.com/better-auth/better-auth/pull/9138">#9138</a>)</li> </ul> <p>For detailed changes, see <a href="https://github.com/better-auth/better-auth/blob/9dfceee14021fc15a2fb93023f39635f25b0b5ba/packages/better-auth/CHANGELOG.md"><code>CHANGELOG</code></a></p> <h2><code>@better-auth/drizzle-adapter</code></h2> <h3>Bug Fixes</h3> <ul> <li>Fixed affected row counting for D1 and postgres-js adapters (<a href="https://redirect.github.com/better-auth/better-auth/pull/10257">#10257</a>)</li> </ul> <p>For detailed changes, see <a href="https://github.com/better-auth/better-auth/blob/9dfceee14021fc15a2fb93023f39635f25b0b5ba/packages/drizzle-adapter/CHANGELOG.md"><code>CHANGELOG</code></a></p> <h2><code>@better-auth/stripe</code></h2> <h3>Bug Fixes</h3> <ul> <li>Fixed organization subscription actions (cancel, upgrade, restore, and the billing portal) that could act on the wrong organization.</li> </ul> <p>For detailed changes, see <a href="https://github.com/better-auth/better-auth/blob/9dfceee14021fc15a2fb93023f39635f25b0b5ba/packages/stripe/CHANGELOG.md"><code>CHANGELOG</code></a></p> <h2><code>auth</code></h2> <h3>Bug Fixes</h3> <ul> <li>Fixed string default values not being properly escaped in the generated Drizzle schema (<a href="https://redirect.github.com/better-auth/better-auth/pull/10259">#10259</a>)</li> </ul> <p>For detailed changes, see <a href="https://github.com/better-auth/better-auth/blob/9dfceee14021fc15a2fb93023f39635f25b0b5ba/packages/cli/CHANGELOG.md"><code>CHANGELOG</code></a></p> <h2>Contributors</h2> <p>Thanks to everyone who contributed to this release:</p> <p><a href="https://github.com/bytaesu"><code>@bytaesu</code></a>, <a href="https://github.com/vladflotsky"><code>@vladflotsky</code></a></p> <p><strong>Full changelog:</strong> <a href="https://github.com/better-auth/better-auth/compare/v1.6.22...v1.6.23"><code>v1.6.22...v1.6.23</code></a></p> <h2>v1.6.22</h2> <h2><code>better-auth</code></h2> <h3>Bug Fixes</h3> <ul> <li>Fixed unproven credentials not being revoked during magic link and email OTP sign-in (<a href="https://redirect.github.com/better-auth/better-auth/pull/10239">#10239</a>)</li> <li>Fixed server-side OAuth requests to refuse redirect responses instead of following them (<a href="https://redirect.github.com/better-auth/better-auth/pull/10241">#10241</a>)</li> </ul> <p>For detailed changes, see <a href="https://github.com/better-auth/better-auth/blob/a90d061de7cdbd60e796230aadf5d1082add1fe2/packages/better-auth/CHANGELOG.md"><code>CHANGELOG</code></a></p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/better-auth/better-auth/blob/main/packages/better-auth/CHANGELOG.md">better-auth's changelog</a>.</em></p> <blockquote> <h2>1.6.23</h2> <h3>Patch Changes</h3> <ul> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/9138">#9138</a> <a href="https://github.com/better-auth/better-auth/commit/8581f97ea0000e03edd6aa7911efabf694a9ff95"><code>8581f97</code></a> Thanks <a href="https://github.com/vladflotsky"><code>@vladflotsky</code></a>! - Add a pre-configured Yandex provider helper for the generic OAuth plugin.</p> </li> <li> <p>Updated dependencies [<a href="https://github.com/better-auth/better-auth/commit/930b260cfd402e9f8886719a3ced503b9ceff7f6"><code>930b260</code></a>]:</p> <ul> <li><code>@better-auth/drizzle-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/core</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/kysely-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/memory-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/mongo-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/prisma-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> <li><code>@better-auth/telemetry</code><a href="https://github.com/1"><code>@1</code></a>.6.23</li> </ul> </li> </ul> <h2>1.6.22</h2> <h3>Patch Changes</h3> <ul> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/10239">#10239</a> <a href="https://github.com/better-auth/better-auth/commit/c06a56d83a40bbaeac12d3a8b8b67e59f92a9110"><code>c06a56d</code></a> Thanks <a href="https://github.com/gustavovalverde"><code>@gustavovalverde</code></a>! - Magic-link and email-OTP sign-in now reset the credentials on an account whose email had never been confirmed. When verification resolves to such an account, any existing password on it is removed and its sessions are revoked before the user is signed in, so proven control of the mailbox is the source of truth for the account.</p> <p>If you signed up with email and password but first signed in through a magic link or email OTP rather than confirming the verification email, your password is cleared and you will need to set a new one through password reset.</p> </li> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/10240">#10240</a> <a href="https://github.com/better-auth/better-auth/commit/3a035e968e27bfdee1e53ad857e5569090d9f2d1"><code>3a035e9</code></a> Thanks <a href="https://github.com/gustavovalverde"><code>@gustavovalverde</code></a>! - Add account-level lockout for two-factor verification. The attempt limit applies per account across sign-in challenges and across factors: TOTP, email-OTP, and backup codes share one counter, and a successful verification resets it.</p> <p>Enabled by default: an account locks for 15 minutes after 10 consecutive failed verifications, and locked attempts return <code>429</code> with the <code>ACCOUNT_TEMPORARILY_LOCKED</code> error code. Configure it with <code>twoFactor({ accountLockout: { enabled, maxFailedAttempts, durationSeconds } })</code>.</p> <p>Run a database migration after upgrading: this adds <code>failedVerificationCount</code> and <code>lockedUntil</code> columns to the <code>twoFactor</code> table.</p> </li> <li> <p>Updated dependencies [<a href="https://github.com/better-auth/better-auth/commit/8bd43d9d8312fd9ddbfb8fb5c827cf0a0e55132d"><code>8bd43d9</code></a>]:</p> <ul> <li><code>@better-auth/core</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/drizzle-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/kysely-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/memory-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/mongo-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/prisma-adapter</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> <li><code>@better-auth/telemetry</code><a href="https://github.com/1"><code>@1</code></a>.6.22</li> </ul> </li> </ul> <h2>1.6.21</h2> <h3>Patch Changes</h3> <ul> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/10212">#10212</a> <a href="https://github.com/better-auth/better-auth/commit/e0762a127ce351a96614e60866b3455e6eddffa1"><code>e0762a1</code></a> Thanks <a href="https://github.com/bytaesu"><code>@bytaesu</code></a>! - In root-mounted deployments, requests whose path does not start with the configured <code>basePath</code> now return 404 instead of resolving to an endpoint.</p> </li> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/10187">#10187</a> <a href="https://github.com/better-auth/better-auth/commit/882cf9e592d1d305b5b78cadbb10aaeee7acd6dc"><code>882cf9e</code></a> Thanks <a href="https://github.com/ping-maxwell"><code>@ping-maxwell</code></a>! - Admin permission changes and bans now take effect immediately for admin APIs, even when session cookie cache is enabled. Sensitive session checks also continue to work in stateless apps where signed cookies are the session record.</p> </li> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/9939">#9939</a> <a href="https://github.com/better-auth/better-auth/commit/f52e1ab50b60d289b64d6b06f1bff5a4358cdfd0"><code>f52e1ab</code></a> Thanks <a href="https://github.com/benpsnyder"><code>@benpsnyder</code></a>! - fixes a bug causing deviceAuthorization() throwing a ZodError at construction when called without a schema option</p> </li> <li> <p><a href="https://redirect.github.com/better-auth/better-auth/pull/10196">#10196</a> <a href="https://github.com/better-auth/better-auth/commit/b5bec193a56cec2f7b71c84d71dacb632f0b96a0"><code>b5bec19</code></a> Thanks <a href="https://github.com/Paola3stefania"><code>@Paola3stefania</code></a>! - OAuth sign-up and account-link profile sync now ignore provider profile values for user fields marked <code>input: false</code>. Input-allowed additional fields still persist from <code>mapProfileToUser</code>, and schema defaults still apply when OAuth creates a user. Apps that used <code>mapProfileToUser</code> to fill <code>input: false</code> fields should set those fields in server-side provisioning code instead.</p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/better-auth/better-auth/commit/9dfceee14021fc15a2fb93023f39635f25b0b5ba"><code>9dfceee</code></a> chore: release v1.6.23 (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10260">#10260</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/8581f97ea0000e03edd6aa7911efabf694a9ff95"><code>8581f97</code></a> feat(oauth): add Yandex social provider (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/9138">#9138</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/a90d061de7cdbd60e796230aadf5d1082add1fe2"><code>a90d061</code></a> chore: release v1.6.22 (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10245">#10245</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/3a035e968e27bfdee1e53ad857e5569090d9f2d1"><code>3a035e9</code></a> fix(two-factor): add account-level verification lockout (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10240">#10240</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/c06a56d83a40bbaeac12d3a8b8b67e59f92a9110"><code>c06a56d</code></a> fix: revoke unproven credentials on magic-link/email-OTP sign-in (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10239">#10239</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/414169d95a88a6e1fac41688bc7011e96feb0d2a"><code>414169d</code></a> chore: release v1.6.21 (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10184">#10184</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/f52e1ab50b60d289b64d6b06f1bff5a4358cdfd0"><code>f52e1ab</code></a> fix(device-authorization): make <code>schema</code> option optional under Zod v4 (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/9939">#9939</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/882cf9e592d1d305b5b78cadbb10aaeee7acd6dc"><code>882cf9e</code></a> fix(admin): use authoritative session reads for authorization (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10187">#10187</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/b5bec193a56cec2f7b71c84d71dacb632f0b96a0"><code>b5bec19</code></a> fix(oauth): apply user input rules to provider profiles (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10196">#10196</a>)</li> <li><a href="https://github.com/better-auth/better-auth/commit/471f81c1ab088368de07383491535c1288fa6543"><code>471f81c</code></a> refactor: centralize request IP resolver in core (<a href="https://github.com/better-auth/better-auth/tree/HEAD/packages/better-auth/issues/10216">#10216</a>)</li> <li>Additional commits viewable in <a href="https://github.com/better-auth/better-auth/commits/v1.6.23/packages/better-auth">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c4eb5339e7 |
fix(worktree): require target attestation before repair (#9414)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work > - Local agent execution uses isolated git worktrees with worktree-specific config, environment, storage, and ports > - Legacy worktree repair and runtime-port persistence must mutate only the worktree they are serving > - A leaked ambient `PAPERCLIP_IN_WORKTREE=true` could be combined with config resolution pointing at the default instance > - The server test suite reproduced that combination and repeatedly rewrote the live default instance `.env` with its old fixture name > - Existing PR #3071 guards configs under Paperclip home, but does not require the target itself to attest worktree ownership and does not cover runtime-port persistence > - This pull request requires both a worktree config layout and target-local persisted worktree attestation before either writer adopts the target > - The benefit is that ambient process state can never turn the main instance into a worktree on its next restart ## Linked Issues or Issue Description Related implementation: Refs #3071 **Pre-submission checklist** - [x] Searched open and closed issues and pull requests; #3071 is the only direct related implementation. - [x] Reproduced on current `master` before applying the fix. - [x] Confirmed the mutation originates in Paperclip's worktree config repair path. **What happened?** A process with leaked `PAPERCLIP_IN_WORKTREE=true` could resolve `PAPERCLIP_CONFIG` to the default instance and cause worktree repair to rewrite that instance's `.env`. The recurring trigger was `server/src/__tests__/worktree-config.test.ts`: an ambient config path from the developer shell survived into a test whose fixture worktree name was `PAP-884-ai-commits-component`, explaining the stale name repeatedly written to the live file. **Expected behavior** Worktree repair and worktree runtime-port persistence must mutate a target only when that target is independently provisioned and persisted as a worktree. Ambient environment flags alone must never authorize writes to the default instance or a normal repository-local `.paperclip` config. **Steps to reproduce on unpatched `master`** 1. Export `PAPERCLIP_CONFIG` pointing to a default instance config and set `PAPERCLIP_IN_WORKTREE=true`. 2. Run `server/src/__tests__/worktree-config.test.ts` from that shell. 3. Observe that the default instance `.env` is rewritten with the test fixture's worktree marker and name. **Environment** - Version: `master` at `e4e12bfb8` - Deployment/install: local source checkout with pnpm - Adapter: not adapter-specific; core server config - Database/access context: not applicable - OS: Linux **Privacy** - [x] All paths and values in this description are generic and contain no credentials or personally identifying data. ## What Changed - Reject config targets unless their parent directory is the worktree-specific `.paperclip` layout. - Require the target's own persisted `.env` to declare `PAPERCLIP_IN_WORKTREE=true` before repair or runtime-port persistence can mutate it. - Scrub ambient `PAPERCLIP_*` variables before every worktree-config test so developer-machine exports cannot escape test isolation. - Add regressions for default-instance config poisoning, runtime-port persistence, and unattested repository-local `.paperclip` targets. - Preserve valid provisioned worktree behavior by adding persisted worktree markers to the existing positive fixtures. ## Verification - `NODE_ENV=test pnpm --filter @paperclipai/server exec vitest run src/__tests__/worktree-config.test.ts` — 12 tests passed. - Branch is based directly on current `origin/master`; only two server files changed. - No `pnpm-lock.yaml`, workflow, migration, UI, or generated asset changes. ## Risks - Low risk: the new guard intentionally refuses repair for targets that lack provisioning evidence. - A manually assembled worktree that sets only ambient flags but never writes its worktree marker will no longer be auto-repaired; the supported provisioning path already writes that marker. - No schema, API, migration, or user-facing command changes. > This is a focused correctness fix and does not overlap with planned core work in `ROADMAP.md`. ## Model Used - Implementation and root-cause investigation: Anthropic Claude through the `claude_local`/Claude Code runtime, reported by the producing agent as “Claude Fable 5”; the runtime did not expose a more specific provider model ID or context-window value. Capabilities used: extended reasoning, shell tool use, code editing, and test execution. - PR preparation and verification: OpenAI Codex CLI runtime; the harness did not expose the exact underlying model ID or context-window value. Capabilities used: repository inspection, shell tool use, Git/GitHub operations, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with all model details exposed by the runtimes - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked #3071 above - [x] I have described the issue in-PR following the bug report template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket ID - [x] I have run the focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have assessed documentation impact; no documentation change is required for this internal guard - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f204605ad |
fix(ui): show source SHA for unreleased builds (#9508)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need to identify the exact build running from the persistent account menu > - Formal releases already have a concise public version, but source builds include a long derived version string > - The derived version identifies a commit but does not expose the source branch or a direct path to inspect the code > - Server Git metadata is auth-sensitive, so the UI must also refresh it when the current session changes > - This pull request shows linked branch and commit metadata for source builds while preserving `v<version>` for formal releases > - The benefit is faster build diagnosis with correct metadata across sign-in and sign-out transitions ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I searched existing open and closed issues and found no duplicate implementing this exact account-menu behavior. - [x] The behavior reproduces on `master`. - [x] The behavior originates in Paperclip's core UI, not an adapter, provider, or local configuration. ### What happened? Source builds displayed the full derived version, such as `2026.626.0+58.git.518fc71ce`, without linking the operator to the corresponding source branch or commit. ### Expected behavior Source builds should show the concise branch and short commit SHA with links to GitHub, while formal releases should continue showing their public version. Auth transitions should refresh the health metadata that supplies those Git details. ### Steps to reproduce 1. Run Paperclip from a commit after a release tag. 2. Open the account menu. 3. Inspect the build label beneath the user identity. 4. Sign in or out and reopen the menu. ### Paperclip version or commit Any source build whose server version uses the `<version>+<count>.git.<sha>[.dirty]` format. ### Deployment mode Local dev (`pnpm dev`) or authenticated deployments. ### Installation method Built from source. ## What Changed - Detect source-derived version strings and render the source branch plus seven-character commit SHA in `SidebarAccountMenu`. - Link source branches and commits to the canonical `paperclipai/paperclip` GitHub repository. - Extend server Git metadata with the full SHA and expose it through health/OpenAPI contracts. - Refresh auth-sensitive health metadata after sign-in and every sign-out entry point. - Preserve the existing `v<version>` label for formal releases and add focused regression coverage. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Auth.test.tsx src/components/SidebarAccountMenu.test.tsx src/components/SidebarServerInfo.test.tsx` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/health.test.ts src/__tests__/server-info.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `git diff --check public/master...HEAD` ## Risks - Low risk: formal release rendering retains the existing fallback behavior when the source-version pattern does not match. - Source links assume the build came from the canonical public repository; fork-only branches or commits may not resolve there. - Health metadata is invalidated after auth transitions, adding one bounded refetch so the displayed Git details match the new session. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4 with medium reasoning, repository/tool access, shell execution, and code editing; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |
||
|
|
0e21a27301 |
fix: forward onSpawn to hermes and process adapters for PID persistence (#8722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter layer (hermes-local, process adapters) delegates agent execution to child processes via `runChildProcess()` > - `runChildProcess()` accepts an `onSpawn` callback to report child PID and process group info, but the hermes and process adapters were not forwarding `ctx.onSpawn` to this call > - Without PID persistence, the orphan reaper cannot distinguish live runs from abandoned processes, causing false-positive reaps and 5-minute timeout errors for active runs > - This pull request adds `onSpawn: ctx.onSpawn` to both adapter call sites and declares the option in the `runChildProcess` wrapper type > - The benefit is that the orphan reaper can now correctly track live child processes, eliminating false-positive reaps ## Linked Issues or Issue Description Fixes #8723 Fixes false-positive orphan reaps in hermes-local and process adapters by forwarding the `onSpawn` callback to `runChildProcess()`. All other adapters (claude-local, codex-local, cursor-local, gemini-local, grok-local, opencode-local, pi-local) already forward `ctx.onSpawn` — these two were the only ones missing it. ## What Changed - `server/src/adapters/utils.ts`: Added `onSpawn?` to the `runChildProcess()` options type so callers can forward the callback - `server/src/adapters/process/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` - `packages/adapters/hermes/src/server/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` ## Verification - `pnpm -r typecheck` passes across all packages - Confirmed all other adapters already forward `ctx.onSpawn` (12 grep matches across 9 adapter files) - The 3-line diff is additive only — no existing behavior is changed, only a previously-ignored callback is now forwarded ## Risks Low risk. This is a 3-line additive change. The `onSpawn` parameter is optional (`?`) so existing callers are unaffected. The callback is already well-established across all other adapters. ## Model Used Hermes Agent (by Nous Research) — xiaomi/mimo-v2.5-pro via OpenRouter, with tool use (file editing, git, GitHub API). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass (typecheck passes) - [x] I have added or updated tests where applicable (N/A — type-level fix only, no behavioral change) - [x] I have updated relevant documentation to reflect my changes (N/A — internal fix) - [x] I have considered and documented any risks above --------- Co-authored-by: Zephyr <zephyr@motoyuki.dev> |
||
|
|
c8253e3641 |
fix(adapters): inject execution contract once per fresh heartbeat (#9469)
## Thinking Path > - Paperclip coordinates AI-agent work through repeated heartbeat runs. > - Adapter prompts combine a default heartbeat template with scoped wake context. > - Fresh heartbeats received the same execution contract from both layers, wasting prompt tokens and obscuring which layer owns the contract. > - Resume deltas and template-less adapters do not share that composition path, so removing the wake-payload copy unconditionally would drop required guidance. > - Empty comment batches also emitted instructions and metadata that only matter when comments exist. > - This pull request makes execution-contract inclusion explicit by prompt path, preserves OpenClaw gateway behavior, and suppresses no-op comment boilerplate. > - The benefit is one contract per heartbeat path and roughly 300 fewer prompt tokens on a fresh zero-comment wake. ## Linked Issues or Issue Description - Fixes #9221 - Refs #9200 - Refs #7634 ## What Changed - Stop emitting the execution-contract paragraph from fresh scoped wake payloads because the default heartbeat template already contains the full contract. - Keep the contract in resume deltas, and add `includeExecutionContract` for adapters that do not render the default heartbeat template. - Opt `openclaw-gateway` into wake-payload contract rendering so template-less gateway runs retain the guidance. - Omit comment-batch acknowledgement/fetch guidance and empty `pending comments` / `latest comment id` metadata when a fresh wake has no pending comments. - Add regression and acceptance coverage proving composed fresh prompts contain `Execution contract` exactly once while resume and template-less paths retain it. Measured effect: the fresh zero-comment wake block drops from 1,840 to 855 characters (about 300 tokens saved per fresh heartbeat; about 220 on comment wakes), and the composed fresh prompt contains `Execution contract` once instead of twice. ## Verification - `npx vitest run packages/adapter-utils/src/server-utils.test.ts` — 63 passed - `npx vitest run server/src/__tests__/codex-local-execute.test.ts` — 13 passed - `npx vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/openclaw-gateway-adapter.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts` — 27 passed - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` — passed ## Risks - Low risk: prompt text and adapter composition only; no database or API migration. - The main compatibility risk is a template-less adapter losing the contract. The explicit option and OpenClaw gateway regression coverage protect the known template-less path. - External adapters that call `renderPaperclipWakePrompt` directly can opt into `includeExecutionContract: true` when they do not render the default template. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, reasoning mode with tool use and code execution; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5dff52631d |
fix(heartbeat): throttle redundant issue re-wakes (#9470)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work > - Heartbeat admission decides when an agent should start another adapter session for an issue > - After process-loss recovery, assignment pollers and reconcilers can repeatedly request another wake while the issue remains `in_progress` > - When the preceding runs succeeded without issue-visible progress, those event-free wakes provide no new information but still pay the full cost of an adapter session > - Existing liveness evidence is too broad for this case because workspace tool calls can make a run look active without moving the issue > - This pull request adds an issue-scoped admission throttle for consecutive no-progress re-wakes while preserving every wake that carries new information or recovery intent > - The benefit is bounded recovery cost without delaying comments, operator actions, failures, or other meaningful events ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** After a process died, external wake drivers could re-wake the same agent for the same `in_progress` issue every few seconds. Each succeeded run that produced no issue-visible progress could be followed by another full adapter session despite no new issue input. In the observed recovery smoke, one recovery consumed 25 sessions and 2.4× the direct-run cost. **Expected behavior** Repeated event-free re-wakes should back off after consecutive successful runs produce no issue-visible progress. Any new information, explicit operator intent, or failed-run recovery should continue immediately. **Steps to reproduce** 1. Start an issue heartbeat and simulate process loss while the issue remains `in_progress`. 2. Allow assignment/reconciliation drivers to request repeated event-free wakes for the same agent and issue. 3. Complete each follow-up run successfully without adding a comment, issue mutation, document, work product, interaction, or continuation. 4. Observe repeated adapter sessions starting every few seconds without new issue input. **Environment** - Version: reproduced on `master` before this change - Deployment: local development, built from source - Adapter scope: core bug; not adapter-specific - Database: reproduced and tested with embedded Postgres ## What Changed - Add a pure issue re-wake throttle that detects consecutive succeeded runs without issue-visible progress and applies a 120-second exponential cooldown capped at 30 minutes. - Gate event-free `enqueueWakeup` requests and return the explicit skip reason `issue_rewake_throttled` while the cooldown is active. - Always bypass throttling for comment wakes, new issue activity, explicit resumes, `forceFreshSession`, event-shaped reasons, and post-failure recovery. - Add focused pure unit coverage and database-backed heartbeat admission coverage for throttle and bypass behavior. ## Verification - `cd server && pnpm vitest run src/__tests__/issue-rewake-throttle.test.ts` — 12 passed. - `cd server && pnpm vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6 passed with embedded Postgres. - `cd server && pnpm run typecheck` — passed. - Neighbor suites previously verified: `heartbeat-dependency-scheduling`, `heartbeat-process-recovery`, `run-continuations`, `heartbeat-issue-liveness-escalation`, `recovery-stale-issue-lock-sweep`, and `heartbeat-comment-wake-batching` — 131 tests passed. ## Risks - A progress classifier that is too narrow could defer a legitimate event-free poll; the cooldown is bounded and new issue activity bypasses it immediately. - A progress classifier that is too broad could allow the original heartbeat storm; tests intentionally distinguish issue-visible mutations from workspace-only activity. - Low compatibility risk: no schema, API contract, or migration changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The runtime does not expose the exact underlying model ID or context-window size; reasoning, terminal tool use, code inspection, GitHub CLI access, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e7e84e3fe |
fix(adapters): propagate ACP-lane usage and cost into spend telemetry (#9471)
## Thinking Path > - Paperclip is the open source control plane for running and governing AI-agent companies. > - Adapter executions feed token usage, billing identity, and run cost into the control plane's spend telemetry. > - The default ACP execution lane for local Claude and Codex adapters did not propagate per-turn usage or cost, so paid runs could be recorded with zero spend and no tokens. > - Claude CLI result events could also undercount output tokens by reading only the main-loop usage block instead of the complete per-model ledger. > - The shared executor needs to distinguish per-run usage from session-cumulative usage so the server does not apply the wrong delta heuristic. > - This pull request captures ACP usage and cumulative-cost deltas, resolves adapter billing identity, uses Claude's complete model-usage ledger, and preserves per-run usage in server normalization. > - The benefit is accurate token and cost accounting across the default paid Claude and Codex execution paths. ## Linked Issues or Issue Description ### What happened? Paid `claude_local` and `codex_local` runs using the default ACP engine can complete successfully while the control plane records zero or null cost and missing token usage. Claude CLI result parsing can additionally undercount output tokens when subagent or sidechain usage is present. ### Steps to reproduce 1. Run a paid Claude or Codex local adapter through the ACP engine. 2. Complete a turn that reports usage and cumulative cost through ACP status/events. 3. Inspect the execution result and normalized run telemetry. ### Expected behavior The execution result contains per-turn token usage, a per-run USD cost delta, and the correct billing identity. Server normalization records those per-run values without applying a session-cumulative delta a second time. ### Actual behavior before this change ACP execution results returned no usage and `costUsd: null` with unknown billing. The server therefore recorded zero spend and no tokens for paid runs. Claude CLI parsing could use an incomplete usage block. ## What Changed - Capture ACP usage from runtime status and `usage_update` events, reporting it as `usageBasis: per_run`. - Convert agent-reported cumulative ACP cost into a per-turn delta, including counter-reset and no-report safeguards. - Add a shared billing-identity resolver and map Claude and Codex authentication/provider modes to control-plane billing types. - Prefer Claude result-event `modelUsage` totals so subagent and sidechain tokens are included. - Skip the server's session-cumulative usage delta when an adapter explicitly reports per-run usage. - Add regression coverage for usage capture, event fallback, cost resets, stale reports, billing identities, model-usage totals, and server spend normalization. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/acp.test.ts server/src/__tests__/costs-service.test.ts server/src/__tests__/monthly-spend-service.test.ts` — 6 files, 126 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `pnpm --filter @paperclipai/adapter-claude-local typecheck` — passed. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - A broader Claude-local suite has a pre-existing rate-limit classification failure in `test.probe.test.ts`; it also fails on clean `master` and is unrelated to this change. ## Risks - Cost reporting depends on the agent's cumulative counter semantics; reset handling falls back to the post-turn amount and is covered by regression tests. - Incorrect billing-mode inference could misclassify spend; provider/auth mappings mirror each adapter's existing CLI behavior and have focused tests. - The new `usageBasis` contract changes server normalization only when adapters explicitly opt into `per_run`; existing adapters retain prior behavior. - No database migration, workflow, lockfile, or UI changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation commit: Anthropic Claude Fable 5, tool-enabled coding workflow (exact context window and runtime configuration were not recorded in the commit metadata). - PR preparation and verification: OpenAI Codex, tool-enabled coding agent (runtime model ID and context window are not exposed to this session). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4a40c0cb13 |
feat(routines): gate scheduled runs on external activity (#9436)
## Thinking Path > - Paperclip is the open source app people use to manage AI-agent companies and their recurring work > - Scheduled routines provide native cron-driven execution for recurring agent tasks > - Watcher-style routines currently dispatch a model run even when the control plane has been quiet since their last useful run > - Existing pause, catch-up, and concurrency policies do not distinguish external work from a routine's own bookkeeping > - This pull request adds a generic activity gate that checks company-scoped activity provenance before scheduled dispatch > - The benefit is backward-compatible zero-token quiet skips while real human, agent, or delegated-child activity still wakes the routine ## Linked Issues or Issue Description - Refs #8534 ## What Changed - Added `activity_gate_policy` and `activity_gate_scope` routine columns with backward-compatible `always` / `company` defaults. - Added a company-bounded `evaluateActivityGate()` predicate that uses the last dispatched run as its open window, excludes the routine's own execution runs and scheduler bookkeeping, ignores pure-read actions, and supports company/project scope. - Integrated the predicate into scheduled ticks after pause/worktree eligibility checks; quiet ticks create visible skipped run-history rows with reason `no_external_activity` and gate-window diagnostics without advancing the activity window. - Kept webhook, manual, and API dispatch paths ungated; catch-up schedules evaluate the gate once per scheduler tick. - Added migration-default, provenance predicate, project-scope, quiet-window, scheduler, and webhook-bypass coverage. ## Verification - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts` — 51 tests passed - Embedded Postgres `EXPLAIN` for the company-scope gate scan: ```text Limit (cost=24.56..24.58 rows=1 width=24) -> Incremental Sort (cost=24.56..24.60 rows=2 width=24) Sort Key: activity.created_at, activity.id Presorted Key: activity.created_at -> Nested Loop Anti Join (cost=0.44..24.55 rows=1 width=24) Join Filter: (own_run.id = activity.run_id) -> Index Scan using activity_log_company_created_idx on activity_log activity (cost=0.15..8.19 rows=1 width=40) Index Cond: ((company_id = '00000000-0000-0000-0000-000000000001'::uuid) AND (created_at > (now() - '01:00:00'::interval)) AND (created_at <= now())) ``` ## Risks - The migration adds two non-null text columns, but constant defaults preserve all existing routine behavior and avoid a backfill step. - Project scope resolves activity through issue/run/routine provenance; tests cover in-project and cross-project issue activity, while every top-level and correlated query remains company-bounded. - This is the scheduler/schema foundation. Public API validation and documentation for configuring the new fields are intentionally handled in the next scoped follow-up. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.4` with medium reasoning, repository/tool access, terminal code execution, and test execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR extends the existing Scheduled Routines roadmap item - [x] I have searched GitHub for duplicate or related PRs and linked the related efficiency request above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing configuration is exposed in this scoped foundation PR) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4e12bfb89 |
fix(workspaces): persist readiness state and validate ports (#9408)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces can run managed services that must report reliable lifecycle and readiness state > - Service startup previously waited for readiness before committing the starting row, making concurrent control actions see stale state > - Fixed service ports also needed clearer configuration and ownership diagnostics to avoid cross-workspace collisions > - This pull request persists startup state before readiness, validates port ownership, and exposes configurable service ports in the workspace UI > - The benefit is dependable service controls and actionable diagnostics when workspace runtimes start slowly or compete for ports ## Linked Issues or Issue Description ### What happened? Slow-starting workspace services could remain invisible to concurrent stop/restart controls until readiness completed, and fixed-port conflicts lacked enough ownership context for safe repair. ### Expected behavior A starting service is persisted immediately, control operations can observe it, configured ports are editable, and conflicts identify the owning process/workspace. ### Steps to reproduce 1. Configure a workspace service that delays binding its HTTP port. 2. Start the service and immediately request another control action. 3. Observe stale persisted state before this change. 4. Configure two workspaces for the same fixed port and observe limited conflict diagnostics. ### Paperclip version or commit `origin/master` at `02e2dd271` ### Deployment mode Local dev; built from source; not adapter-specific; database-backed workspace runtime state. ## What Changed - Commit the `starting` runtime-service row before waiting for readiness and transition it after the probe completes. - Add port-owner inspection and cross-workspace conflict details to local service supervision. - Preserve configurable runtime service ports through workspace configuration updates. - Surface service-port editing and validation in the execution workspace details UI. - Add server and UI regression coverage for slow readiness, concurrent controls, port persistence, and conflict diagnostics. ## Verification - `vitest --project @paperclipai/server src/__tests__/workspace-runtime.test.ts src/__tests__/execution-workspaces-service.test.ts` — 118 tests passed. - `vitest --project @paperclipai/ui src/pages/ExecutionWorkspaceDetail.service-ports.test.ts` — 4 tests passed. - `node scripts/check-token-gates.mjs` — all token gates clean. ## Risks - Moderate risk: changes touch workspace service lifecycle persistence and local process/port inspection. - No schema migration is required; tests exercise slow readiness, concurrent control, persisted ports, and cross-workspace conflicts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e0f1905222 |
[codex] Quiet packaged version fallback diagnostics (#9207)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server startup path reports the product version from package metadata and, in source checkouts, Git metadata. > - Packaged installs can run from `node_modules`, where Git metadata is normally unavailable and that absence is expected. > - The fallback path was still attempting Git metadata probing in packaged contexts, which could print scary diagnostic noise during onboarding even though the package version fallback was working. > - This pull request makes the packaged path skip Git probing only when the package does not look like a source checkout, and keeps fallback diagnostics opt-in. > - The benefit is a quieter first-run experience without weakening source-checkout version detection or debug diagnostics. ## Linked Issues or Issue Description No public GitHub issue exists. ### Bug Report #### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in an agent adapter, API provider, or local configuration. #### What happened? When Paperclip starts from a packaged install, server version resolution can fall back from Git metadata to package metadata. That expected fallback path could emit scary Git diagnostic noise during onboarding even though startup could continue normally. #### Expected behavior Packaged Paperclip startup should use package metadata quietly when Git metadata is unavailable. Source checkouts should still use Git-derived versions, and operators who explicitly opt into version-resolution diagnostics should still receive useful Git failure details. #### Steps to reproduce 1. Run Paperclip from a packaged install where the server package is under `node_modules` and does not include package-local Git metadata. 2. Start the server in an environment where `git describe` cannot resolve repository metadata for that package. 3. Observe that version fallback can produce Git diagnostic noise during startup even though the package version fallback is expected. #### Paperclip version or commit Reproduced against the pre-fix server version resolution behavior on `master`-derived builds. #### Deployment mode Self-hosted server / packaged local install. #### Installation method npm / pnpm package install. #### Agent adapter(s) involved Not adapter-specific; this is core server startup/version behavior. #### Database mode Not database-related. #### Access context Unclear / not applicable. #### Relevant logs or output Git fallback diagnostics from `git describe` could appear during packaged startup. The exact path and Git output depend on the operator environment. #### Additional context The fix keeps diagnostics available behind `PAPERCLIP_DEBUG_VERSION_RESOLUTION=1` and preserves source-checkout Git version detection, including source paths that happen to contain a `node_modules` segment. #### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Skip Git metadata probing for packaged installs under `node_modules` only when no package-local Git metadata is present. - Preserve Git-derived version detection for source or linked workspace checkouts, even when their path contains a `node_modules` segment. - Keep fallback diagnostics behind the existing debug/diagnostic opt-in path. - Include useful Git failure details such as stderr/stdout/stack/cause when diagnostics are enabled. - Add version tests covering packaged fallback behavior, source-checkout detection, richer diagnostics, and quiet default output. ## Verification - `pnpm vitest run server/src/__tests__/version.test.ts` passed after the Greptile follow-up changes. - `pnpm --filter @paperclipai/server typecheck` passed. - `git diff --check` passed. - Greptile completed with confidence score 5/5 and no blocking issues on the latest reviewed commit. ## Risks Low risk. The change is scoped to version fallback behavior. Source-checkout Git version detection remains covered, while packaged `node_modules` contexts intentionally rely on package metadata instead of Git probing. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent with shell/tool use in the Paperclip workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local> |
||
|
|
36ec79c196 |
feat: add attention queue and Decisions surface (#9380)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies, where operators need a reliable way to find and act on work awaiting their input. > - The attention and issue-thread interaction subsystems expose those decision points across server APIs and the board UI. > - The previous navigation and interaction presentation left these actions fragmented and did not offer a controlled rollout for the Decisions surface. > - This branch adds the attention feed, richer interaction cards, grouping, dismiss/snooze behavior, and a gated Decisions sidebar entry. > - It also keeps experimental settings and API contracts synchronized, with an idempotent migration for the new dismissal state. > - This pull request delivers the complete, tested attention/Decisions experience as one reviewable unit. ## Linked Issues or Issue Description - Adds an operator-focused attention queue and Decisions experience: grouped decision cards, semantic interaction actions, dismiss/snooze handling, resilient interaction states, and an experimental flag to control the Decisions navigation entry. ## Feature Context ### Problem or Motivation Operators currently have to hunt across approvals, interactions, failed runs, and budget alerts to find decisions that need their action. ### Proposed Solution Provide a gated Decisions attention queue that groups actionable items, supports direct resolution, and preserves operator control through dismiss and snooze actions. ### Alternatives Considered Keep separate, source-specific views only; this leaves cross-cutting operator decisions fragmented and harder to prioritize. ### Roadmap Alignment This improves the V1 control-plane operator workflow by making pending governed actions discoverable in one company-scoped surface. ## What Changed - Added server attention-feed services, routes, interaction handling, dismiss/snooze support, and an idempotent `0145` inbox-dismissal migration. - Added shared attention, inbox-dismissal, and experimental-settings contracts. - Added Decisions/attention UI, interaction-card states, sidebar badge/navigation integration, grouping, keyboard support, and Storybook coverage. - Added tests for attention behavior, thread interactions, settings normalization, dismissals, and API behavior. - Removed generated screenshots from the final PR diff and rebased the branch onto current `master`. ## Verification - `pnpm check:token-gates` — passed. - `pnpm exec vitest run packages/shared/src/issue-thread-interactions.test.ts server/src/__tests__/attention-service.test.ts server/src/__tests__/inbox-dismissals.test.ts server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts ui/src/lib/attention.test.ts ui/src/components/AttentionQueueRow.test.tsx ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx` — passed: 158 tests across 9 focused files. - GitHub Actions for `ad636f560`: build and typecheck/release-registry have passed; remaining general-server and Greptile checks are in progress. ## Risks - Moderate: this is a cross-layer attention/interaction feature with a new migration and navigation behavior. - The `enableDecisions` experimental setting defaults to off, limiting rollout impact. - Existing dismissal data is backfilled to `dismiss`; the migration is idempotent and uses guarded constraint creation. > ROADMAP.md was checked; no duplicate planned core feature was identified. Related open pull requests were searched before opening this PR. ## Model Used - OpenAI GPT-5.5 via Codex CLI, with tool use and local code execution. Context-window size unavailable in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; focused tests pass and the remaining unrelated AWS test failure is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |