mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
ae7790861811323de7ffcedffa9f56cd88b98bc0
297
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ae77908618 |
feat(search): add bulk extract endpoint (#9507)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents and operators need company-scoped search to discover relevant issue history safely > - The interactive search endpoint intentionally returns compact excerpts and low pagination caps for UI use > - Automation that inventories repeated references, such as pull-request URLs, needs exhaustive distinct matches without loading full issue objects into an LLM context > - Client-provided regular expressions would create an unsafe and expensive query surface, so extraction must remain literal with server-owned expansion modes > - This pull request adds a bounded agent-oriented extraction endpoint with explicit truncation > - The benefit is deterministic, compact bulk discovery across issues, comments, and documents while preserving company authorization and rate limits ## Linked Issues or Issue Description ### Subsystem affected `server/` REST API and `packages/shared/` contracts. ### Problem or motivation The existing interactive company search caps issue pagination and snippets, so automation cannot reliably enumerate every distinct literal or pull-request URL across issue descriptions, comments, and linked documents without fetching large full issue payloads. ### Proposed solution Add `GET /api/companies/:companyId/search/extract` with escaped literal matching, optional server-owned URL token expansion, issue/comment/document scopes, status/date filters, higher issue-level pagination caps, compact source references, and explicit pagination/match truncation flags. ### Alternatives considered Reusing `GET /issues?q=` would return unnecessarily large issue objects; increasing interactive-search snippet limits would make the UI API heavier; accepting arbitrary client regex would expose avoidable database cost and ReDoS risk. ### Roadmap alignment `ROADMAP.md` does not currently list a conflicting company-search or bulk-extraction initiative. GitHub searches found no directly duplicative open issue or pull request. ## What Changed - Added shared query validation and response contracts for literal and URL extraction. - Added a company-scoped extraction service that pages issues, gathers matching issue/comment/document sources, expands URL tokens, deduplicates values, and reports truncation explicitly. - Added the authenticated route using the existing company-search authorization decision and rate limiter. - Added targeted Vitest coverage for URL extraction, multi-source dedupe, date/status filters, match caps, cross-company denial, and rate limiting. - Documented the extraction surface in the implementation specification. ## Verification - `pnpm exec vitest run server/src/__tests__/company-search-extract-service.test.ts server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-rate-limit-routes.test.ts server/src/__tests__/company-search-service.test.ts` — 30 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Bulk substring search can scan large text columns. The endpoint mitigates this with a minimum literal length, bounded issue pagination, a 20-distinct-match cap per issue, explicit truncation, existing company-search rate limiting, and no client-provided regex. - URL expansion uses a fixed server-owned pattern plus an escaped literal. A security review is requested as part of PR review to confirm the pattern and abuse controls. - No database migration or existing API response shape changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact runtime model ID and context-window size were not exposed to the session. Tool-enabled code execution and repository editing were used with medium reasoning effort. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9af96461d5 |
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work. > - Its server recovery layer classifies blocked issue graphs and restores interrupted heartbeat execution. > - A dependent issue could remain dispatch-suppressed by a cancelled blocker without producing operator-visible attention when the dependent still displayed as todo or backlog. > - Separately, a monitor-triggered run that lost its process before disposition could consume the monitor's one-shot wake without scheduling the existing bounded continuation. > - Both gaps strand useful work even though Paperclip already has the relevant blocker-attention and process-loss recovery mechanisms. > - This pull request widens the existing classification path and reuses the single process-loss retry for monitor dispatches with no future wake. > - The benefit is visible, routable recovery without weakening dependency checkout rules or introducing an unbounded retry loop. ## Linked Issues or Issue Description No matching public GitHub issue or pull request was found. ### What happened? Two server recovery cases could leave work stranded: 1. A non-terminal, agent-assigned issue with an unresolved cancelled blocker remained ineligible for checkout, but blocked-chain liveness classification only inspected issues already displaying `blocked` or `in_review`, so the existing `blocked_by_cancelled_issue` attention was not surfaced. 2. A one-shot issue monitor cleared its next check when dispatched. If that monitor-triggered run ended as `process_lost` without a tracked local child, the existing bounded retry gate rejected it and no future monitor wake remained. ### Expected behavior - Cancelled blockers continue to be unresolved dependencies, and their dependents receive blocker attention regardless of whether the dependent currently displays as backlog, todo, blocked, or in review. - A monitor-triggered run lost before disposition receives exactly one bounded continuation when no future monitor check exists; a second loss follows the normal recovery-action escalation path. ### Steps to reproduce 1. Create an agent-assigned todo issue blocked by a cancelled issue and run issue-graph liveness classification. 2. Observe that no cancelled-blocker finding appears before this change. 3. Dispatch a due issue monitor, clear its one-shot `monitorNextCheckAt`, and mark the resulting untracked run `process_lost`. 4. Observe that no retry is queued before this change. ### Environment - Paperclip commit: `3e348b96b` - Deployment: built from source / local test environment - Adapter: not adapter-specific; core server recovery - Database: embedded test database ## What Changed - Inspect non-terminal, agent-assigned issues with unresolved blocker edges during blocked-chain liveness classification. - Include cancelled dependents in the existing blocked-inbox attention query while preserving company-scoped relation checks. - Allow monitor-triggered `process_lost` runs with no future monitor wake to use the existing single bounded retry. - Mark monitor recovery retries as continuation-needed context and retain the existing second-loss escalation behavior. - Document cancelled-blocker and monitor-dispatch recovery semantics. - Add focused regressions for liveness findings, attention propagation, one retry, and second-loss escalation. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-blocker-attention.test.ts server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Low risk and server-only. The liveness scan inspects more unresolved dependency shapes, which can produce additional existing attention entries for previously invisible cancelled blockers. - Monitor recovery remains bounded by `processLossRetryCount < 1`, and the extra path only applies when the dispatch was monitor-triggered and no future monitor check exists. - No schema, migration, authorization, API-contract, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool use, command execution, and test execution capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
89ce36d7af |
feat(skills): open-by-default company skill policy and core UX (#9564)
## Thinking Path > - Paperclip uses company skills to make agent capabilities reusable across an organization. > - Skill operations currently mix capability availability with permission checks, which creates avoidable setup friction and inconsistent denial handling. > - The policy contract needs to remain open by default while allowing company-scoped restrictions for governed deployments. > - Core owns the canonical policy actions, persistence, evaluation, API behavior, safe import boundaries, and generic denial/read-only UI. > - Enterprise policy-editor implementation belongs in the separate `paperclip-ee` repository and is intentionally excluded from this PR. ### Problem or motivation Company skill operations can encounter permission dead ends even when no explicit restriction has been configured, and import-source classification can drift between policy evaluation and execution. ### Proposed solution Define eight canonical skill policy actions, default all actions to allowed, persist company-scoped restrictions, expose policy evaluation APIs, normalize import sources at the boundary, and update Skill Studio to present actionable restriction states without embedding Enterprise Edition implementation in the core repository. ### Alternatives considered Keeping capability checks distributed across routes and UI surfaces was rejected because it duplicates policy logic and makes denial behavior inconsistent. Shipping the Enterprise policy editor in this repository was rejected because `paperclip-ee` is a separate repository and must receive its own PR. ### Roadmap alignment Extends the completed **Skills Manager** roadmap area by adding coherent governance and removing workflow dead ends. ### Additional context The core API contract remains suitable for a separate Enterprise Edition editor, but this PR contains no `paperclip-ee` package or EE-specific UI integration code. ## What Changed - Added the company skill policy contract to product and implementation documentation, including the open-by-default rule, eight canonical actions, decision shape, and core/EE ownership boundary. - Added the company-scoped policy schema, migration `0170`, shared validators, policy service, REST routes, OpenAPI coverage, and focused server tests. - Hardened import policy enforcement by normalizing import sources and keeping source classification consistent between policy evaluation and execution. - Updated core Skill Studio behavior to remove generic permission dead ends and show actionable policy/platform denial states only when an operation is actually denied. - Removed the `plugin-paperclip-ee` package, Docker wiring, EE discovery/deep-link helpers, and EE-specific UI tests/stories from this PR so that implementation can be submitted separately to the EE repository. - Preserved open-by-default behavior when no explicit company restriction exists. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/skill-studio/SkillPolicySurfaces.test.tsx src/lib/skill-policy-denial.test.ts` — 20/20 passed. - `pnpm --filter @paperclipai/ui exec tsc --noEmit` — passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/worktree-config.test.ts` — 12/12 passed. - `pnpm check:token-gates` — passed with all gates clean. - `git diff --check` — passed. - `git diff --name-only origin/master | rg 'paperclip-ee|ee-skill-policy'` — no matches. ## Risks - Migration `0170` introduces company policy persistence; rollout depends on the migration applying before policy routes are exercised. - Open-by-default is an intentional behavioral policy: deployments expecting implicit denials must configure explicit restrictions. - Import normalization is security-sensitive and should retain focused review. - The separate EE editor must stay contract-compatible with the core policy API as policy actions evolve. ## Model Used - OpenAI Codex CLI, runtime model identifier and context-window size not exposed by this execution environment; reasoning, repository tool use, shell execution, and code review capabilities enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details available to this runtime) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked or described the result above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green on the latest head - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Evyatar Bluzer <bluzername@users.noreply.github.com> |
||
|
|
3db2e6bdd2 |
feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f712cedbc2 |
feat(mcp) [split 7/8]: activate Apps and gateway UI (#9562)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 7/8 and focuses on Apps/Gateways UI and feature-flagged activation > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The standalone UI foundation needs feature-flagged routes, navigation, app detail flows, gateways, Storybook scenarios, and QA configuration. - Proposed solution: Adds Apps/Gateways pages and components, navigation/route activation, remaining page integrations, Storybook stories, and QA Vite configuration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/06-ui-tools-foundation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: UXDesigner sanity pass on Apps/Gateways flows and flag-off behavior; Greptile on every PR. ## What Changed - Adds Apps/Gateways pages and components, navigation/route activation, remaining page integrations, Storybook stories, and QA Vite configuration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `pnpm check:token-gates` — all gates clean - Focused UI Vitest run with `NODE_ENV=test` — 17 files, 147 tests passed ## Risks - Navigation or flag regressions could expose incomplete experiences; activation remains controlled by existing experimental settings. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. ## UI Evidence QA captured these from the live Garden MCP split stack at 1440px and verified clean rendering:    --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a116713b9a |
feat(mcp) [split 6/8]: add Tools and Profiles UI foundation (#9561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 6/8 and focuses on UI API, shared components, and Tools/Profile surfaces > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: Operators need typed clients and administration surfaces that compile independently before navigation exposes them. - Proposed solution: Adds UI APIs, hooks, libraries, shared components, Tools/Profiles pages, and the plugin settings consumer required by the new company-scoped API. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/05-runtime-integration`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: UXDesigner sanity pass on Tools/Profile surfaces; Greptile on every PR. ## What Changed - Adds UI APIs, hooks, libraries, shared components, Tools/Profiles pages, and the plugin settings consumer required by the new company-scoped API. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `pnpm check:token-gates` — all gates clean - Focused UI Vitest run with `NODE_ENV=test` — 28 files, 183 tests passed ## Risks - Large dead-code UI additions can drift from activation routes; PR 7 supplies the registration layer and top-level parity catches omissions. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. ## UI Evidence QA captured these from the live Garden MCP split stack at 1440px and verified clean rendering:    --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
98d9360658 |
feat(adapters): confine local coding processes (#9504)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work. > - Local coding adapters currently spawn their CLI processes directly on the Paperclip host. > - CLI-native approval and sandbox flags do not provide a reliable host filesystem or network boundary. > - An agent can therefore inspect unrelated host files or fetch external material when an operator needs stronger isolation. > - The confinement must stay opt-in so existing local adapter behavior does not change unexpectedly. > - This pull request adds a shared Linux Bubblewrap spawn layer for workspace filesystem and deny/allowlist network scopes. > - The benefit is enforceable defense in depth around Codex and Claude local runs while preserving explicit provider connectivity. ## Linked Issues or Issue Description ### What happened? `codex_local` and `claude_local` processes could read arbitrary host paths and make unrestricted outbound network requests because Paperclip did not impose a spawn-level boundary. ### Expected behavior Operators can opt into a workspace-only filesystem view and either deny network egress or allow exact provider/API hosts, independently of CLI approval flags. ### Steps to reproduce 1. Run current `master` on Linux and configure a Codex or Claude local adapter. 2. Ask the agent to read a canary file outside its active workspace. 3. Ask the agent to `curl` a public host. 4. Observe that both operations succeed without a Paperclip-level confinement option. ### Environment - Paperclip commit: `c36f1a4af` / current `master` base. - Deployment mode: Linux local dev or self-hosted server. - Installation: built from source. - Adapters: Codex and Claude Code. - Database: not related. ## What Changed - Added a shared Bubblewrap process wrapper with opt-in `filesystemScope: "workspace"`, managed/extra path mounts, private `/tmp`, and Linux-only validation. - Added `networkScope: "deny" | "allowlist"`; both use a private network namespace, while allowlist mode exposes an exact-host HTTP(S) proxy over a Unix-socket bridge. - Wired Codex and Claude local CLI execution through the wrapper and forced scoped auto runs onto the CLI lane because ACP processes are not covered. - Added unit and gated Bubblewrap canaries for outside-file denial, workspace writes, direct network denial, allowlisted forwarding, and rejected destinations. - Documented both scopes, provider allowlist examples, Bubblewrap requirements, and default-off behavior. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts` — 34 passed, 4 gated Bubblewrap tests skipped by default. - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-utils build` - `pnpm --filter @paperclipai/adapter-codex-local build` - `pnpm --filter @paperclipai/adapter-claude-local build` - Attempted the gated tests with a vendored Bubblewrap binary; this container blocks unprivileged namespace setup (`setting up uid map: Permission denied` / loopback `RTM_NEWADDR: Operation not permitted`), so kernel-level execution remains for CI or a namespace-enabled Linux host. ## Risks - Bubblewrap must be installed and unprivileged user/mount/network namespaces must be enabled on the host; scoped runs fail clearly if the prerequisite is missing. - Allowlist mode depends on the coding CLI honoring standard `HTTP_PROXY` / `HTTPS_PROXY` variables; custom providers must list every required exact hostname and port. - Exact-host allowlists intentionally reject wildcards, which is safer but may require operators to enumerate multi-host provider setups. - No behavior changes unless an operator enables `filesystemScope` or `networkScope`. > This aligns with the ROADMAP direction toward safer remote and sandboxed agent environments and does not duplicate an open PR or issue found in the repository search. ## Model Used - OpenAI GPT-5.5 (`gpt-5.5`) via Codex CLI, with reasoning, repository tool use, shell execution, code editing, and test execution. The serving context-window size is not exposed to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
90f85a7d11 |
Add telemetry proposal extractor (#9544)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - It emits telemetry events to understand product usage — registered event names are gated by a generated `PAPERCLIP_EVENTS` registry so the client only enqueues known, schema-approved events > - When product teams want to instrument a new behaviour, they must first register the event name — but schema registration is a commit-and-release cycle, which creates friction in fast-moving product iterations > - A proposal lane is needed: let developers mark a `track()` call with a typed `@ts-expect-error` proposal marker so the event name can be reviewed and tracked in CI before the schema is formally registered > - The existing client had no guard against unregistered event names, so any call with an out-of-registry name (or a prototype-inherited key) would silently enter the queue, state, and network flush path > - This PR adds an `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard at the entry point of `track()` to swallow unregistered calls before any side effects, adds `scripts/extract-proposed-events.mjs` to scan source for proposal markers and emit a v2 JSON manifest with provenance and rationale, and documents the complete proposal workflow > - The benefit is that new instrumentation can be proposed and reviewed in code without touching the registered schema, and tooling can surface missing rationale before events graduate to stable ## Linked Issues or Issue Description No existing GitHub issue covers this change. This PR introduces a new feature. **Feature motivation:** Paperclip's telemetry schema is intentionally stable — registered event names are code-generated and gated. Product engineers who want to instrument a new behaviour today must land a schema change first, creating a two-step process that slows iteration. A proposal lane lets developers write the instrumentation call ahead of schema registration, protected by a compile-time `@ts-expect-error` marker that an extractor script can surface for review. This PR implements both the client-side safety gate and the extraction tooling. Refs: #9518 (closed predecessor — docs-only; this PR supersedes it with the full implementation) ## What Changed - Added `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard at the top of `TelemetryClient.track()`: unregistered event names (including prototype-inherited keys) are now swallowed before any state, queue, or network operation - Added `scripts/extract-proposed-events.mjs`: scans TypeScript source for `@ts-expect-error -- proposed-telemetry(<issue>): <rationale>` markers; emits a v2 JSON manifest per proposed event including name, rationale, provenance (repo-relative file + line), and a `rationale_missing` flag for CI enforcement - Added `scripts/extract-proposed-events.test.mjs`: test suite covering marker parsing, multi-line markers, path validation, out-of-repo rejection, and the v2 schema output contract - Added `doc/TELEMETRY_WORKFLOW.md`: documents the proposal workflow, the canonical multi-line marker example, rationale requirements, and how to graduate a proposed event to stable schema - Updated `packages/shared/src/telemetry/README.md`: added "Proposed Events" section to the Telemetry Data Contract per the contributing guide requirement for telemetry changes ## Verification Run all of the following from the repo root: ```sh # Extractor unit tests node --test scripts/extract-proposed-events.test.mjs # Telemetry client + types tests pnpm exec vitest run --config vitest.config.ts \ src/telemetry/client.test.ts src/telemetry/client-types.test.ts \ --reporter=verbose # (run from packages/shared) # Type-check pnpm --filter @paperclipai/shared typecheck # Smoke-run the extractor in local-test mode node scripts/extract-proposed-events.mjs --ref local-test ``` All four commands pass locally. ## Risks - **Silent drop on unregistered events:** The `Object.hasOwn` guard fails closed — any event name not in `PAPERCLIP_EVENTS` is silently dropped. If the generated registry is missing an event that was previously tracked, those calls will be silently lost. Mitigation: the extractor script surfaces proposed events that need registration; the TypeScript type system already enforces `TelemetryEventName ⊆ PAPERCLIP_EVENTS` at compile time. - **Extractor is read-only:** `extract-proposed-events.mjs` reads source and emits JSON; it does not modify any files. No runtime or schema risk. - Overall risk: **low**. The guard is additive and defensive; the extractor and docs are additive only. ## Model Used - Provider: Anthropic - Model ID: `claude-sonnet-4-6` - Context window: 200 K tokens - Capabilities: tool use, extended context, code generation ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8b6a06ee25 |
[codex] Add built-in agents and Reflection Coach bundle (#9206)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators need first-party agent capabilities for repeatable company work, not just manually created one-off agents. > - Built-in agents need to behave like normal company-scoped agents while preserving approval gates, permissions, budgets, and audit trails. > - Reflection and coaching work also needs bundled instructions, skill content, and a routine so the feature can be installed and reset predictably. > - The API, database, UI, portability, and tests all need to agree on the built-in lifecycle from not provisioned through setup, approval, ready, paused, and reset. > - This pull request adds built-in agent provisioning and the Reflection Coach bundle end-to-end. > - The benefit is a safer first-party path for Paperclip-managed agents without bypassing the same governance model used for operator-created agents. ## Linked Issues or Issue Description No public GitHub issue was found for this exact built-in agent and Reflection Coach bundle work. Problem/motivation: - Paperclip did not have a first-party built-in agent lifecycle for product-owned agents. - Bundled agent resources such as default instructions, skills, and routines needed managed ownership and reset semantics. - Approval-gated companies needed built-in setup to preserve requested adapter, budget, manager, and permission state through board approval. - The board UI needed clear built-in badges, setup affordances, readiness state, and bundle status without exposing secrets. Proposed solution: - Add a company-scoped built-in agent registry, provisioning/reset/reconcile/status APIs, and Reflection Coach bundled resources. - Track bundled managed resources in the database with idempotent migration behavior. - Reuse existing agent approval, authorization, budget, and activity-log paths instead of creating a bypass. - Add UI setup, badges, gates, bundle panels, and route coverage for built-in agents. Duplicate search: - Searched GitHub PRs for `built-in agents Reflection Coach repo:paperclipai/paperclip`; only this PR was returned. - Searched GitHub issues for the same query; no public issues were returned. ## What Changed - Added built-in agent definitions, lifecycle state derivation, provisioning, reset, reconcile, status, and routine-control routes. - Added the `built_in_managed_resources` migration and schema exports for bundled instructions, skill, and routine ownership. - Added the Reflection Coach built-in bundle with default instructions, skill catalog content, routine template, default permissions, and managed-resource drift handling. - Added approval-aware provisioning behavior that preserves requested adapter config, budgets, manager assignment, and built-in permissions through hire approval. - Added authorization and mutation gates for built-in agent and skill changes, including consented Reflection Coach change paths. - Added UI surfaces for built-in agent setup, roster/detail badges, readiness gates, bundle status, routine controls, and route filtering. - Added company import/export and validator coverage for built-in managed resources and low-trust/red-team presets. - Addressed Greptile follow-ups for pending approval reconciliation, consent-gate error propagation, config-read authorization fallback, approval-path manager preservation, and non-model adapter provisioning. ## Verification Local verification: - `git diff --check public/master..HEAD` passed. - `pnpm check:token-gates` passed with all gates clean. - `pnpm exec vitest run ui/src/components/ConfigureBuiltInAgentModal.test.tsx` passed: 1 file, 4 tests. - `pnpm exec vitest run ui/src/components/EntityRow.test.tsx ui/src/pages/Agents.test.tsx ui/src/components/BuiltInAgentGate.test.tsx ui/src/components/ConfigureBuiltInAgentModal.test.tsx ui/src/components/BuiltInBundlePanel.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx ui/src/pages/Routines.test.tsx` passed: 7 files, 64 tests. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts src/__tests__/authorization-service.test.ts src/__tests__/company-skills-routes.test.ts` passed: 3 files, 91 tests. - `pnpm --filter @paperclipai/db check:migrations` passed. - `pnpm -r typecheck` passed after the rebase; `pnpm --filter ui typecheck` passed after the final UI review fix. Remote verification on latest head `1c61f693a4ec881d739022b0e75a8ca8bf8c2cd8`: - Merge state: `CLEAN`. - Greptile: `5/5`, zero unresolved Greptile threads. - PR check rollup: all checks successful, neutral, or skipped as expected. - Passing gates include Build, Typecheck + Release Registry, all server shards, all workspace shards, all serialized server suites, e2e, Canary Dry Run, policy, review, verify, Socket, Superagent, and Snyk. ## Risks - This adds a new managed-resource table and migration; the migration uses idempotent create/add/index guards and passed migration safety checks. - Built-in agent provisioning touches approval and authorization paths; tests cover pending approval preservation, stale retry rejection, consent gates, and config-read fallback behavior. - Reflection Coach creates managed instructions, skill, and routine resources; drift/reset behavior is covered by service tests and redacted API responses. - Non-model adapter setup now provisions a `needs_setup` built-in row before command/endpoint fields are complete; this matches the server lifecycle and is covered by the setup modal regression test. ## Model Used OpenAI Codex coding agent based on GPT-5. Exact hosted model ID, context-window size, and reasoning-mode labels are not exposed in this runtime; tool use, shell execution, GitHub CLI/API access, and local code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53d09d4c34 |
fix(ui): inbox/task list parity, nesting alignment, hover perf, and routine detail polish (#9317)
## Thinking Path > - Paperclip's UI is governed by the design system merged in #9134 and the component convergence in #9240 — one Card, one Badge, one nav row, one `IssueRow`, a single multiplicative radius ladder. > - With those primitives in place, the remaining rough edges were interaction and alignment details on the surfaces people use every day: the inbox, the task list, the sidebar, and the routine/task detail pages. > - Each item here was reported from live use and fixed against a running instance, then verified by measurement (pixel alignment, frame timing) rather than by eye alone. > - The result is that the inbox and task lists now behave as one system (same hover, keyboard nav, tree-guide, and archive language), list nesting reads correctly, list hover is smooth, and the routine detail page scrolls and aligns like the rest of the app. ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template. - **Problem**: after the design-system foundation (#9134) and component convergence (#9240), the inbox/task lists still had interaction and alignment gaps — hover lag on long lists, keyboard navigation that only partly matched between the two lists, workspace/parent nesting whose guides and chevrons didn't line up, an inbox that sat offset from the task list, and a routine detail page with an odd double-scroll and a bespoke sub-nav. - **Proposed behavior**: the inbox and task lists share one interaction contract (hover, keyboard nav, collapse, archive), list nesting aligns to the status column with clean chevrons, list hover is CSS-only (no per-hover re-render), and the routine detail page uses a fixed header/sub-nav with a single scrolling content region and a sub-nav that matches the primary nav. ## What Changed - **List hover performance**: hover is now painted purely by CSS `:hover` and records the hovered row in a ref, instead of writing list-selection state on every `mouseenter` (which re-rendered 100–300 non-memoized rows per hover). Keyboard nav reads the ref so it still continues from the hovered row; the keyboard band clears on the first real mouse move so hover and keyboard selection never show two bands at once. The row's `transition-colors` fade was removed so the highlight snaps (no comet-tail). Measured on a 100-row scrub: worst frame **333 ms → 33 ms**, long frames (>50 ms) **12 → 0**, avg **41 → 60 fps**. - **Inbox ↔ task-list parity**: keyboard navigation works on every inbox tab (archive/read keys stay scoped to the archivable tab); group headers and parent tasks collapse/expand with the arrow keys in both lists; the task list gains the same j/k / arrows / Enter selection model as the inbox; hover selection bands match; the `g` then `i` go-to-inbox chord works app-wide. - **List nesting & alignment**: the workspace group-header chevron lines up exactly with the task chevrons below it (both lists); the parent→child connector line drops from under the parent's **status icon** rather than its chevron and breaks with a 14px gap around a nested row's own chevron; and inbox rows line up with the task list (read rows no longer reserve a mark-read column). - **Inbox archive affordance**: moved from a bare `x` left of the status icon to an `Archive` icon + label button on the right (before the timestamp), revealed on row hover; the left slot now carries only the unread dot; swipe-to-archive is unchanged. - **Sidebar**: when no agent has a live run, the AGENTS section shows 3 recent agents (was 5) plus "See all agents"; the working-agents view is unchanged. - **Routine detail page**: the layout is bounded to the main scroll area so the header (Run now + automation toggle) and the sub-nav stay fixed and only the section content scrolls (was a page-level scroll competing with a `sticky` sub-nav). The sub-nav items adopt the primary nav's rhythm — row padding/height, inset rounded pill, type scale, and 16px icons — and its background matches the main nav. - **Task detail**: dropped the redundant `🔵` prefix the breadcrumb added for live/in-progress tasks; the status glyph already conveys that state. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2255/2255 (assertions updated in lockstep where behavior changed) - `pnpm --filter @paperclipai/ui build` → exit 0 - Storybook visual regression: run locally throughout (the baseline-manifest archive is still unpublished, so CI cannot run this suite — pre-existing condition from #9134). Every visible delta was reviewed against a live instance and, where it was intentional (nesting alignment, routine sub-nav restyle, unread-row shift), the affected snapshots were re-baselined locally. - Manual/measured: inbox + task lists (grouped and nested, light + dark), hover-scrub frame timing, keyboard navigation, and the routine detail scroll/alignment were exercised on a running instance. ## Risks - Behavior-and-alignment changes concentrated in `IssueRow` / `IssuesList` / `Inbox` (the shared task-row surfaces). The riskiest area — the hover/keyboard-selection model — is covered by unit tests (updated in lockstep) and was measured and driven live. - The routine detail scroll change restructures that page's layout container; verified the page's own scroll stays fixed while only the section content scrolls. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing); snapshot coverage is local-only until that lands. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright measurement/screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c90e66bdd6 |
feat(ui): design-system component convergence — Card/Badge adoption, multiplicative radius ladder, unified list surfaces (#9240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is governed by a design system (`DESIGN.md` + the token layer in `ui/src/index.css`, merged in #9134) whose first principle is "one way to say each thing" — one Card, one Badge, one nav row > - After the token extraction landed, ~35 files still hand-rolled card containers, ~45 files hand-rolled pill spans, the sidebar agents section duplicated the nav-row chrome, and the inbox and tasks lists rendered the same task rows two subtly different ways > - Each divergence is a place where a future design change (radius, hover language, status vocabulary) silently misses surfaces, defeating the "edit tokens + run checks" model the design system exists for > - This pull request converges those surfaces onto the shared primitives, codifies the radius scale as the modern multiplicative shadcn ladder, and unifies the row/hover/tree-guide language across the inbox and tasks lists — every visible delta was human-reviewed screen-by-screen against a live instance across nine feedback rounds > - The benefit is that the app's look is now steerable from single knobs (one `--radius` anchor, one Card, one Badge, one row component), and the visual regression suite covers the result (514 snapshots including a new AgentDetail page story) ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template: - **Problem**: after the design-token foundation (#9134), component-level drift remained — hand-rolled cards/pills, duplicated sidebar row chrome, and two different renderings of task rows (inbox vs tasks list) meant design changes had to be applied per-surface and frequently missed spots (e.g. status glyphs rendered 16px in the inbox but 20px in the tasks list because a slot override silently beat the component default). - **Proposed behavior**: all card-shaped containers render via `Card`, all label pills via `Badge`, sidebar rows via `SidebarNavItem`, and both task-list surfaces via one `IssueRow` configuration; the radius scale is a single multiplicative ladder anchored at `--radius: 0.5rem`. - **Alternatives considered**: converting interactive `<button>`/`<Link>` cards to `Card` divs (rejected — breaks semantics; documented inline with `design-allow` comments instead); keeping the legacy additive radius ladder (rejected in favor of the standard shadcn multiplicative mapping). ## What Changed - `Card` adoption across ~35 files (settings pages, auth/board flows, dashboards, list containers, KPI tiles); non-adoptable sites (interactive cards, `<li>` rows, class-string props, chart tooltip) carry documented `design-allow(card-pattern)` comments - `Card` gains an `interactive` prop — one quiet hover affordance for clickable cards (cursor, border darken, shadow lift, focus ring), applied to skills tiles, artifact cards, and the company selector; cards carry no resting shadow - `Badge` adoption for 113 hand-rolled pill spans across ~45 files; `PropertyChip` wraps `Badge` internally; `StatusBadge`, external-object chips, and match chips stay bespoke by documented decision (WCAG-tuned status mechanics) - Radius ladder becomes the multiplicative shadcn mapping (`sm/md/lg/xl/2xl/3xl/4xl = 0.6/0.8/1.0/1.4/1.8/2.2/2.6 × --radius`, anchor `0.5rem`); every card surface unifies on `rounded-lg`; the orphaned 8px literal token is deleted - Sidebar: agent rows render via `SidebarNavItem` (new additive props: `iconNode`, `active`, `trailing`, `liveAccessory`); live dots use `--status-agent-running`; one row rhythm and inset rounded pill highlight; right-aligned trailing badges; every labeled section is collapsible - Inbox + tasks lists unified: md status glyphs, `accent/50` rounded row hovers, vertical tree guides under parent rows (opaque underlay so dark-mode translucent borders don't stack), no horizontal dividers under expanded parents; the swipe-to-archive reveal layer shows only mid-swipe; board toggle uses the `SquareKanban` glyph - Kanban: every column carries a status-hued tint; lanes default expanded (including empty); compact mode collapses empty lanes to labeled rails (fixes a clipped, label-less empty-column state) - Storybook: new AgentDetail page story (realistic fixtures, light+dark) joins the visual suite; suite captures with `reducedMotion: 'reduce'` and the ux-lab reasoning ticker honors `prefers-reduced-motion`; a stale lexical alias in `storybook/main.ts` is fixed (build was broken since the lexical 0.46 bump) - Keyboard navigation, from live review of the unified lists: inbox navigation keys work on every tab (archive keys stay scoped to the archivable tab); keyboard-driven scrolling no longer hands the selection to whatever row lands under the stationary cursor (hover selects only after real pointer movement); the tasks list view gains the same j/k / arrows / Enter selection model as the inbox; and the `g` then `i` go-to-inbox chord works app-wide instead of only on the issue detail page - Token gates restored to 3/3 CLEAN (tokenized a post-#9134 regression in the recovery card); decisions recorded in `doc/design/DECISION-SHEET.md` and `doc/design/COMPONENT-INVENTORY.md` (investigation verdicts: FileTree vs WorkspaceFileBrowser and the four entity pickers stay separate — evidence included) ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2106/2106 (assertions updated in lockstep where they documented superseded decisions; new tests for the global go-to-inbox chord) - `pnpm --filter @paperclipai/ui build` → exit 0 - `pnpm build-storybook` → succeeds (also fixes the lexical-alias break on master) - Visual regression: 514-snapshot Playwright suite green against the updated baseline (zero diffs from the keyboard-navigation round — those changes are purely behavioral). Note: baselines live outside git per the suite design and the baseline-manifest archive is not yet published, so CI cannot run this suite — it was run locally throughout; every visible delta was reviewed screen-by-screen in a live instance across nine review rounds. Review evidence (before/after triplets) intentionally kept out of the repo for size; available on request. - Manual: exercised dashboard, tasks (list + board), inbox, agents, skills, costs, settings, and artifact surfaces in light and dark themes ## Risks - Wide but shallow visual surface: most changes are class-string substitutions with behavior preserved (props, handlers, roles, test ids). The riskiest areas — dnd-kit card refs (React 19 ref-as-prop), inbox swipe-to-archive, and sidebar overlays — are covered by existing unit tests (all green) and were manually exercised. - Intentional visual deltas (rounded cards, tinted kanban columns, md status glyphs, unified hovers) are design decisions recorded in `doc/design/DECISION-SHEET.md`; each maps to a re-baselined snapshot set locally. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing condition from #9134); until that lands, snapshot coverage is local-only. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending first review) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3919713bc |
[codex] Document Storybook visual baseline platform lock (#9216)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Storybook visual baselines protect UI surfaces from unintended visual drift. > - Pixel-perfect screenshot baselines are sensitive to OS, font rasterization, and browser environment. > - The suite already stores external baseline artifacts and has an opt-in CI path. > - Local runs on non-matching platforms can report false-positive diffs unless the platform lock is explicit. > - This pull request documents the Linux/Ubuntu baseline constraint and makes the local static server port explicit. > - The benefit is clearer visual-review guidance and more predictable Playwright web server startup. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? The Storybook visual baseline suite requires a matching Linux capture environment for pixel-exact comparisons, but the docs did not clearly warn local users that non-Linux environments can produce false-positive diffs. The Playwright web server command also relied on the static server's default port instead of passing the configured port explicitly. ### Expected behavior Developers should see clear Linux/Ubuntu baseline guidance before running the visual suite locally, and Playwright should start the Storybook static server on the same explicit port that the test config expects. ### Steps to reproduce 1. Review the Storybook visual docs before this PR. 2. Run or inspect the Storybook visual Playwright config. 3. Notice the missing platform guidance and implicit static server port coupling. ### Paperclip version or commit Reproducible on `master` before this branch. ### Deployment mode Local dev (pnpm dev) / built from source. ## What Changed - Documents the Linux/Ubuntu-only baseline limitation in the developer docs and visual-suite README. - Adds `--port` parsing and validation to the Storybook static server helper. - Adds regression coverage for `--port` followed by another flag. - Passes the Playwright web server port explicitly from the Storybook visual config. ## Verification - Passed: `node --check scripts/serve-storybook-static.mjs` - Passed: `node --test scripts/__tests__/serve-storybook-static.test.mjs` - Passed: `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` - Greptile: 5/5 with no unresolved review threads after commit `94a649755a2ae7c4a34a3e8a1f16ec4d26d738fd`. - Not run: full `pnpm test:storybook-visual`, because it builds Storybook and runs the browser visual suite; this PR only changes docs plus server port plumbing. ## Risks Low risk. The server still defaults to port 6106 when no explicit port is provided, and invalid port values now fail fast with a clear error before the Playwright server waits for an unreachable URL. ## Model Used OpenAI GPT-5 Codex coding agent with local command execution and repository editing tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c07e650cd7 |
feat(ui): single-source design tokens, visual regression suite, and theme retune (#9134)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be821a4f7e |
Fix DB backup health alerts (#9147)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators depend on `/api/health` and OpenAPI status surfaces to know whether the local control plane is healthy. > - Database backups are a safety-critical background process, but backup failures were not represented in health responses. > - That gap means an instance can look healthy while backup state is stale, failing, or unavailable. > - This pull request adds backup-health evaluation and exposes it through the health route, server startup wiring, and OpenAPI contract. > - The benefit is earlier operator visibility when automatic backups stop protecting instance data. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug report: **Pre-submission checklist** - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip (or can reproduce on `master`). - [x] I have confirmed the error originates in Paperclip itself — not in my agent adapter, API provider, or local configuration. **What happened?** Automatic database backup health was not included in the app health response, so backup failures or stale backups could be missed while `/api/health` still looked otherwise usable. **Expected behavior** The health endpoint should include backup-health details that let operators identify disabled, stale, failing, or healthy backup states. **Steps to reproduce** 1. Configure a Paperclip instance with automatic database backups. 2. Force backup status into a stale or failing state. 3. Call `/api/health` and inspect whether backup state is represented. **Paperclip version or commit** `master` at the PR base. **Deployment mode** Local dev (`pnpm dev`) and self-hosted server deployments. **Installation method** Built from source (`pnpm dev` / `pnpm build`). **Agent adapter(s) involved** - [x] Not adapter-specific (core bug) **Database mode** Embedded development Postgres and external Postgres backup paths. **Access context** Board/operator health checks. **Relevant logs or output** Covered by the added `server/src/__tests__/health.test.ts` cases. **Relevant config (if applicable)** Not applicable. **Additional context** This surfaces backup status only; it does not change backup execution scheduling. **Privacy checklist** - [x] I have reviewed all pasted output for PII (usernames, file paths, API keys, tokens, company names) and redacted where necessary. ## What Changed - Added a database backup health service that classifies backup recency, status, and failure conditions. - Wired backup health into app/server startup and the health route response. - Documented the backup-health behavior in development docs and OpenAPI output. - Added focused health route tests for healthy, stale, disabled, and failing backup states. ## Verification - `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/vitest run server/src/__tests__/health.test.ts` ## Risks Low-to-medium risk. This changes health response content and may affect external health consumers that parse fields strictly. It should not alter backup execution itself. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5.5 coding agent with repository tool use and local shell execution. Context window was not surfaced by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
329652a2dd |
refactor(db): add migration authoring checklist (#9122)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip stores agent work in a PostgreSQL database that evolves via numbered sequential migrations > - Migrations that run large unbounded table scans can block the server's `listen()` call during upgrades, causing multi-minute startup stalls on large databases > - Migration 0126 ran a full sequential scan backfill over `issue_comments` for derived attribution columns — O(n²) due to unindexed `LIMIT`/`OFFSET` batching, pegging CPU for ~5 minutes on a 3–4 M row table > - The safe fix (landed in #9108) replaced 0126 with a new forward-only migration using a partial index + keyset pagination backfill > - But the root cause is the absence of author-time guidance: contributors have no documented rules for writing bounded, indexed migration backfills before they land > - This PR adds a migration authoring checklist to `doc/DATABASE.md` so future contributors have those rules at hand before opening a PR > - The benefit is a durable, discoverable guide that prevents the same class of startup-blocking slowness before it reaches production ## Linked Issues or Issue Description This PR is a documentation follow-on to #9108, which landed the fast 0132 migration fix. It adds author-time guidance that captures the root-cause lesson from that incident. No separate public issue exists for the doc addition; the motivation is described above. Refs #9108 ## What Changed - `doc/DATABASE.md`: Added a **Migration authoring checklist** section with rules for indexed, bounded backfill batches — keyset pagination over `LIMIT`/`OFFSET`, mandatory partial index, idempotent guards, and split-phase schema-vs-data changes. The `check:migrations` CI gate is referenced as the enforcement backstop. ## Verification - `git diff --check -- doc/DATABASE.md` passes (no whitespace errors). - No executable code changed; the checklist is an additive documentation section. ## Risks Low. The change is additive text in `doc/DATABASE.md`. No schema, migration, or code changes. No behavioral diff. ## Model Used Claude claude-sonnet-4-6 (Anthropic, 200 K context, tool use) — used to author the migration authoring checklist and coordinate the PR workflow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with \`Fixes: #\` / \`Closes #\` / \`Refs #\` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub \`#NNN\` / \`github.com/paperclipai/paperclip\` URLs) - [x] My branch name describes the change (e.g. \`docs/...\`, \`fix/...\`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5cdf5103c9 |
docs(spec): humans/permissions granularity is V1, not OOS (#6744)
## Thinking Path > - `doc/SPEC-implementation.md` §5.2 (Out of Scope V1) conflates two distinct concerns into a single bullet: _"Multi-board governance or role-based human permission granularity"_. > - Role-based human permission granularity has been V1 for a while — the `humans-and-permissions` plan and the `principal_permission_grants` table shipped, the `PERMISSION_KEYS` set covers `users:invite`, `users:manage_permissions`, `tasks:assign`, `tasks:assign_scope`, `tasks:manage_active_checkouts`, `tasks:view_all`, `agents:view_all`, `joins:approve`, `agents:create`, `environments:manage`. > - The remaining OOS item from that bullet is _multi-board governance_ — running multiple board UIs against one company. That's a deployment-topology concern, separate from permission granularity, and stays V1 OOS. > - This PR splits the bullet so the OOS list reflects reality and contributors don't read the SPEC and conclude that per-user permission scoping is unplanned. ## What Changed Single-file docs edit in `doc/SPEC-implementation.md` §5.2: - Drop the conflated "Multi-board governance or role-based human permission granularity" bullet. - Keep "Multi-board governance (multiple board UIs for a single company)" as OOS — that part is still out of scope. - Add a short paragraph below the OOS list pointing readers to the `humans-and-permissions` plan, `principal_permission_grants`, and the existing `tasks:view_all` + `agents:view_all` opt-out scoping primitives so anyone reading §5.2 finds the V1 surface immediately. ## Verification - N/A — docs-only change, no code/schema/test impact. ## Risks - None. Single bullet rewording. ## Model Used Claude (Anthropic). Model ID: `claude-opus-4-7`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work — it documents already-shipped V1 work that the SPEC was lagging on - [x] I have run tests locally and they pass — N/A docs-only - [x] I have added or updated tests where applicable — N/A - [x] If this change affects the UI, I have included before/after screenshots — N/A docs-only - [x] I have updated relevant documentation to reflect my changes — this PR IS the doc update - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Mike <ms@moar.tools> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
574b4e71db |
[codex] Bundle UI webfonts with the app (#9020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The browser UI depends on the Inter font family for its intended visual baseline > - The app previously referenced remote Google Fonts stylesheets at runtime > - That meant self-hosted, offline, or privacy-sensitive deployments could lose the intended typography or depend on an external request > - This pull request bundles the required Inter variable font files with the UI and serves them from the app > - A normal UI test now covers the static assets and CSS wiring without adding a package script or build step > - The benefit is a more reliable, self-contained UI that does not rely on third-party webfont hosting ## Linked Issues or Issue Description No public GitHub issue was found for this exact gap. **Subsystem affected** ui/ — React + Vite board UI **Problem or motivation** Paperclip's UI should ship the webfont assets it references so production and self-hosted deployments render consistently without reaching out to Google Fonts at runtime. **Proposed solution** Bundle the Inter variable font files under the UI public assets, load them with local `@font-face` declarations, document the bundled assets, and cover the source assets/CSS wiring with a normal UI Vitest test. **Alternatives considered** Keeping the remote stylesheet dependency is simpler, but leaves deployments dependent on external font hosting. Using system fonts only would avoid the asset footprint, but changes the intended UI typography. **Roadmap alignment** This is a focused UI reliability/polish fix, not a roadmap-level core feature. **Additional context** Searched public GitHub issues and PRs for `webfonts repo:paperclipai/paperclip`; no duplicates or closely related open items were found. ## What Changed - Added bundled Inter variable font assets and their notice under `ui/public/fonts/`. - Replaced remote Google Fonts imports with local `@font-face` declarations using relative public-asset URLs that remain subpath-safe from built CSS. - Documented the local font asset expectation in development and UI spec docs. - Removed the follow-up font asset checker scripts and package/build wiring after review feedback clarified they are not required for building the UI. - Added `ui/src/lib/ui-font-assets.test.ts` to verify the shipped WOFF2 files, notice text, and CSS font references through the normal UI test suite. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/ui-font-assets.test.ts --config vitest.config.ts` - Passed. - Confirms the bundled font files exist, are WOFF2 files, have notice coverage, and are referenced by `ui/src/index.css`. - `pnpm --filter @paperclipai/ui build` - Passed. - Confirmed the UI still builds after removing the checker from `ui/package.json`. - Confirmed `ui/dist/fonts/` contains `InterVariable.woff2`, `InterVariable-Italic.woff2`, and `NOTICE.md` after the build. - The UI build emitted existing warnings about `::highlight(...)`, a dynamic/static import overlap for `MarkdownEditor.tsx`, unresolved relative public font URLs left for runtime resolution, and large chunks, but completed successfully. ## Risks Low risk. This adds static font assets and swaps the font source from a remote stylesheet to same-origin files. The main tradeoff is a larger repository/UI asset footprint from the bundled `.woff2` files. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/GitHub workflow with reasoning support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad961227f5 |
feat(secrets): add user-specific runtime secrets (#8825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs often need provider credentials, API tokens, and other environment-bound secrets. > - Company-level secrets work for shared credentials, but they do not model values that should differ by human operator. > - Without a user-scoped model, a run can dispatch without knowing whether the responsible human has supplied the needed value. > - Paperclip also needs run attribution to make those user-scoped runtime checks deterministic and auditable. > - This pull request adds user-specific secret definitions, per-user values, environment bindings, responsible-user attribution, and runtime resolution gates. > - The benefit is that teams can define the secret once, let each user provide their own value, and block runs before dispatch when required user secrets or active definitions are unavailable. ## Linked Issues or Issue Description Refs #224 Refs #6057 This PR implements user-specific secret support as a core secret-management capability rather than a one-off adapter setting. It is related to existing public work on company secrets UI and runtime secret refs, but is distinct because the value is owned by the responsible user and resolved at run dispatch time. Related PR search before opening found existing secrets work such as #1550, #8256, #8614, #8634, and #8647; none of those add the full user-secret definition/value/runtime gate covered here. ## What Changed - Added user-secret definitions and per-user "My secrets" values, keeping stored values out of access metadata. - Added `user_secret_ref` environment bindings and UI affordances to pick them alongside existing secret refs. - Added responsible-user runtime resolution so user-secret refs resolve against the human responsible for the run. - Added pre-dispatch missing-secret gates so runs fail before adapter dispatch when required user values are absent or definitions are inactive. - Added low-trust allowlist hardening for user-secret runtime access. - Added issue, routine, run, and agent API key responsible-user attribution and fail-closed dispatch behavior when attribution cannot be resolved. - Added denial-copy mapping so responsible-user authorization failures surface as actionable run outcomes instead of opaque setup failures. - Added OpenAPI documentation for the user-secret routes. - Rebases cleanly on current `master`; migrations were renumbered incrementally as `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant` after upstream `0126`/`0127` migrations. - Removed previously committed local design screenshots so the PR contains code/docs/tests only. ## Verification - PASS: PR head `2527febd106bcf3ca264ca0da7fca491084192d6` is based on `paperclipai/paperclip:master`. - PASS: `git diff --check` - PASS: `git diff --name-only public/master...HEAD | rg '^(pnpm-lock\\.yaml|\\.github/workflows/|screenshots/)' || true` produced no files. - PASS: migration journal audit confirmed unique indexes through `130` with tail entries `0126_issue_comment_derived_attribution`, `0127_environment_custom_images_instance_scoped`, `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant`. - PASS: `pnpm --filter @paperclipai/ui typecheck` - PASS: `pnpm --filter @paperclipai/server typecheck` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-active-run-output-watchdog.test.ts src/__tests__/heartbeat-stale-queue-invalidation.test.ts src/__tests__/heartbeat-workspace-finalize-branch.test.ts src/__tests__/issue-monitor-scheduler.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-comment-wake-batching.test.ts src/__tests__/heartbeat-retry-scheduling.test.ts src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts src/__tests__/heartbeat-plugin-environment.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/low-trust-red-team-routes.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` (55 tests) - PASS: `pnpm vitest run server/src/__tests__/secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` (89 tests after final Greptile cleanup fixes) - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-issue-liveness-escalation.test.ts` (17 tests after the final rebase CI fix) - PASS: focused server Vitest batches covering heartbeat recovery, project env, plugin env, routines, low-trust, pipelines, monitors, watchdog, and stale queue paths. - PASS: GitHub checks are green on `2527febd106bcf3ca264ca0da7fca491084192d6`, including Typecheck + Release Registry, Build, General tests, serialized server suites, e2e, Canary Dry Run, verify, security checks, and Greptile Review. - PASS: Greptile Review completed successfully on `2527febd106bcf3ca264ca0da7fca491084192d6` with Confidence Score 5/5, and GraphQL review-thread audit returned zero unresolved non-outdated threads. ## Risks - Runtime behavior now depends on a run having a correct responsible user; missing or incorrect responsibility assignment can block runs before adapter dispatch. - `user_secret_ref` bindings intentionally expose metadata without values, but UI/API callers may need to handle the new binding kind explicitly. - External secret providers and IAM policies are not automatically provisioned by this PR; operators still need to configure provider-side access for non-local vaults. - The PR is broad across db/shared/server/UI/runtime paths, so release validation should include both API and UI secret workflows before merge. - The migration renumbering is intentionally incremental after upstream migrations; the branch migrations use guarded column/table/index/constraint creation so users who tested the older draft numbering should not hit duplicate DDL for the existing objects. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), Codex local adapter with shell/tool use and code execution. Context window and internal reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
936687ca55 |
fix(workspace): restore clean branch drift on finalize (#8914)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs can execute inside reusable, runtime-created git worktree execution workspaces. > - Those managed worktrees record the expected branch so later dispatches do not accidentally run an agent in the wrong checkout. > - Successful run finalization already checked branch coherence, but it treated every unrecorded branch switch as fatal. > - A common publishing flow can briefly switch a clean worktree to a PR/publish branch that points at the same commit as the recorded issue branch, leaving no divergent work to protect. > - This pull request keeps the strict finalization guard for unsafe drift, but lets finalization restore the recorded branch when same-commit repair is provably safe. > - The benefit is fewer false failed runs after harmless branch switches while preserving hard failures for divergent or dirty worktrees. ## Linked Issues or Issue Description No public issue exists for this exact finalization failure. Related public worktree-recovery context: #3087 and #3056, but those address different worktree realization/reuse recovery paths rather than successful-run finalization branch repair. Bug report details: **What happened?** When an adapter run succeeded after switching a managed git worktree from its recorded issue branch to a publish/PR branch, finalization failed with a managed worktree branch mismatch even when the publish branch and recorded branch pointed at the same commit and the worktree was clean. **Expected behavior** Finalization should restore the recorded branch only when it can prove the worktree is clean, registered, and the recorded branch points at the current `HEAD`. If the actual branch has different commits or unsafe state, finalization should continue to fail with bounded validation evidence. **Steps to reproduce** 1. Create a runtime-managed `git_worktree` execution workspace for an issue run. 2. During the adapter run, create and check out a new publish branch without committing new changes. 3. Return adapter success and let heartbeat finalization run. 4. Before this change, finalization records a failed branch check and fails the run even though the branches point at the same commit. 5. With this change, finalization records the repair operation, restores the recorded branch, and records a successful finalize row. 6. Repeat with a commit on the publish branch; finalization still fails because the branch heads differ. **Paperclip version or commit** Reproduced against `master` at `bac7307ec`; fixed by this PR at `64ec605cf`. **Deployment mode** Local dev / built from source. **Agent adapter(s) involved** Not adapter-specific. This is core heartbeat/workspace finalization behavior. **Database mode** Embedded test Postgres in the focused server test. **Access context** Agent run finalization. **Node.js version** `v25.6.1` **Operating system** `Darwin 24.6.0 arm64` **Relevant logs or output** The new focused test intentionally exercises both outcomes: ```text Test Files 1 passed (1) Tests 3 passed (3) ``` **Relevant config** Runtime-created `git_worktree` execution workspace. **Additional context** The unsafe divergent branch case still fails with `workspace_validation_failed` and `git_worktree_branch_incoherence` evidence. **Privacy checklist** Reviewed; this description avoids internal task links, local workspace paths, credentials, and instance-specific URLs. ## What Changed - Reused the existing guarded branch-coherence repair helper during heartbeat finalization when the final branch inspection finds clean same-commit branch drift. - Recorded repair metadata in the `workspace_finalize` operation so reviewers/operators can audit whether finalization repaired branch drift. - Preserved failure behavior for divergent branch heads and surfaced the bounded workspace validation evidence from the repair helper. - Added focused server coverage for safe finalization repair and unsafe divergent branch failure. - Updated execution semantics docs to describe the narrower finalization rule. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks Low to medium risk. The change affects successful-run finalization for runtime-created git worktree execution workspaces. The repair path is constrained to clean, registered, same-commit branch drift, and the focused test confirms divergent branch heads still fail instead of being restored silently. ## Model Used OpenAI Codex, GPT-5-based coding agent. Exact hosted model ID was not exposed in the runtime; tool use and local shell execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2eba718bef |
Fix sandbox bridge credentials and stalled review recovery (#8844)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local adapter and heartbeat recovery systems decide whether an agent has a real control-plane mutation path. > - Sandboxed local adapters split execution between the trusted host process and the sandbox shell/tool surface. > - A host-side adapter can still reach Paperclip while the sandbox shell surface cannot, which leaves agents thinking no endpoint or credentials are configured even though the host can still post comments. > - Execution-policy review stages can also remain pending after a reviewer run finishes without recording a decision. > - This pull request makes the sandbox bridge available to the actual shell mutation surface and adds bounded recovery for terminal-but-still-pending review participants. > - The benefit is that agents get a real reachable Paperclip API path where they need it, and stalled review stages become visible recovery work instead of silently drifting. ## Linked Issues or Issue Description No exact public GitHub issue matched this combined failure. I searched for exact and related terms including `cannot reach the Paperclip control plane`, `execution_review_participant_recovery`, `sandbox callback bridge`, `review participant in_review`, and `control plane sandbox`. Related public issues: - Refs #8482 for `in_review` liveness invariant recovery. - Refs #863 for prior agent API-key reachability confusion. - Refs #248 for the broader sandboxed agent execution model. Bug summary: - What happened: a sandboxed local-adapter run could have host-side Paperclip access while the sandbox Bash/tool surface lacked a reachable API endpoint or usable run credentials. Separately, a reviewer run could finish while its execution-review stage remained pending, leaving the source issue in `in_review` with no decision and no live participant run. - Expected behavior: the mutation surface that agents actually use should receive a run-scoped Paperclip bridge, and pending review participants should get one bounded normal-model recovery wake before moving to explicit blocked/source-scoped recovery. - Steps to reproduce: run a sandbox-backed local adapter that needs Bash/curl/tooling to call Paperclip from inside the sandbox, or finish an execution-policy reviewer run without submitting the pending review decision. - Deployment mode: local/authenticated private development instance with sandbox-backed local adapters. ## What Changed - Changed sandbox callback bridge startup so bridge credentials are passed through the sandbox runner environment instead of embedded in the visible `nohup env ...` command string. - Added adapter-utils coverage proving the sandbox shell can call Paperclip through the bridge, forwards the host run JWT with `X-Paperclip-Run-Id`, and does not leak host or bridge tokens into stdout/stderr, runner command text, or runtime files. - Added one bounded execution-review participant recovery path for terminal reviewer runs whose `executionState` remains pending. - Escalated exhausted or non-invokable review participant recovery to blocked/source-scoped recovery with dedicated evidence, activity, and next-action text. - Documented the mutation-surface reachability contract in `doc/execution-semantics.md` and updated the Paperclip skill authentication guidance for sandbox bridge env vars. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism --maxWorkers=1` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - `curl -fsS $PAPERCLIP_API_URL/api/health` returned `status: ok` on the local instance. ## Risks - Medium behavioral risk: more `in_review` issues with terminal-but-pending reviewer runs will now be retried once and then blocked explicitly instead of remaining quiet. - Low sandbox bridge risk: credential delivery moved from command text to the runner environment, which is less leaky but depends on sandbox providers honoring the env payload for startup commands. - No database migration is included. - Full repo build and CI were not run locally before opening the PR; targeted server/adapter tests and typechecks passed. ## Model Used OpenAI GPT-5 via the Codex local agent, with repository tool use and shell-based code execution. The runtime did not expose a precise context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d68c34f2cc |
Fix managed workspace branch coherence recovery (#8826)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed issue workspaces are part of the control-plane runtime boundary: the server records which git worktree and branch an agent run is allowed to use. > - Existing reuse checks validated the worktree path and cleanliness, but did not fully validate that the actual checked-out branch still matched the recorded execution workspace branch. > - That gap let an agent run switch a managed worktree onto a publishing branch without updating the execution workspace record, then later reuse or finalize the workspace as though it were coherent. > - The runtime needs a bounded repair path for provably safe mismatches and a hard validation failure for dirty, divergent, or unrecorded branch transitions. > - This pull request adds branch coherence to managed git worktree validation, records explicit recovery evidence, and prevents finalize success when a run silently changes branches. > - The benefit is that branch drift becomes either safely repaired or visibly recoverable instead of silently corrupting managed workspace state. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Bug report: - What happened: a managed agent workspace could be recorded for one branch while the underlying git worktree was actually checked out on another branch. Reuse and finalization could still treat the workspace as healthy. - Expected behavior: managed git worktrees should verify the actual branch against the recorded execution workspace branch. Safe same-HEAD clean mismatches may be repaired, while dirty, divergent, or unrecorded branch transitions should fail into explicit workspace validation recovery. - Reproduction outline: create a runtime-managed issue worktree, switch its checkout to another branch without updating the execution workspace record, then attempt reuse or run finalization. - Deployment mode: local/self-hosted Paperclip server using managed git workspaces. - Related public work: Refs #7644 and #7579. Related but not duplicate: #8275 and #5851. ## What Changed - Added managed git worktree branch inspection, formatted validation evidence, and safe same-HEAD repair logic to the workspace runtime service. - Validated recorded managed workspace branch state before reuse and during heartbeat setup. - Added finalization-time branch guards so runs that silently switch branches fail with `workspace_validation_failed` instead of recording a successful finalize. - Added recovery fingerprints and evidence for `git_worktree_branch_incoherence`, including manual-repair next actions for unsafe branch drift. - Documented branch coherence as part of runtime-created git worktree workspace coherence. - Added focused tests for safe branch repair, dirty/divergent recovery evidence, heartbeat setup validation, and finalize failure/success paths. ## Verification - `pnpm install --frozen-lockfile` - `git diff --check origin/master...HEAD` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-recovery-actions.test.ts server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` Notes: - An initial full `pnpm test:run` attempt hit a transient `socket hang up` in one `plugin-routes-authz` case. The exact case passed when rerun directly, the full `plugin-routes-authz` file passed, and the subsequent full `pnpm test:run` passed. - `pnpm build` still emits existing Vite CSS pseudo-element and chunk-size warnings unrelated to this change. ## Risks - This intentionally changes behavior for managed runs that switch branches without recording the transition: they now fail during workspace validation/finalization instead of silently proceeding. - The automatic repair path is intentionally narrow. It only repairs clean branch mismatches when both branches point at the same commit; dirty or divergent worktrees require manual recovery. - Recovery fingerprints now include workspace-validation evidence, so duplicate recovery-action grouping is more precise for branch-incoherence failures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 Codex CLI/API coding agent, with shell/git/test execution and reasoning mode enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
a8f0ebaa80 |
Refresh run config before reusing workspaces (#8797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs are assembled by the heartbeat service from agent config, project workspaces, environment config, secret bindings, skills, and runtime session state. > - The heartbeat service intentionally reuses adapter sessions, execution workspaces, and sandbox leases when that preserves useful state. > - Reuse becomes incorrect when the effective next-run config changes after a saved session, workspace, or lease was created. > - Stale reuse can make a later run appear pinned to old agent, environment, secret, instruction, or workspace settings. > - This pull request records non-sensitive fingerprints for the effective session, workspace, and lease config at run boundaries. > - When those fingerprints drift, Paperclip refreshes persisted runtime config or starts fresh execution instead of reusing stale state. > - The benefit is predictable next-run config freshness without storing raw secret values, full env maps, provider credentials, or private path details. ## Linked Issues or Issue Description - Refs #8058 - Related PRs checked during dedup search: #4968, #4155, #84, #8480. These cover nearby workspace/session routing or model-config freshness areas, but do not duplicate this effective run config fingerprinting path. ## What Changed - Added effective run config fingerprinting for session, workspace, and lease reuse decisions, with canonicalization that ignores generated runtime noise and redacts sensitive values. - Updated heartbeat reuse logic to compare stored and next-run fingerprints, reset stale saved sessions, refresh persisted workspace config snapshots, replace stale reused workspaces when required, and avoid stale sandbox lease reuse. - Included plain environment value drift via value hashes, without storing the raw env values. - Root-bound instruction content hashing so legacy direct absolute instruction paths are represented but not read for config fingerprints. - Batched secret/version metadata lookups for environment lease fingerprinting. - Added workspace operation/run result freshness metadata so operators can inspect non-sensitive decision categories. - Surfaced config freshness labels and next-run copy in the UI and docs. - Added focused coverage for fingerprint redaction, session reset decisions, workspace refresh/replace behavior, environment lease drift, and persisted workspace restoration. ## Verification - `git diff --check` - Sensitive-data scan before push: - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "(AWS_ACCESS_KEY_ID|AWS_SECRET_ACCESS_KEY|ghp_[A-Za-z0-9_]{20,}|github_pat_[A-Za-z0-9_]{20,}|sk-[A-Za-z0-9]{20,}|-----BEGIN (RSA |OPENSSH |EC |DSA )?PRIVATE KEY-----|AKIA[0-9A-Z]{16})"` - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}"` - `pnpm exec vitest run server/src/__tests__/effective-run-config-fingerprints.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/environment-runtime.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm --filter @paperclipai/db clean` - `pnpm test:run` - `pnpm build` - UI screenshots from Cutter: - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-01.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-02.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-03.png ## Risks - Medium: overly broad fingerprints could start fresh sessions, workspaces, or sandbox leases more often than necessary. - Medium: missing a config category would allow stale reuse to persist for that category. - Medium: legacy direct absolute instruction paths are no longer content-hashed unless they are paired with an absolute managed instructions root. - Low data risk: fingerprint metadata stores hashes and category names, not raw secrets, raw env values, provider credentials, or private path details. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex CLI / Codex coding agent, tool-enabled with shell, Git, GitHub CLI, local test execution, and code editing. The exact deployed model variant and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
fd2f82ac5b |
[codex] Add built-in Hermes adapters (#8543)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters are the boundary between the control plane and the runtimes that actually do work. > - Hermes support needs to be available as first-class local and gateway adapters while still preserving the adapter-manager override path for external packages. > - The adapter work touches runtime execution, UI adapter metadata, onboarding prompts, scoped credentials, release packaging, and smoke coverage, so the handoff needs concrete verification rather than only unit tests. > - This pull request adds built-in Hermes local and Hermes gateway support, keeps external adapter overrides compatible, and documents/tests the gateway flow end to end. > - The benefit is that operators can hire Hermes-backed agents without a manual plugin install, while self-hosted installs can still override/shadow the built-ins through Adapter manager packages. ## Linked Issues or Issue Description No public GitHub issue exists for this exact Hermes built-in adapter, gateway onboarding, and release-source work. Problem description: - Hermes local and gateway adapters need a public, reviewable source path in the monorepo so package artifacts and built-in adapter behavior match the application source. - Operators need built-in `hermes_local` and `hermes_gateway` adapter choices without losing the ability to install external Hermes packages as overrides. - Gateway onboarding needs secure defaults for API server URLs, API keys, and generated agent setup text. - Hermes-originated task bridge credentials need narrower API-key scope configuration. - Related public PRs found during duplicate search include #3027, #2363, #7544, #7950, #8095, and #8543. ## What Changed - Added the unified Hermes adapter package with local and gateway server/UI/CLI exports, config schemas, transcript parsing, model detection, and package metadata. - Registered `hermes_local` and `hermes_gateway` as built-in adapters across shared constants, server registries, CLI packaging, and UI adapter registries. - Kept the external adapter override path compatible so installed Hermes packages can shadow built-ins and restore the built-in parser when disabled. - Added Hermes gateway onboarding docs, board-operator docs, Docker smoke assets, and shell smoke harnesses for join/e2e validation. - Added scoped task-bridge API-key support, authorization checks, issue-origin handling, and tests for Hermes-created Paperclip tasks. - Hardened gateway transport and redaction behavior for API keys, headers, session data, and smoke diagnostics. - Updated release packaging/bootstrap checks for the Hermes packages while leaving `pnpm-lock.yaml` out of the PR per repository policy. ## Verification Targeted local verification recorded before PR handoff: - `pnpm --filter @paperclipai/hermes-paperclip-adapter exec vitest run src/gateway/server/execute.test.ts` — 14/14 passed. - `pnpm test:hermes-gateway-smoke` — 6/6 passed. - Hermes package typecheck/build checks passed. - Focused server/UI adapter tests passed — 31/31. - Release helper Node tests passed — 18/18. - `git diff --check origin/master..HEAD` passed. Fresh Docker E2E smoke evidence: - Ran `pnpm smoke:hermes-gateway-e2e` on 2026-06-26 with a fresh state directory and fresh Docker container against a live Paperclip dev server. - Hermes direct execution reached `completed`. - Hermes stop/cancel path reached `cancelled`. - Hermes gateway created a Paperclip task, Paperclip ran the Hermes agent, and the task reached `done` with the expected marker response. - Temporary board auth keys, token files, smoke state, and Docker containers were cleaned up after the run. PR checks on head `b5eae40ce`: - GitHub Actions passed: `policy`, `review`, `Typecheck + Release Registry`, all general test shards, all serialized server shards, `Build`, `Canary Dry Run`, `e2e`, and aggregate `verify`. - External checks passed: Snyk and Socket Project Report. - External Socket Pull Request Alerts remained pending after the first-party CI matrix completed. ## Risks - Medium risk: this spans adapter registration, package publishing, gateway execution, onboarding docs, API-key scoping, and UI adapter metadata. - Migration risk is low: the scope-config migration adds a nullable column and does not rewrite existing keys. - Gateway execution depends on operator-provided Hermes API configuration; the smoke covers the Docker gateway path but real deployments may differ by network/auth setup. - Direct Greptile review on the latest expanded diff is file-count limited, although the commitperclip review gate passed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool use enabled in a local repository workspace. Context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Commitperclip review gate is green; direct Greptile review is file-count limited on the latest expanded diff - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed65d08d57 |
[codex] Gate skill mutations with skills:create permission (#8616)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents and board users operate inside a company-scoped control plane where permissions decide which mutating actions they can perform > - Company skills are part of the reusable agent-company setup surface, but skill mutation had been coupled to broader agent-creation authority > - That coupling meant importing or managing skills required a permission that also implies hiring power, which is broader than the operation needs > - Paperclip already has a grant-based permission vocabulary, so skill mutation should be authorized through a dedicated `skills:create` capability while preserving existing default behavior for trusted agents > - This pull request adds the skill creation permission contract, enforces it on company skill mutations, exposes it in agent permission management, and documents the changed CLI/API expectations > - The benefit is a narrower, auditable permission path for skill import/create/update/delete flows without forcing agents to receive broader agent-creation authority ## Linked Issues or Issue Description No public issue is linked. Problem: company skill mutation APIs were effectively tied to broader agent creation authority. This PR splits skill mutation authorization onto the public `skills:create` permission while keeping existing default skill creation behavior for agents unless explicitly disabled. Related public PR found during duplicate search: #5330. That PR uses an older `canManageSkills` shape; this PR implements the `skills:create` grant path instead. ## What Changed - Added `skills:create` to shared permission constants and agent permission types/validators as `canCreateSkills`. - Backfilled default human/member role grants for `skills:create`. - Updated company skill mutation routes to require board/user or agent access to `skills:create`, while preserving legacy/default agent behavior through `canCreateSkills` unless explicitly disabled. - Updated agent permission update handling, UI permission controls, duplicate-agent payloads, plugin SDK fixtures, and agent detail API surfaces for `canCreateSkills`. - Added regression coverage for skill route authorization, permission schema/default behavior, invite grants, omitted permission updates, and duplicate-agent payloads. - Updated CLI and Paperclip skill documentation for the new skill creation permission. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-permissions-service.test.ts server/src/__tests__/agent-permissions-routes.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/invite-join-grants.test.ts ui/src/lib/duplicate-agent-payload.test.ts` — 5 files, 90 tests passed. - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/server typecheck && pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm test:run ...changed files...` was attempted first, but the stable wrapper rejects explicit file arguments; direct Vitest was used for the same targeted files. ## Risks - Moderate authorization risk: this changes the gate for company skill mutations, so the tests cover board grant checks, agent explicit grant checks, legacy default allowance, and explicit denial. - Migration/backfill risk is low: the migration only grants `skills:create` to existing human roles that already need broad management capability. - UI/API compatibility risk is low: `canCreateSkills` remains default-on for full agent permissions, and the update validator preserves omitted values so unrelated permission edits do not re-enable disabled skill creation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with terminal/tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1951c80237 |
feat(cli): surface plugin install target host + add plugin target (#8575)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The CLI (`paperclipai plugin ...`) installs and manages plugins against a Paperclip server resolved from `--api-base` / `PAPERCLIP_API_URL` / the active profile / an inferred default > - During local plugin development you can have more than one Paperclip running (a released host plus a branch build on another port), and nothing told you *which* instance a command actually talked to > - So a plugin that depends on a route or response field only present on a feature branch could be silently installed/tested against a stale host, returning `API route not found`, and look broken when the real problem was the test target > - This pull request makes the install target explicit: it probes `GET /api/health` and prints the resolved API URL + server status/version/mode/exposure before installing, and adds a `plugin target` command plus docs for running and verifying against a branch service > - The benefit is that local plugin authors can confirm they are exercising the runtime they intend to, instead of debugging phantom plugin bugs caused by hitting the wrong server ## Linked Issues or Issue Description No public GitHub issue exists, so the underlying problem is described inline following the feature-request template. **Problem or motivation** Local plugin development assumes a single Paperclip on `http://127.0.0.1:3100`. When a plugin depends on server code that only exists on a feature branch (a new scoped route, a new response field, a new managed-resource capability), installing it into a long-lived host still on older code makes the route/field missing there. The plugin falls back or errors and *looks* broken, when the real cause is that it was tested against the wrong runtime. The CLI already let you point at any server, but it never surfaced which server you ended up on — so the mistake was invisible. **Proposed solution** Make the install target explicit. Before `plugin install` runs, probe `GET /api/health` and print the resolved API URL plus server status/version/deploymentMode/exposure, so the developer can confirm which Paperclip they are installing into. Add a standalone `plugin target` command to inspect the target without installing, a `--no-verify-target` escape hatch, and docs covering how to run a branch service on its own port and verify a branch route end-to-end. **Alternatives considered** - Do nothing and rely on the existing `--api-base` / `PAPERCLIP_API_URL` resolution — rejected because the gap was never the inability to point at a branch server, it was the lack of feedback about which server was actually hit. - Fail the install when the target looks stale — rejected as too aggressive; the probe is advisory and degrades gracefully when health details are not exposed or the server is unreachable. ## Dedup Search - [x] I searched the open and recently closed GitHub PRs for similar or duplicate PRs — this is not a duplicate ## What Changed - Add `probeTargetDiagnostics` / `formatTargetDiagnostics` helpers (`cli/src/commands/client/plugin.ts`) that read `GET /api/health` and report the resolved API URL plus server `status` / `version` / `deploymentMode` / `deploymentExposure`. - `plugin install` now prints these target diagnostics before installing, so you can confirm which instance you are installing into. Skippable with `--no-verify-target`. - `plugin install --json` keeps its original flat `PluginRecord` shape (top-level `id` / `pluginKey` / `version` / `status` are unchanged); when the target was probed it gains an additional top-level `target` field. Existing automation that reads the plugin fields keeps working. - Add a standalone `paperclipai plugin target` command to inspect the install target without installing anything. - Update `doc/plugins/LOCAL_PLUGIN_DEVELOPMENT.md`: how the CLI resolves its target, how to run a branch service on its own port and point the CLI at it explicitly, an end-to-end check that the branch route is actually served, and a troubleshooting entry for the stale-target symptom. - Unit tests for the diagnostics helpers (reachable + unreachable probe, and both render paths). ## Verification - `npx vitest run cli/src/__tests__/plugin-init.test.ts` — 10/10 pass (covers `probeTargetDiagnostics` success/failure and `formatTargetDiagnostics` rendering). - CLI typecheck (`tsc --noEmit` in `cli/`) — clean. - Manual: with a server running, `paperclipai plugin target` prints `Target Paperclip: <url>` and the health line; `plugin install` prints the same block before installing and `--no-verify-target` skips it. ## Risks Low risk. The probe is read-only (`GET /api/health`) and runs before install; if the server does not expose details it degrades to `ok (no details exposed)`, and an unreachable target prints a remediation hint rather than failing the command. The `--json` output keeps its original flat shape, so existing scripts are unaffected. No server or schema changes. ## Model Used Claude Opus 4.7 (`claude-opus-4-7`), extended thinking + tool use, via Claude Code. |
||
|
|
2dbaf4a7fa |
External object references across issue surfaces (#8512)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents, issues, approvals, comments, and work products. > - The involved subsystem is issue context: markdown links, issue properties, related work, lists, filters, inbox/sidebar status, and plugin-provided external context. > - The gap is that URLs to external systems currently remain mostly plain links, so humans and agents must manually open them to understand status, identity, and liveness. > - This matters because external work objects such as GitHub issues and pull requests are part of the operational state of a Paperclip company. > - The implementation keeps core provider-neutral: shared contracts, storage, sync, routes, and UI surfaces live in core while providers can contribute detection and status resolution. > - This pull request adds the external object reference foundation, GitHub provider support, issue-surface rendering, filters, sidebar/list/inbox signals, and test/story coverage. > - The benefit is that linked external work becomes inspectable Paperclip context without hardcoding every provider directly into the UI. ## Linked Issues or Issue Description No public GitHub issue exists for this work. Feature request: - Problem: URLs in Paperclip issues, comments, documents, and related surfaces do not expose provider status or object identity inline. - Proposed behavior: detect supported external object URLs, persist normalized references, refresh provider status, and render concise status-aware links across issue surfaces. - Users affected: board users, agents, and maintainers who triage issues containing external work links. - Acceptance: external object references are company-scoped, provider-extensible, visible in key issue surfaces, filterable where relevant, and covered by focused shared/server/UI tests. Related PR search: - No open duplicate PRs found for `external object references`. - Closed related prior attempt: #4556. ## What Changed - Added shared external-object contracts, validators, status/liveness helpers, and plugin protocol declarations. - Added database schema and additive migrations for external objects, source mentions, and display metadata. - Added server services/routes for detecting, syncing, summarizing, refreshing, and resolving external objects across issues, documents, comments, projects, and plugins. - Added a GitHub external-object provider plus plugin SDK authoring docs. - Wired UI presentation across markdown links, comments, issue chat, documents, properties, related work, issue rows, filters, inbox/sidebar badges, and Storybook stories. - Rebasing cleanup: moved the branch onto current `master`, repaired stale worktree provision config, hardened environment-sensitive tests/mocks, and removed committed screenshot artifacts from the PR branch to keep the reviewable file set below tool limits. ## Verification - `pnpm exec vitest run packages/shared/src/external-objects.test.ts server/src/__tests__/external-object-routes.test.ts server/src/__tests__/external-objects-service.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/lib/external-objects.test.ts` passed after rebasing: 5 files, 56 tests. - Historical branch verification before this PR creation included `pnpm test:run`, `pnpm -r typecheck`, and `pnpm build`; this PR body does not claim those were rerun after the final rebase. ## Risks - Medium: this adds a new cross-surface sync path on issue/document/comment writes. The implementation uses safe sync wrappers so external-object failures warn instead of blocking core mutations. - Medium: the migrations introduce new tables and indexes. They are additive and company-scoped. - Medium: provider-specific URL parsing can miss or misclassify edge cases. Shared canonicalization tests and provider tests cover current GitHub shapes. - Low: UI badge/filter behavior could add visual noise for object-heavy issues; component tests and Storybook stories cover the intended surfaces. > Roadmap checked: `ROADMAP.md` references the plugin system as the current extension path and does not list a duplicate core feature. Related long-range docs discuss external references, work products, preview URLs, and plugin extension points; this PR implements the scoped external-object reference foundation. ## Model Used OpenAI Codex, GPT-5 coding-agent runtime, with shell and GitHub CLI tool use. Reasoning mode: medium. Exact deployed runtime model ID and context window were not exposed in the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fce3b439af |
fix: warn operators that experimental features may break (#8382)
## Thinking Path > - Paperclip is the control plane operators use to manage AI-agent companies. > - Board operators rely on the settings UI and CLI docs to understand which product surfaces are stable to depend on. > - Experimental features already existed in the product, but the operator-facing contract around them was too soft and too fragmented. > - That created a risk that users would enable experiments without being told clearly that they can break, change, or disappear. > - The docs and the in-product settings page both needed the same explicit warning language so the contract is visible at the moment of decision. > - This pull request adds that warning to the board-operator guide, CLI references, and the experimental settings page. > - The benefit is clearer operator expectations without changing the underlying feature flags or rollout behavior. ## Linked Issues or Issue Description No public GitHub issue exists for this docs/polish gap. Problem description: - Board operators could enable experimental features without a clear operator-facing statement that those features are opt-in and come without compatibility guarantees. - The docs site, repo CLI reference, and in-product experimental settings page did not present one consistent warning contract. - This PR closes that gap by documenting the risk explicitly where operators discover and enable those settings. Related public search: - Searched public issues/PRs for related work with `gh search issues --repo paperclipai/paperclip 'experimental features warning'` and `gh search prs --repo paperclipai/paperclip 'experimental features warning'`. - Reviewed open PR #6165 during that search and found it unrelated; it changes experimental auth/routing flags rather than documenting experimental-feature risk. ## What Changed - Added a new board-operator guide at `docs/guides/board-operator/experimental-features.md` that defines the Paperclip contract for experimental features. - Registered that guide in `docs/docs.json` so it appears in the public docs navigation. - Added matching caveat language next to `instance settings:experimental` in `docs/cli/control-plane-commands.md`. - Added the same caveat to `doc/CLI.md` so the repo CLI reference does not drift from the published docs. - Added a single page-level warning banner to `ui/src/pages/InstanceExperimentalSettings.tsx` stating that experimental features are opt-in, carry no compatibility guarantees, and may change, break, or be removed. - Added a targeted UI test in `ui/src/pages/InstanceExperimentalSettings.test.tsx` that asserts exactly one page-level warning renders with the new risk language. ## Verification - `jq empty docs/docs.json` - `git diff --check` - `cd ui && pnpm vitest run src/pages/InstanceExperimentalSettings.test.tsx` - Manual review of the warning contract across: - `docs/guides/board-operator/experimental-features.md` - `docs/cli/control-plane-commands.md` - `doc/CLI.md` - `ui/src/pages/InstanceExperimentalSettings.tsx` UI note: - This is a copy-level warning addition rather than a layout rework. I did not attach before/after screenshots in this PR body. ## Risks - Low risk: this changes operator-facing documentation and warning copy, not feature-flag behavior. - The main failure mode is wording drift across docs and UI in future edits, which is why this PR adds the same contract to all relevant operator-facing surfaces. > I checked `ROADMAP.md` before opening this PR. This is docs/UI polish around an existing experimental surface, not overlapping roadmap-level core feature work. ## Model Used - OpenAI Codex Local using `gpt-5.4` with high reasoning and tool use for coordination, review, docs changes, and PR preparation. - Anthropic Claude Local using `claude-opus-4-8` with high reasoning and tool use for the in-product warning and targeted UI test. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
277a9a43d6 |
fix(recovery): convert review-parked continuations into dependency waits (#8371)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The productivity-recovery subsystem watches for "stranded" assigned issues — claims whose live run disappeared — and repairs, resumes, or visibly blocks them > - When an executor decomposes an umbrella issue into sub-tasks, it parks its own continuation as "waiting on review/approval" (error code `issue_continuation_waiting_on_review`) — a deliberate pause, not a lost run > - Recovery's staleness gate mistook that deliberate park for a disappeared run: it retried once, then escalated the issue to `blocked` with a recovery action and an operator-facing failure notice — even though nothing had failed and there was nothing for a human to do > - The user is left staring at an inscrutable, over-technical "stranded" error on a task they did nothing wrong with, with no idea what action to take > - This pull request teaches recovery to recognize a review-parked continuation and, when the issue has a real waiting target (open sub-tasks or unresolved blockers), convert it into a first-class dependency wait: `blocked`-by-children, original assignee kept, plus a plain-language comment saying it will resume automatically > - The benefit is that post-decomposition umbrellas sit on a real waiting path and self-resume through the normal blockers-resolved flow, while genuine strands (no waiting target) still escalate exactly as before ## Linked Issues or Issue Description Refs #6503 ## What Changed - `server/src/services/recovery/service.ts`: add `resolveContinuationWaitingOnReview`. When a continuation was cancelled with `issue_continuation_waiting_on_review` and the issue has a real waiting target — open (non-terminal) sub-tasks or existing unresolved blockers — recovery sets the issue `blocked` by those issues, keeps the original assignee, posts a plain-language `system` comment, and logs the activity. Wired into `reconcileStrandedAssignedIssues` ahead of the escalation path, with a new `waitingOnReviewResolved` counter on the result. - With no waiting target, the code falls through to the existing escalation, preserving genuine stranded-run detection. - `server/src/__tests__/heartbeat-process-recovery.test.ts`: two new tests — (1) a review-parked continuation converts into a dependency wait on its open sub-tasks (done children excluded, no recovery issue opened, plain-language comment, raw error code never leaks), and (2) it still escalates when no open dependency remains. - `doc/execution-semantics.md`: document the "Deliberate wait is not a lost run" recovery rule and the requirement that a post-decomposition umbrella hold a first-class waiting path rather than relying on `parentId` rollup. ## Verification - `cd server && npx tsc --noEmit` — passes against current `master`. - New tests in `server/src/__tests__/heartbeat-process-recovery.test.ts` (describe: "heartbeat orphaned process recovery"): - "converts a continuation parked for review into a dependency wait on its open sub-tasks" - "still escalates a continuation parked for review when no open dependency remains" - Run with the repo's vitest setup, e.g. `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts` (requires the embedded-postgres test harness). ## Risks Low. The change adds a single guarded pre-check ahead of the existing escalation path; behavior is unchanged when the cancellation error code is not `issue_continuation_waiting_on_review` or when the issue has no open sub-task / unresolved blocker to wait on. No schema or migration changes. ## Model Used Claude (Anthropic), Opus-class model, via the Claude Code agent harness — extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — server-only) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a71c4b6782 |
[codex] feat(watchdog): add task watchdog control plane (#8339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task lifecycle and recovery subsystems decide when agent work is still productive, stalled, or ready for review. > - Existing recovery paths can observe stopped or incomplete work, but there was no first-class per-task watchdog model with scoped review permissions. > - Watchdog follow-ups also need strict boundaries so recovery/status-only runs cannot mutate approvals or perform deliverable work. > - This pull request adds the task watchdog data model, API/service layer, scheduler/review flow, adapter wake context, UI configuration surfaces, and docs. > - The branch has been rebased onto current `paperclipai/paperclip` `master`; the watchdog migration is now ordered after master's latest migrations as `0104_issue_watchdogs`. > - The benefit is a more explicit task-review loop that preserves Paperclip's single-assignee and governance invariants while making stalled work easier to route. ## Linked Issues or Issue Description No linked GitHub issue. Paperclip task: [PAP-11275](/PAP/issues/PAP-11275). ## Problem or motivation Task recovery needs a first-class watchdog path that can inspect stopped work and create scoped follow-ups without bypassing normal task ownership. Board/UI users need a way to configure watchdogs on tasks and see watchdog-related live work. Recovery/status-only runs must remain limited to status reporting and must not create approvals, link approvals, or submit approval comments. ## Proposed solution Add a task-watchdog data model, scheduler/classifier, scoped mutation guard, adapter wake context, API/UI configuration surfaces, and documentation so watchdog agents can review stopped task subtrees under explicit boundaries. ## Alternatives considered Reuse the existing recovery-action flow only. That would keep stopped-work detection implicit, make per-task watchdog assignment harder to expose in the UI, and would not provide a durable scoped-review issue for stalled task trees. ## Roadmap alignment This is Paperclip control-plane lifecycle infrastructure for task execution and recovery. I checked `ROADMAP.md`; this PR does not duplicate an existing planned core item. ## What Changed - Added issue watchdog schema, migration, shared contracts, validators, CRUD API, and service support. - Added task watchdog scheduler/classifier behavior, scoped mutation enforcement, adapter wake context, and default watchdog mandate guidance. - Added UI surfaces for configuring watchdogs on new/existing tasks, viewing watchdog activity, and exposing the experimental setting. - Added docs for the user-facing task watchdog workflow and implementation semantics. - Gated new-task watchdog setup behind `enableTaskWatchdogs` and blocked cheap status-only recovery runs from approval mutations. - Rebased onto current `master` and renumbered the idempotent watchdog migration from the branch-local `0102_issue_watchdogs` slot to `0104_issue_watchdogs`. - Addressed Greptile feedback by loading watchdog classifier input with a recursive subtree query and centralizing the watchdog origin-kind constant. - Added and updated focused server/UI tests for watchdog routes, scheduler/classifier behavior, scope boundaries, live task visibility, settings, and new issue dialog behavior. ## Verification - `pnpm vitest run server/src/__tests__/task-watchdogs-scheduler.test.ts server/src/__tests__/task-watchdogs-classifier.test.ts` - `pnpm vitest run server/src/__tests__/approval-routes-idempotency.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts` - `pnpm vitest run ui/src/components/NewIssueDialog.test.tsx` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - Verified the PR diff does not include `pnpm-lock.yaml` or `.github/workflows`. ## Risks - Medium risk: this introduces a new task lifecycle surface touching DB schema, server routes/services, adapter wake context, and UI task configuration. - Watchdog scheduling behavior depends on the new experimental setting and runtime context checks behaving consistently across local and production agents. - The watchdog migration is idempotent (`IF NOT EXISTS` / duplicate-object guards) so users who tried the previous branch-local migration number should not get duplicate-object failures. - CI and the second Greptile pass are pending after the latest review-fix push. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in the Paperclip workspace. Exact runtime model id and context window were not exposed to the agent; tool use and local command execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots — N/A per Paperclip task instruction: do not add screenshots/images to this PR unless they are specifically part of the work. - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7069053a1f |
[codex] Add ask issue work mode (#8334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue work mode controls how a task starts and how the conversation composer frames the operator's intent. > - Paperclip already supports standard agent execution and planning mode, but there is no lightweight mode for asking a question without immediately implying execution or plan drafting. > - That gap makes low-commitment clarification workflows look like normal task execution. > - This pull request adds an explicit Ask mode and threads it through shared contracts, server heartbeat context, and the issue composer UI. > - The benefit is that operators can create or switch a task into a question-oriented mode while preserving existing agent and planning flows. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Inline feature request follows the repository feature request template. ### Subsystem affected Cross-cutting: `packages/shared`, `server/`, and `ui/`. ### Problem or motivation Issue conversations currently distinguish standard agent work from planning work, but question-first conversations do not have a clear public mode in the shared contract or UI. Operators who want to ask an agent a focused question have to use standard mode, which can imply normal task execution, or planning mode, which asks for a plan rather than an answer. ### Proposed solution Add Ask as a first-class issue work mode. It should be selectable from issue creation and issue chat, cycle alongside Standard and Planning from the keyboard shortcut/menu, appear distinctly in composer styling, and be included in heartbeat context so agents know to answer directly instead of executing or drafting a plan. ### Alternatives considered - Keep using standard mode for questions: rejected because it does not communicate answer-only intent to the agent or the UI. - Reuse planning mode for questions: rejected because planning mode asks for a plan and is semantically different from asking a question. - Add only local UI copy: rejected because the mode needs to be represented in the shared contract and server heartbeat context to be reliable. ### Roadmap alignment This is a focused issue-workflow improvement. `ROADMAP.md` was checked and no duplicate planned core work was found. ### Additional context Related public searches performed before opening this PR: - GitHub PR search for `"ask mode" repo:paperclipai/paperclip` - GitHub issue search for `"ask mode" repo:paperclipai/paperclip` - GitHub PR search for `"work mode" "ask" repo:paperclipai/paperclip` No duplicate PR was found. ## What Changed - Added `ask` to the shared issue work-mode contract and validation coverage. - Included issue work mode in heartbeat context summaries so agents can see standard, planning, and ask state. - Added Ask mode metadata, styling, composer tone handling, and selection/cycling behavior in the issue chat/new issue UI. - Updated focused tests for shared validators, heartbeat context, and affected UI work-mode flows. ## Verification - `NODE_ENV=test pnpm exec vitest run ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts` - `NODE_ENV=test pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/issues-service.test.ts ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts ui/src/pages/IssueDetail.test.tsx` The broader targeted command passed 8 test files / 245 tests. Visual reference for Standard/Planning/Ask composer states: https://gist.github.com/cryppadotta/714d8590bac55500a65e7e16de5bb4b8 It emitted an expected warning from an existing server test fixture about a missing run-log fixture while verifying derived issue comment metadata. ## Risks Low to moderate risk. This adds a new enum value that crosses shared, server, and UI contracts. Existing standard and planning modes are preserved, but any downstream code assuming only two non-terminal work modes may need to handle `ask`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent in Paperclip CodexCoder mode, with shell, git, GitHub connector, and local test execution tools. Context window and exact hosted model snapshot are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`, `feat/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc95699fde |
fix(server): enforce agent secret binding sync across lifecycle flows (#8307)
## Thinking Path > - Paperclip is the control plane people use to create, configure, and run AI agents for work. > - This change sits in the server-side agent lifecycle and secret-binding subsystem, where adapter config `env` entries can reference company secrets. > - An incident (while trying to configure a Novita sandbox) showed that an agent can reach a broken runtime state if `adapterConfig.env` contains `secret_ref` entries but the matching `company_secret_bindings` rows are missing. > - The immediate run-path guard and error-surfacing work made the failure diagnosable, but they did not fully prevent new broken agents from being created. > - The risk came from create and approval flows being responsible for remembering to sync bindings at each call site, which is easy to miss as new flows are added. > - This pull request moves the invariant into `agentService` create/update/activate paths, keeps the existing hire-flow fix, and adds regression coverage for create, update, and legacy pending-approval recovery. > - The benefit is that agent secret binding integrity is enforced closer to the data mutation point, so future callers inherit the protection automatically. ## Linked Issues or Issue Description Refs #8309 ### What happened? A Paperclip agent could persist `adapterConfig.env` `secret_ref` entries without matching agent-scoped `company_secret_bindings` rows. When that happened, the config UI could still look configured, but the real run path failed pre-dispatch because the secret was not actually bound to that agent. ### Expected behavior Every normal agent create, config-update, and pending-approval activation flow should leave the agent with secret bindings that match its persisted secret-ref env config. ### Steps to reproduce 1. Create or activate an agent through a flow that persists `adapterConfig.env` secret refs without synchronizing `company_secret_bindings`. 2. Observe that the config state can still appear populated. 3. Start a run for that agent. 4. Observe that pre-dispatch binding validation fails because the secret reference exists but the agent binding does not. ### Deployment mode Local dev (`pnpm dev`) ### Installation method Built from source (`pnpm dev` / `pnpm build`) ### Agent adapter(s) involved - Claude Code - Not adapter-specific (core bug) ### Database mode Embedded PGlite / embedded local dev database flow ### Access context Board (human operator) created or approved the agent; agent runtime later consumed the config. ### Additional context This PR focuses on preventing new broken states from normal service flows and on backfilling the covered legacy pending-approval activation path. ## What Changed - Kept the existing branch-local hire-flow fix that synchronized bindings for route and approval paths. - Moved the binding integrity invariant into `agentService.create()`, `agentService.update()` when `adapterConfig` changes, and `agentService.activatePendingApproval()`. - Added `server/src/__tests__/agents-service-secret-bindings.test.ts` covering create-time sync, update-time resync, and backfill for legacy pending-approval agents. - Removed now-redundant route-layer and approval-layer binding sync calls once the service layer became authoritative. - Simplified the affected unit tests so route/approval tests no longer assert service-owned binding writes directly. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/agents-service-secret-bindings.test.ts server/src/__tests__/approvals-service.test.ts server/src/__tests__/agent-skills-routes.test.ts` ## Risks - Low to medium risk. - This changes where secret-binding synchronization is enforced, so any unexpected caller that relied on upper-layer manual sync behavior could behave differently. - Agent create/update/activation flows now perform binding synchronization consistently, which adds binding-table writes at those mutation points. - This PR does not retroactively scan and heal every already-broken historical agent row; it prevents and backfills through the covered service flows. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex / GPT-5 Codex class model via `codex_local` - Session model family: GPT-5 Codex - Tool-assisted coding with shell, git, HTTP, and local test execution - Reasoning mode: medium interactive tool-use workflow ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5320a44088 |
Guard codex_local agents from shared OpenAI key (#8272)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The `codex_local` adapter runs local Codex CLI processes and builds their environment from persisted agent config plus host process env. > - A host-level `OPENAI_API_KEY` or shared Codex auth home can silently make new agents spend through shared credentials. > - Existing agents can be repaired manually, but new and updated agents need a persistent guard at the agent configuration boundary. > - This pull request isolates new and updated `codex_local` agents with per-agent `CODEX_HOME` and an empty `OPENAI_API_KEY` override. > - The benefit is that future agent creation or adapter updates cannot silently fall back to shared OpenAI credentials. ## Linked Issues or Issue Description Paperclip work item: [ZOL-5477](/ZOL/issues/ZOL-5477). No matching GitHub issue exists, so the bug is described inline following `.github/ISSUE_TEMPLATE/bug_report.yml`. **Pre-submission checklist** - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest `master` commit for this PR branch. - [x] I have confirmed the error originates in Paperclip's `codex_local` adapter configuration boundary, not in a provider outage. **What happened?** New or updated `codex_local` agents could inherit a host-level `OPENAI_API_KEY` or use a shared Codex home when their adapter config did not explicitly isolate those values. That made it possible for future agents or manual adapter edits to silently fall back to shared OpenAI credentials. **Expected behavior** Creating, hiring, or updating a `codex_local` agent should either persist isolated per-agent configuration or reject unsafe shared Codex home configuration with a clear 422 response. The guard must not print secret values. **Steps to reproduce** 1. Create or update a `codex_local` agent without an explicit `adapterConfig.env.OPENAI_API_KEY` override. 2. Run it on a host where the Paperclip server process has `OPENAI_API_KEY` set. 3. Observe that the adapter process can inherit the host key unless Paperclip persists a blocking empty override. 4. Set `adapterConfig.env.CODEX_HOME` to a shared path such as `~/.codex` or the company-level `codex-home`. 5. Observe that the old code allowed the shared auth home instead of returning a validation error. **Paperclip version or commit** - Reproduced by inspection against `master` before this PR. **Deployment mode** - Local dev / self-hosted server with `codex_local` agents. **Installation method** - Built from source. **Agent adapter(s) involved** - Codex. **Database mode** - Not database-related. **Access context** - Board and agent configuration paths. **Relevant logs or output** - No secret-bearing logs included. **Relevant config** - Unsafe shape: missing `adapterConfig.env.OPENAI_API_KEY`, or shared `adapterConfig.env.CODEX_HOME`. - Fixed shape: per-agent `CODEX_HOME` plus empty `OPENAI_API_KEY` override. **Additional context** Related PR search for `codex_local OPENAI_API_KEY CODEX_HOME` found: - #3681 `fix: preserve managed Codex auth and repo-root env loading` - #5621 `fix: copy worktree codex auth locally` Those are adjacent auth-handling changes, but they do not add the agent create/update guard implemented here. **Privacy checklist** - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added a `codex_local` config guard in agent create, hire, and update routes. - The guard assigns `adapterConfig.env.CODEX_HOME` to `companies/<companyId>/agents/<agentId>/codex-home` when missing. - The guard persists `adapterConfig.env.OPENAI_API_KEY = ""` when missing, preventing host env inheritance. - Shared `CODEX_HOME` values for the company codex-home, host `$CODEX_HOME`, or `~/.codex` now fail with a 422 error. - Added route tests for create, hire, update, and rejected shared host Codex home. - Updated `codex_local` and development docs to describe the per-agent home contract. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/agent-skills-routes.test.ts` - `pnpm typecheck` - `git diff --check upstream/master...HEAD` - `gh pr list --repo paperclipai/paperclip --state all --search "codex_local OPENAI_API_KEY CODEX_HOME" --limit 20 --json number,title,state,url` - `rg -n "codex|OPENAI_API_KEY|CODEX_HOME|adapter" ROADMAP.md` returned no roadmap overlap. ## Risks - Existing legacy `codex_local` agents with shared `CODEX_HOME` will get a clear 422 when their adapter config is updated until the shared path is replaced. This is intentional because silent fallback is the bug being guarded. - Low migration risk: no database migration and no secret values are printed or persisted beyond the empty override. ## Model Used - OpenAI GPT-5.5 Codex, Codex coding-agent session with repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Paperclip - Issue: [ZOL-5477](/ZOL/issues/ZOL-5477) - Owner: Разработчик (`6625498c-66c9-429f-b578-4463ddc3ba16`) - Status: waiting reviewer - Next action: merge after approval and green CI --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d47b4da655 |
Auto-build bundled plugins on install (#8254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Plugins extend the server with worker/UI surfaces, and bundled local plugins under `packages/plugins/**` ship as TS source — their compiled `dist/` is not checked in > - On a fresh checkout, installing a bundled local plugin via the in-app **Install** button failed because `paperclipPlugin.manifest` points at `./dist/manifest.js`, which does not exist until the package is built > - The error surfaces as `Package … does not appear to be a Paperclip plugin (no manifest found)`, which is misleading — the manifest is real, the dist is just missing — and forces every contributor to run `pnpm --filter … build` by hand before the bundled-plugin installer works at all > - This pull request teaches the install path to detect that case and run the package's build (plus standalone runtime bootstrap for plugins outside the root workspace) before manifest resolution, gated by a kill switch and a bounded timeout > - The benefit is bundled plugins like `@paperclipai/plugin-workspace-diff` install in one click on a fresh checkout, with a clear error message and manual fallback when the autobuild itself fails ## Linked Issues or Issue Description No existing GitHub issue. Underlying bug, following the bug-report template: **What happened?** Installing a bundled local plugin from a fresh checkout fails with `Package @paperclipai/plugin-workspace-diff at packages/plugins/plugin-workspace-diff does not appear to be a Paperclip plugin (no manifest found)`. The manifest is declared in `package.json` (`paperclipPlugin.manifest = ./dist/manifest.js`) but `dist/` is not built/committed, so the loader cannot find it. **Expected behavior** Clicking **Install** on a bundled plugin builds it if needed and registers it, without a manual build step. **Steps to reproduce** 1. Fresh checkout of `master` 2. Start the server, open Plugin Manager 3. Click **Install** next to `@paperclipai/plugin-workspace-diff` 4. Observe the "no manifest found" failure **Scope** Same failure mode affects every bundled plugin without a checked-in `dist/` (`plugin-llm-wiki`, examples, sandbox-provider plugins, etc.). ## What Changed - `server/src/services/plugin-loader.ts`: added `ensureLocalPluginBuilt(packageRoot, pkgJson)` — when the package lives under `packages/plugins/**` and its declared paperclipPlugin entrypoints (`manifest`, `worker`, `ui`) are missing, run `pnpm --filter <name> build` (and a standalone runtime-deps bootstrap for plugins outside the root pnpm workspace) before manifest resolution - `server/src/routes/plugins.ts`: invoke the autobuild from the local-path install path; surface a `hasBuiltEntrypoints` boolean on the `AvailableBundledPlugin` listing; invalidate the bundled-plugins cache after a successful install so a freshly built plugin no longer reports `hasBuiltEntrypoints: false` - `ui/src/api/plugins.ts` + `ui/src/pages/PluginManager.tsx`: type and consume `hasBuiltEntrypoints` so the installer can show that an autobuild will run on install - `server/src/__tests__/plugin-install-autobuild.test.ts`: new suite — 9 tests covering success, kill-switch, build failure, timeout, manifest still missing after build, standalone variant, and the existing `plugin-routes-authz` listing assertion - `doc/plugins/LOCAL_PLUGIN_DEVELOPMENT.md`: documents the autobuild, the `PAPERCLIP_DISABLE_PLUGIN_AUTOBUILD=1` kill switch, and the manual fallback command - Detect the autobuild timeout via the child-process `killed` flag rather than string-matching the error message, so the "after timing out" context is actually emitted Knobs: - `PAPERCLIP_DISABLE_PLUGIN_AUTOBUILD=1` — skip autobuild entirely; restore prior behavior - Build timeout: 120s, with a clear error that points at the manual `pnpm --filter <name> build` recovery command ## Verification - `cd server && pnpm vitest run src/__tests__/plugin-install-autobuild.test.ts src/__tests__/plugin-routes-authz.test.ts` → 44/44 pass - End-to-end on a clean checkout: `rm -rf packages/plugins/plugin-workspace-diff/dist`, invoke `ensureLocalPluginBuilt()` against the real package, all declared entrypoints (`dist/manifest.js`, `dist/worker.js`, `dist/ui/index.js`) regenerated. The original `no manifest found` symptom no longer reproduces. ## Risks Low. The autobuild only fires when (a) the package sits under `packages/plugins/**`, (b) at least one declared entrypoint is missing, and (c) the kill switch is not set. In a packaged production server the `packages/plugins/**` path does not exist on disk, so the helper short-circuits and never shells out to `pnpm`. Failures from the spawned build are surfaced as an install error with the exact manual command to retry, so the worst-case is the same UX as before plus a clearer message. ## Model Used Claude Opus 4.7 (claude-opus-4-7), extended thinking enabled, tool use (filesystem + bash). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
362c30ccdc |
feat(server): opt-in OpenTelemetry auto-instrumentation (#3735)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Production self-hosters increasingly expect telemetry out of the box — Jaeger, Tempo, Honeycomb, Datadog, Grafana Cloud, Dynatrace all speak OTLP > - Today there is no OpenTelemetry bootstrap in the server, so operators who want traces have to patch their fork or run a sidecar that captures only HTTP-level info > - An opt-in bootstrap that costs nothing when disabled is the minimum-viable surface for this audience > - The OpenTelemetry packages are heavyweight enough that we don't want them in the default dependency graph — they should load only when the operator configures an OTLP endpoint > - This pull request adds a self-contained `server/src/instrumentation.ts` that dynamically imports the OTel SDK and starts it when `OTEL_EXPORTER_OTLP_ENDPOINT` is set, and is a complete no-op otherwise ## Linked Issues or Issue Description No existing issue covers this directly — feature-gap description following the feature-request template: **Problem or motivation** Production self-hosters increasingly expect telemetry out of the box — Jaeger, Tempo, Honeycomb, Datadog, Grafana Cloud, Dynatrace all speak OTLP — but the server has no OpenTelemetry bootstrap. Operators who want traces today must patch their fork or run a sidecar that captures only HTTP-level information. **Proposed solution** An opt-in OTel bootstrap gated on `OTEL_EXPORTER_OTLP_ENDPOINT`, loaded via dynamic `import()` only when configured, so the heavyweight OTel packages stay out of the default dependency graph. **Alternatives considered** Related open PRs found during the duplicate-PR search approach observability differently: #4894 adds OTLP instrumentation to Paperclip core unconditionally, and #3752 proposes an observability plugin. Not duplicates — different layering: this PR keeps the default install dependency-free via opt-in dynamic import. ## What Changed - New `server/src/instrumentation.ts` — opt-in OpenTelemetry auto-instrumentation. Gated on `OTEL_EXPORTER_OTLP_ENDPOINT`. Respects the standard OTel env vars (`OTEL_SERVICE_NAME`, `OTEL_SERVICE_VERSION`, `OTEL_EXPORTER_OTLP_ENDPOINT`). Skips the fs/dns/net auto-instrumentations (too chatty). `sdk.start()` is wrapped in try/catch so a bad endpoint or missing native bindings doesn't crash the server. `process.once("SIGTERM" / "SIGINT", …)` for clean shutdown on the first signal only. OTel packages are loaded via dynamic `import()` so they are true optional runtime dependencies — no entries in `package.json`, no lockfile churn. - `server/src/index.ts` — import `./instrumentation.js` as the very first statement so auto-instrumentation can patch `http` / `express` / `pg` before they are evaluated by downstream modules. ## Verification - `OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 pnpm start` after `pnpm install @opentelemetry/{sdk-node,auto-instrumentations-node,exporter-trace-otlp-grpc,resources,semantic-conventions}` in `server/` — traces show up in the configured collector; HTTP, Express, and Postgres spans are populated. - `OTEL_EXPORTER_OTLP_ENDPOINT` unset — server starts with no OTel-shaped output in logs, no behavior change. - `OTEL_EXPORTER_OTLP_ENDPOINT=…` set but packages not installed — single `console.warn` at startup telling the operator which packages to install. ## Risks Low. No behavior change unless the env var is set. The bootstrap never throws into the caller; every failure path ends in `console.warn` / `console.error` and falls through to non-traced operation. ## Model Used Claude Opus 4.6 (1M context), extended thinking mode. ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] Thinking path traces from project context to this change - [x] Model used specified - [x] Tests run locally and pass - [x] CI green - [x] Greptile review addressed --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
d7049e0cae |
fix(server): adopt stale checkout run ownership (#5413)
## Thinking Path > - Paperclip is a control plane for autonomous AI-agent companies. > - Issue checkout ownership is part of the execution-control layer that prevents two runs from mutating the same task at the same time. > - The current lock model should preserve `409` conflicts for live competing owners, but it should not strand the rightful assignee behind a stale terminal run. > - A same-agent follow-up run can encounter an existing `checkoutRunId` from a failed, timed-out, succeeded, or missing heartbeat run. > - In that case, the new run should safely adopt ownership instead of failing with an ownership conflict. > - This pull request makes stale checkout adoption transactional and keeps live checkout owners protected. > - The benefit is safer run recovery without weakening single-owner checkout semantics. ## Linked Issues or Issue Description - Fixes #5350 - Closes #1508 - Closes #1970 - Closes #2083 - Closes #3158 - Closes #3190 - Related stale-lock PRs reviewed during dedup search: #7536, #6658, #5660, #5442, #6223, #7048, #6824, #6799 ## What Changed - Updated issue checkout ownership recovery so the current assignee can adopt a stale terminal or missing checkout run. - Added row locking around stale checkout adoption to avoid races while replacing `checkoutRunId` / `executionRunId`. - Preserved `409` behavior when a different live checkout owner is still active. - Prevented terminal actor runs from reclaiming an unowned checkout lock after the newer eager stale-checkout clear path. - Fixed the stale checkout test fixture so same-assignee cases do not insert duplicate agent rows. - Added/kept focused coverage for stale checkout adoption and live-owner conflict behavior. - Fixes #5350. ## Verification - Focused tests: ```sh pnpm exec vitest run server/src/__tests__/issues-service.test.ts server/src/__tests__/issue-stale-execution-lock-routes.test.ts ``` Result: ```text 2 passed, 84 tests passed ``` - Server typecheck: ```sh pnpm --filter @paperclipai/server typecheck ``` Result: ```text passed ``` - Live curl smoke confirmed same-agent stale checkout adoption returns `200` instead of `409`. ```text old_run_status=succeeded checkout_http=200 patch_http=200 ``` The PATCH response showed `checkoutRunId` and `executionRunId` updated to the new run id. ### Live curl smoke result <img width="1498" height="570" alt="Live curl smoke showing stale checkout adoption returned 200" src="https://github.com/user-attachments/assets/4bf834de-e3cd-4495-ac5a-74767b439eeb" /> ### Server request log <img width="631" height="131" alt="Server logs showing heartbeat, checkout, and patch requests succeeded" src="https://github.com/user-attachments/assets/ceaaa403-110e-44e8-bac8-5d8506e79cc3" /> ## Risks - Low to medium risk: this touches issue execution lock ownership. - The behavioral shift is intentionally narrow: only the current assignee can adopt stale terminal or missing checkout ownership. - Live checkout owners remain protected with `409`. - No database migration or API contract change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.5 Codex coding agent with repository tool use, shell execution, code review, and local verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Cross-references and status (maintainer) - Closes #1508 - Closes #1970 - Closes #2083 - Closes #3158 - Closes #3190 - Status: rebased onto current master; focused tests and server typecheck pass locally; all required CI is green; Greptile is 5/5; master drift verified. --------- Co-authored-by: Devin Foley <devin@paperclip.ing> |
||
|
|
f3db7b88ea |
Clear stale checkoutRunId on run finalization and add backstop sweeper (#6008)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue subsystem holds per-row lock columns (`checkoutRunId`, `executionRunId`, `executionAgentNameKey`, `executionLockedAt`) that gate checkout, ownership, and release > - When a heartbeat run terminates, `releaseIssueExecutionAndPromote` clears the execution-lock columns but stale checkout locks could remain attached to dead runs in edge paths > - The original fix closed the finalization, checkout, release, and sweeper paths, but PR CI exposed one more process-loss retry path where a queued retry advanced `executionRunId` while leaving `checkoutRunId` pinned to the failed run > - This pull request closes the asymmetry: terminal-run cleanup and process-loss retry recovery release dead checkout locks while preserving live execution ownership > - The benefit is permanent, automatic self-heal of stale lock columns and fewer false checkout 409s requiring board intervention > - Related upstream issue: #6007 ## Linked Issues or Issue Description Refs #6007. Duplicate/related PR search performed on 2026-06-10 with query `checkoutRunId process loss retry stale checkout lock repo:paperclipai/paperclip`. Related PRs found and reviewed for overlap: - #7727 `fix(heartbeat): atomically advance checkoutRunId on process-loss retry` - #7707 `test: cover same-agent stale checkout adoption` - #3068 `fix: clear checkoutRunId when releasing issue execution lock` ## What Changed - `server/src/services/heartbeat.ts` `releaseIssueExecutionAndPromote`: extend the per-issue update to also null `checkoutRunId` when it matches the terminating run id. WHERE clause scoped to `executionRunId = run.id OR checkoutRunId = run.id` for idempotence. - `server/src/services/heartbeat.ts` process-loss retry: when queuing the retry run, move `executionRunId` to the retry and clear the failed run's `checkoutRunId` so the dead run no longer owns checkout. - `server/src/services/issues.ts`: add `clearCheckoutRunIfTerminal` helper, symmetric to `clearExecutionRunIfTerminal`. No assignee/status precondition. Wired into `checkout`, `assertCheckoutOwner`, and `release`. Exported on the issue service. - `server/src/services/recovery/service.ts`: add `sweepStaleIssueLocks`. Scans `issues` where `checkoutRunId IS NOT NULL OR executionRunId IS NOT NULL`, joins each referenced run, and clears all lock columns on issues whose referenced runs are all terminal or missing. Emits one `issue.stale_lock_cleared` activity log row per cleared issue. - `server/src/services/heartbeat.ts`: re-export the sweeper on the heartbeat facade. - `server/src/index.ts`: invoke `sweepStaleIssueLocks` in both the startup recovery sequence and the periodic heartbeat timer chain. - Tests: route-level coverage of the new self-heal path on the next checkout attempt, service-level sweeper coverage, and heartbeat recovery assertions that terminal process-loss cleanup releases `checkoutRunId`. ## Verification ```bash pnpm --filter @paperclipai/server typecheck pnpm --filter @paperclipai/server exec vitest run \ src/__tests__/recovery-stale-issue-lock-sweep.test.ts \ src/__tests__/issue-stale-execution-lock-routes.test.ts NODE_ENV=test pnpm exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t "queues exactly one retry when the recorded local pid is dead|does not block paused-tree work when immediate continuation recovery is suppressed by the hold" NODE_ENV=test pnpm exec vitest run src/__tests__/heartbeat-process-recovery.test.ts ``` All listed local checks pass. The new and updated tests cover: - Run termination clears `checkoutRunId` when it points at the terminating run. - Process-loss retry clears the failed run's `checkoutRunId` while assigning `executionRunId` to the queued retry. - A different agent calling `POST /api/issues/:id/checkout` on an issue whose prior owner died self-heals via `clearCheckoutRunIfTerminal` and succeeds. - Sweeper clears stale lock columns for issues whose run row is terminal. - Sweeper leaves issues alone while the referenced run is still running. - Sweeper leaves issues alone when `executionRunId` is still running even if `checkoutRunId` is terminal. - Sweeper is idempotent; second pass clears nothing. Manual reproduction of the original bug shape: 1. Create an issue assigned to agent A, set `status='in_progress'`, `checkoutRunId=R1`, `executionRunId=null`, where `heartbeat_runs.status = 'failed'` for `R1`. 2. Reassign to agent B and move to `status='todo'`. 3. Before this PR: agent B `POST /checkout` returns `409 Issue checkout conflict` indefinitely. After this PR: succeeds, lock columns rewritten to agent B's current run id. ## Risks - Low. All clears are scoped by run id, so they only fire when the lock column unambiguously points at the terminating or terminal run. No schema change. No migration. No API surface change. - Behavioral shift: an issue that previously stayed `in_progress` with a dead `checkoutRunId` after run termination now self-heals. Downstream code that reads stale `checkoutRunId` as a proxy for recent run history should already be reading `executionRunId` or the `heartbeat_runs` table. - Sweeper cost: one indexed scan per recovery tick over rows where `checkoutRunId IS NOT NULL OR executionRunId IS NOT NULL` plus a single batched `heartbeatRuns` lookup per candidate. Negligible at expected cardinality; further bounded by the existing recovery cadence. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. This is a bug fix, not a feature. No roadmap overlap. ## Model Used - Claude (Anthropic), model ID `claude-opus-4-7`, extended-thinking off, tool use enabled. - OpenAI Codex, GPT-5-based coding agent, tool use enabled, used for the follow-up process-loss retry fix and PR body update. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dotta <bippadotta@protonmail.com> |
||
|
|
67b22d872f |
[codex] Clarify interrupt handoffs and scoped wake semantics (#7855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread is the operator surface where comments, assignee changes, pauses, resumes, and wakeups turn human intent into agent execution. > - Interrupting a live run and handing work to another assignee needs clear semantics so the product does not accidentally keep work alive, wake the wrong participant, or hide why an agent stopped. > - Comment-driven wakes also need strict boundaries so closed, blocked, and dependency-driven work only resumes when there is real actionable input. > - This pull request codifies the interrupt handoff contract, implements backend scheduling behavior, and gives the UI clearer handoff/pause language. > - The benefit is a more inspectable and predictable task lifecycle for both operators and agents. ## Linked Issues or Issue Description Paperclip issue: `PAP-10664` / `PAP-10751`. Problem: interrupting or reassigning live agent work could be ambiguous in the UI and backend. Operators needed clearer feedback about whether a handoff wakes an agent, what pause/cancel affects, and when comments should revive execution. The backend also needed stronger tests around comment wake boundaries, retry supersession, and structured agent mention dispatch. Related GitHub PR search found broad workflow-adjacent PRs #5082, #6359, and #4083, but no exact duplicate for this head branch or interrupt-handoff scope. ## What Changed - Added an interrupt handoff semantics document covering destination behavior, wake expectations, and live-run interruption states. - Implemented backend interrupt handoff behavior and comment wake/reopen handling in issue routes/services and heartbeat scheduling. - Hardened structured agent mention dispatch so mentions resolve through the intended dispatch path. - Added UI helpers and components for handoff chips, wake rows, interrupt banners, pause-affects summaries, and composer guidance. - Updated the issue properties assignee picker and issue chat/composer surfaces to make interrupt/reassign behavior clearer. - Added backend, UI utility, component, and Storybook coverage for the new behavior. - Stabilized the new UI component tests with a local `flushSync`-backed act helper matching existing repo practice in this dependency set. - Addressed Greptile feedback by threading historical run `errorCode` through issue-run data and operator-interrupted chat labels. - Addressed Greptile's cancel ordering concern by terminating/deleting in-memory heartbeat processes before cancellation status persistence, with regression coverage for DB update failure. ## Verification - `git diff --check $(git merge-base HEAD origin/master)..HEAD` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/interrupt-handoff.test.ts src/lib/issue-chat-messages.test.ts src/components/IssueProperties.test.tsx src/components/interrupt-handoff/InterruptHandoffViews.test.tsx --no-file-parallelism --maxWorkers=1` — 4 files / 91 tests passed before the Greptile follow-ups. - `pnpm run preflight:workspace-links && pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-retry-scheduling.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts server/src/__tests__/issue-tree-control-service.test.ts server/src/__tests__/issue-update-comment-wakeup-routes.test.ts server/src/__tests__/issues-service.test.ts --no-file-parallelism --maxWorkers=1` — 6 files / 191 tests passed before the Greptile follow-ups. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts --no-file-parallelism --maxWorkers=1` — 1 file / 24 tests passed after the historical `errorCode` follow-up. - `pnpm exec vitest run server/src/__tests__/activity-service.test.ts server/src/__tests__/activity-routes.test.ts --no-file-parallelism --maxWorkers=1` — 2 files / 11 tests passed after the historical `errorCode` follow-up. - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism --maxWorkers=1` — 1 file / 52 tests passed after the cancel ordering follow-up. - Greptile is green for head `272647636287d034bab8d981eaf5305865aa0f96`; the old inline P2 is resolved/outdated. - GitHub Actions, Socket, security-review, and Greptile checks are green for head `272647636287d034bab8d981eaf5305865aa0f96`. The external `security/snyk (cryppadotta)` status was still pending at `https://app.snyk.io/org/cryppadotta/pr-checks/85b3e8f4-04e1-4f8e-9362-899c8148c23c` after a bounded wait. ## Risks - Medium: changes touch issue comments, wake scheduling, and live-run interruption semantics, so regressions could affect when agents resume or stay stopped. - Medium: UI copy and state grouping for assignee changes may need reviewer tuning after product review. - Low migration risk: no database schema migration is included. - The branch was created before the latest `origin/master` commits; reviewers should confirm CI merge-base behavior and resolve any merge conflicts if GitHub reports them. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent, tool use and local command execution enabled. Exact hosted model build and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Screenshot note: this PR includes Storybook coverage for the new interrupt handoff UI states rather than captured before/after browser screenshots in this PR-creation heartbeat. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
468edd8b22 |
Add workspace file viewer and artifact links (#7681)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent work is issue-centered, and reviewers often need to inspect files, artifacts, and path references produced during that work. > - Before this branch, workspace-relative paths and artifact file references were not first-class inspectable objects in the board UI. > - Safe file viewing needs shared resource contracts, server-side workspace boundary checks, and UI that opens files without exposing arbitrary host paths. > - The workspace file viewer branch needed to stay as one active PR and be rebased onto current `paperclipai/paperclip:master` for review. > - This pull request adds the workspace file resource API, issue-page file viewer and browser, markdown file-reference links, and artifact file chips. > - The benefit is that board users can inspect relevant files from issue context while preserving workspace boundaries and auditability. ## Linked Issues or Issue Description No public GitHub issue exists for this branch. Internal Paperclip issues: `PAP-1953`, `PAP-10539`, `PAP-10733`. Problem / motivation: - Board users need to open workspace-relative files mentioned by agents or attached as work-product metadata without switching to a terminal. - The UI needs to support both direct file-path opening and workspace browsing/searching from an issue page. - The server must enforce company access, workspace boundaries, size limits, rate limits, and safe audit logging. Related PR: - Prior closed attempt: #4442 - Single active PR for this branch: #7681 ## What Changed - Added shared workspace file resource types, validators, and workspace-file `resourceRef` metadata validation for work products. - Added server routes/services for resolving, listing, and previewing workspace-relative files with access checks, scan caps, list-specific limits, and audit logging. - Added the issue file viewer provider, sheet, workspace browser, command-palette action, markdown workspace-file autolinks, and artifact file chips. - Updated issue workspace UI and stories/tests for file browsing and workspace file opening. - Rebased the branch onto current `paperclipai/paperclip:master` and updated the existing single PR branch. - Addressed current-head Greptile follow-ups by applying `offset` consistently across search/recent/changed file listings, restoring stopped-service port ownership checks before auto-port reuse, and stabilizing the workspace browser pagination test. ## Verification Current local verification after rebase to `public/master`: - `pnpm exec vitest run packages/shared/src/work-product.test.ts server/src/__tests__/file-resources.test.ts server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/instance-settings-service.test.ts server/src/__tests__/workspace-runtime.test.ts ui/src/components/FileViewerSheet.test.tsx ui/src/components/FileViewerSheet.copy.test.tsx ui/src/components/WorkspaceFileBrowser.test.tsx ui/src/components/WorkspaceFileMarkdownBody.test.tsx ui/src/context/FileViewerContext.test.ts ui/src/lib/remark-workspace-file-refs.test.ts ui/src/lib/workspace-file-parser.test.ts ui/src/components/IssueWorkspaceCard.test.tsx` - 13 files passed, 197 tests passed. - `pnpm -r --filter @paperclipai/shared --filter @paperclipai/server --filter @paperclipai/ui typecheck` - passed. - `pnpm exec vitest run ui/src/components/WorkspaceFileBrowser.test.tsx` - 1 file passed, 25 tests passed. - `pnpm exec vitest run server/src/__tests__/file-resources.test.ts server/src/__tests__/workspace-runtime.test.ts` - 2 files passed, 90 tests passed. - `pnpm -r --filter @paperclipai/server typecheck` - passed. - Confirmed branch is `0` behind and `46` ahead of current `public/master` after rebase and follow-up commits. - Confirmed the PR diff does not include `pnpm-lock.yaml`. - Confirmed the PR diff does not include `.github/workflows` changes. - Searched GitHub for duplicate or related workspace file viewer PRs/issues; #4442 is the prior closed attempt and this PR is the single active PR for the branch. - No screenshots were committed; the task explicitly asked not to add design screenshots or images unless they were part of the work. Current remote verification on head `a698a7bc10137baf7d25bd5722e1d6e0343387c1`: - Greptile Review - success, 64 files reviewed, 0 comments added, no unresolved Greptile review threads. - PR workflow `verify` - success. - Typecheck + Release Registry, General tests, workspace test shards, serialized server suites, Build, Canary Dry Run, e2e, Socket, and Snyk - success. - `security-review` - neutral, with output saying a draft advisory was filed for maintainer review and is not a merge block. - `commitperclip PR Review / review` - cancelled after the security gate detected flags and timed out while creating/reviewing the advisory. I reran it once and it cancelled the same way; no actionable code/test failure was exposed in the job logs. ## Risks - This is a broad UI/server feature PR, so review needs to pay attention to route authorization, workspace boundary handling, and markdown autolink false positives. - Workspace browsing intentionally caps list results and scan depth; very large workspaces may require users to refine search terms. - Remote workspace preview remains unavailable until remote file-access support is implemented. - The neutral commitperclip security-review advisory needs maintainer review, but the check output says it is not a merge block. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent in a Paperclip/Codex local tool-use environment, medium reasoning, with shell/GitHub CLI tool use for branch inspection, verification, rebase, PR update, Greptile review, and CI inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
50bff3b274 |
feat(ui): add collapsible sidebar rail and takeover panes (#7824)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents, work, and company context. > - The board UI sidebar is the main way operators keep orientation across companies, projects, agents, issues, and settings. > - The existing fixed expanded sidebar competes with route-specific navigation, especially company settings and plugin routes that bring their own contextual sidebar. > - A collapsible primary rail preserves global navigation while giving contextual pages more horizontal room. > - This pull request adds a persisted collapsed rail, hover/focus peek, keyboard toggle, and a secondary sidebar takeover model for settings and plugin `routeSidebar` surfaces. > - The benefit is a denser board shell that keeps the app rail available without replacing it when a route needs its own navigation. ## Linked Issues or Issue Description Paperclip issue: PAP-10638 Create collapsible sidebar branch. Related GitHub PR found during duplicate search: #3838 (`feat/collapsible-sidebar`) covers a similar sidebar area but is a different head branch and implementation. This PR intentionally packages the work from `PAP-10638-collapsable-sidebar` into one reviewable branch. Problem description: The board shell needs a first-class collapsed sidebar mode. Contextual surfaces such as company settings and plugin route sidebars should not replace the global app sidebar; they should collapse the app sidebar to a rail and render their contextual navigation beside it. ## What Changed - Added desktop collapsed/sidebar-peek state to `SidebarContext`, including persisted user pins, route collapse requests, and forced collapse for secondary-sidebar routes. - Replaced the old resizable sidebar pane with `SidebarShell`, which supports a fixed 64px rail, persisted expanded width, keyboard/pointer resizing, and hover/focus peek overlay behavior. - Updated `Sidebar`, sidebar nav items, project/agent sections, badges, and account/company menu presentation for expanded, collapsed, and peeking states. - Added `RequestCollapsedSidebar` and `SecondarySidebar` so routes and plugin `routeSidebar` slots can request contextual sidebar layouts without replacing the primary app sidebar. - Wired company settings and plugin route sidebars into the secondary-pane takeover model. - Added focused Vitest coverage for sidebar state precedence, shell sizing, nav item rail rendering, keyboard shortcuts, layout takeover behavior, and route collapse requests. - Updated plugin authoring docs/spec references for route sidebar behavior. ## Verification Targeted local verification passed: ```sh NODE_ENV=test pnpm run preflight:workspace-links && NODE_ENV=test pnpm exec vitest run ui/src/context/SidebarContext.test.tsx ui/src/components/SidebarShell.test.tsx ui/src/components/Sidebar.test.tsx ui/src/components/Layout.test.tsx ui/src/components/RequestCollapsedSidebar.test.tsx ui/src/components/SidebarNavItem.test.tsx ui/src/components/SidebarAgents.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/KeyboardShortcutsCheatsheet.test.tsx ui/src/hooks/useKeyboardShortcuts.test.tsx ``` Result: 10 test files passed, 88 tests passed. Additional follow-up verification passed after review fixes: ```sh NODE_ENV=test pnpm run preflight:workspace-links && NODE_ENV=test pnpm exec vitest run ui/src/components/Layout.test.tsx ui/src/context/SidebarContext.test.tsx && pnpm --filter /ui typecheck ``` Result: 2 test files passed, 28 tests passed, and UI typecheck passed. Latest PR-head remote checks: Paperclip PR workflow, Snyk, Socket, and Greptile are green; commitperclip `review` is cancelled in its security-gate step after filing a non-blocking neutral `security-review` check. Notes: - A direct run without `NODE_ENV=test` loads React's production build in this workspace, where `act` is unavailable; the command above matches the repo stable runner's test environment. - I did not run Playwright/browser e2e or full workspace build/typecheck in this PR-creation heartbeat. - QA screenshots are attached in https://github.com/paperclipai/paperclip/pull/7824#issuecomment-4661968387 for expanded, collapsed rail, hover peek, and settings secondary-sidebar states. ## Risks - Medium UI layout risk: this changes the board shell and primary sidebar composition across many routes. - Local storage migration risk is low: new collapsed state uses a new key and existing width storage remains scoped to the sidebar width. - Plugin route risk: plugin `routeSidebar` slots now render as secondary panes on desktop, so plugin authors should confirm their route sidebar content fits a 240px contextual pane. - Mobile risk appears low because mobile keeps the drawer model and gates collapsed/peek behavior to desktop. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with local shell/git/GitHub CLI tool use. Exact service-side model identifier and context window were not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a2230b2ec |
[codex] Guard document comment wake boundaries (#7766)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The execution control plane uses issue comments, assignments, monitors, blockers, and interactions to decide when agent-owned work should wake and run. > - Top-level issue comments are actionable issue-thread feedback for the assignee, but document-scoped comments are review context unless they are converted into an explicit routing primitive. > - Document annotation comments were still wired into the same `issue_commented` wake path as top-level issue comments. > - That made document activity capable of waking an assignee and looking like an execution path even when no issue-level handoff happened. > - This pull request narrows the wake boundary so document annotation activity stays document-scoped while normal issue comments continue waking the assignee. > - The benefit is fewer spurious wakeups and clearer non-terminal issue liveness semantics. ## Linked Issues or Issue Description Internal Paperclip work: [PAP-10613](/PAP/issues/PAP-10613), [PAP-10640](/PAP/issues/PAP-10640) Problem description: - Document annotation thread creation and annotation comments were treated as assignee wake sources. - Document-scoped activity should remain visible as document/review context, but should not by itself act as a queued issue wake, monitor, approval, interaction response, blocker, or terminal disposition. - Top-level issue comments should still wake the assignee on agent-assigned, non-terminal issues. Related PR search performed: - Found related prior document annotation work: #6733. - Found related prior issue-comment wake work and revert context: #7678, #7765. - No existing PR for `PAP-10613-why-is-this-task-not-running`. ## What Changed - Removed the document annotation comment assignee wake helper from issue routes. - Kept document annotation reference sync and activity logging intact. - Documented the distinction between top-level issue comments and document-scoped comments in `doc/execution-semantics.md`. - Added route tests proving document/document annotation activity does not wake the assignee. - Added route coverage proving top-level board issue comments still wake the assignee. ## Verification - `pnpm exec vitest run server/src/__tests__/document-annotation-routes.test.ts server/src/__tests__/issue-update-comment-wakeup-routes.test.ts` — 2 files passed, 9 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git status -sb` — clean branch tracking `origin/PAP-10613-why-is-this-task-not-running`. ## Risks - Low to moderate behavior change: document annotation comments no longer wake the issue assignee automatically. - Operators who want document feedback to route work must use an explicit primitive such as assignment, issue-thread comment, agent mention, issue-thread interaction, approval, blocker, or delegated follow-up. - No database migration or public API shape change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent with shell/tool use enabled. Exact hosted runtime model identifier beyond GPT-5 was not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7428fb956f |
[codex] Guard git-sensitive adapter workspaces (#7644)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The affected subsystem is the heartbeat execution path that turns issue assignment into adapter-backed work in a selected workspace. > - PAP-10409 and sibling follow-ups failed before useful adapter output because project/workspace identity became incoherent. > - A project-workspace-linked child issue could keep `projectWorkspaceId` / execution workspace state while losing `projectId`, then a git-sensitive local adapter could fall through toward an invalid fallback cwd. > - Paperclip needs to treat coherent workspace identity as part of the live-path contract, not only as post-failure cleanup. > - This pull request documents that rule, repairs issue inheritance, and blocks git-sensitive adapter launch before it can run from the wrong cwd. > - The benefit is a bounded recovery path: affected issues are repaired explicitly, future malformed workspaces fail fast with a clear recovery action, and the UI surfaces that reason. ## Linked Issues or Issue Description Refs #7646 Bug report fields: - Summary: adapter-backed follow-up issues can fail before doing work when issue creation/inheritance preserves workspace ids but drops project identity. - Affected issues: internal Paperclip issues PAP-10408 through PAP-10412, especially PAP-10409. - Steps to reproduce: create a project-scoped parent/follow-up tree where a child issue keeps `projectWorkspaceId` or an inherited execution workspace but has `projectId: null`, then launch a git-sensitive local adapter such as `codex_local`. - Expected behavior: Paperclip derives or preserves coherent project identity during issue creation, and heartbeat refuses malformed git-sensitive workspace launches with one clear recovery action. - Actual behavior before this PR: the run could reach adapter bootstrap with an incoherent workspace context and fail with git errors such as `fatal: not a git repository (or any parent up to mount point /srv)`. - Root cause: child/follow-up issue inheritance preserved workspace execution context without coherent project context. That let heartbeat workspace resolution/adapter launch reach a fallback cwd instead of refusing the malformed workspace state up front. ## What Changed - Documented the adapter workspace-coherence live-path precondition in `doc/execution-semantics.md`. - Updated issue creation/inheritance so workspace-inheriting issues preserve or derive project identity, while existing mismatch validation still rejects incoherent project/workspace combinations. - Added a heartbeat preflight guard for git-sensitive local adapters that validates effective cwd, persisted workspace identity, project workspace identity, and required git metadata before launch. - Added `workspace_validation` recovery actions for this failure class and ensured the source issue gets a visible, idempotent recovery comment. - Surfaced workspace-validation recovery state in issue rows, blocked notices, and recovery action cards, including the manual-repair wake policy label. - Added focused regression coverage for issue inheritance, all heartbeat workspace-validation guard branches, recovery display helpers, and UI recovery components. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` - Result: 1 test file passed, 68 tests passed. - `pnpm exec vitest run ui/src/components/IssueRecoveryActionCard.test.tsx` - Result: 1 test file passed, 12 tests passed. - `pnpm exec vitest run ui/src/components/IssueBlockedNotice.test.tsx ui/src/components/IssueRecoveryActionCard.test.tsx` - Result: 2 test files passed, 18 tests passed. - `pnpm --filter @paperclipai/ui typecheck` - Result: passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/issues-service.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts ui/src/components/IssueBlockedNotice.test.tsx ui/src/components/IssueRecoveryActionCard.test.tsx ui/src/lib/recovery-display.test.ts` - Result: 7 test files passed, 200 tests passed before the final guard-branch additions; the changed server file was re-run above. - UI coverage: `ui/storybook/stories/source-issue-recovery.stories.tsx` contains rendered scenarios for the generic recovery chip, workspace-validation recovery chip, blocked notice indicator, recovery action card, and issue-row chip. - Screenshot capture attempt: Storybook started successfully on `http://127.0.0.1:6016/`, but screenshots could not be captured in this runner because `agent-browser` launched an unusable Chrome binary and Playwright Chromium failed on missing system library `libatk-1.0.so.0`; the runner is non-root and lacks passwordless sudo for installing browser dependencies. - Hosted CI on final commit `969594e7` is green, including `verify`, `Build`, `Typecheck + Release Registry`, `General tests (server)`, workspace suites, serialized server suites, `Canary Dry Run`, and `e2e`. - Roadmap checked: no duplicate roadmap item; this is a tightly scoped reliability fix for existing heartbeat/workspace behavior. - Duplicate PR search checked: no open PR matched `workspace coherence adapter cwd`. ## Risks - Medium risk: heartbeat launch is stricter for git-sensitive local adapters and can now block malformed workspace states before adapter execution. - Mitigation: the guard is limited to local git-sensitive adapters and records a source-scoped recovery action with structured evidence instead of retrying indefinitely. - Compatibility: valid project/workspace execution paths continue normally; explicit project/workspace mismatches remain rejected. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based `codex_local` coding agent with terminal/tool use. Work was produced through Paperclip issue execution with focused local test runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dbebf30c89 |
Add low-trust review containment (#7530)
## Thinking Path > - Paperclip is a control plane for AI-agent companies, so execution policy and trust boundaries are part of the product's safety contract. > - Low-trust review work needs narrower authority than normal same-company agents because hostile PRs, comments, attachments, and generated output can carry prompt-injection payloads. > - The current V1 shape gives trusted workers broad company context, which is useful for normal execution but too permissive for a reviewer assigned to hostile content. > - This branch adds a `low_trust_review` preset, source-trust tagging, route-level containment, and quarantine handling so low-trust output does not automatically flow into higher-trust wake context. > - The branch has been rebased onto current `origin/master`, and the low-trust migration was renumbered to `0097_low_trust_source_trust.sql` to avoid collisions with existing `0091` through `0096` migrations. > - Greptile feedback was addressed by tightening low-trust detection, preserving project-level trust policy checks, fixing issue-kind promotion lookup, removing duplicate post-lease isolation assertion, documenting fail-closed source-trust behavior, bounding ancestry checks, enforcing runtime issue context for CEOs, awaiting accepted-plan monitor authorization, and making low-trust issue source-trust tagging atomic. > - The benefit is a first production slice of deny-by-default review containment with regression coverage for the main control-plane pivot surfaces. Fixes #7531. ## What Changed - Added shared trust-policy types and validators, plus database/source-trust fields for issues, comments, documents, and work products. - Implemented server enforcement for low-trust issue scope, agent self-view redaction, secret/plugin/runtime denial paths, promotion checks, and quarantined continuation/wake context. - Added focused low-trust regression tests for resolver behavior, source trust, route authorization, heartbeat preflight ordering, runtime containment, and quarantine redaction. - Added board UI affordances for selecting/reviewing the low-trust preset and surfacing source-trust badges in relevant issue views. - Added `doc/LOW-TRUST-PRESETS.md`, updated `doc/SPEC-implementation.md`, and committed the low-trust review contract plan under `doc/plans/`. - Rebasing note: the original `0097_low_trust_source_trust.sql` migration was renamed to `0097_low_trust_source_trust.sql`; the SQL uses `ADD COLUMN IF NOT EXISTS` so users who already applied the old-numbered migration are not broken by the renumbered migration. ## Verification - Rebased branch onto current `origin/master` and force-pushed with lease to `origin/PAP-10211-low-trust-agent` at head `2719f31e3`. - Confirmed the PR diff does not include `pnpm-lock.yaml` or `.github/workflows` changes. - Resolved upstream UI/comment conflicts by preserving deleted-comment tombstone behavior and low-trust source-trust badges/metadata. - Renumbered the low-trust source-trust migration to `0097_low_trust_source_trust.sql`; the SQL uses `ADD COLUMN IF NOT EXISTS` so users who already applied an old-numbered copy are not broken. - `pnpm exec vitest run ui/src/lib/issue-chat-messages.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts` - `pnpm exec vitest run server/src/__tests__/source-trust.test.ts server/src/__tests__/workspace-runtime-service-authz.test.ts ui/src/lib/trust-policy-ui.test.ts ui/src/components/TrustPresetSection.test.tsx` - `pnpm run typecheck:build-gaps` - `git diff --check` - GitHub checks pass on head `2719f31e3`: build, typecheck/release registry, general tests, serialized server suites, e2e, canary, verify, policy/review, Socket, and Snyk. - Greptile Review passes with Confidence Score 5/5 and zero unresolved Greptile review threads. - No design screenshots/images were added because the task explicitly says not to add them unless they are specifically part of the work. ## Risks - Medium risk: this touches shared trust-policy contracts, server authorization paths, heartbeat context generation, migration metadata, and UI preset controls. - Low-trust containment is intentionally deny-by-default; legitimate future review workflows may need explicit allowlisted exceptions. - Plugin/runtime/security surfaces are broad, so regression tests cover the current known routes but future integrations must route through the same containment layer. - The PR is ready for review; GitHub checks are green and Greptile is 5/5. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent, tool-enabled shell and GitHub CLI workflow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] UI changes are covered by focused tests; no screenshots were added per task instruction not to add design images unless specifically required - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fff3832a01 |
[codex] Add teams catalog extraction (#7550)
Fixes #7551 ## Thinking Path > - Paperclip is the control plane for AI-agent companies, and reusable company/team setup is part of making those companies faster to launch. > - The teams catalog work introduces app-shipped team templates that can be browsed, previewed, and installed into a company. > - Catalog installation crosses several contracts: bundled package contents, shared API types, server import/install behavior, CLI workflows, and the board UI. > - Agents also need a safe path through catalog installs: scoped company selection, explicit source policy, approval fallback for agent creation, and preserved catalog provenance. > - This pull request extracts the completed teams catalog branch into one reviewable PR on top of `public-gh/master`. > - The benefit is a reusable teams catalog foundation with server, CLI, package, docs, and hidden UI surfaces kept in sync. ## What Changed - Added the `@paperclipai/teams-catalog` package with bundled/optional team definitions, generated manifest, validators, catalog builder tests, and migration notes. - Added shared teams catalog types/validators plus server routes and services for listing, previewing, and installing catalog teams. - Integrated catalog install with company portability, skill/source policy checks, provenance metadata, origin hashes, target-manager reparenting, and installed/out-of-date detection. - Added CLI `teams` commands and agent-safe company selection behavior, including `company current` and approval fallback for forbidden agent-run installs. - Added hidden Team Catalog UI/API/query surfaces, Storybook fixtures, and targeted UI tests while keeping the UI route out of primary navigation. - Added docs for CLI/company/teams catalog behavior and removed generated screenshot artifacts from the PR diff. ## Verification - `pnpm exec vitest run cli/src/__tests__/company.test.ts cli/src/__tests__/teams.test.ts packages/teams-catalog/src/catalog-builder.test.ts packages/teams-catalog/src/shipped-catalog.test.ts server/src/__tests__/agent-permissions-service.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-skills-service.test.ts server/src/__tests__/teams-catalog-routes.test.ts server/src/__tests__/teams-catalog-service.test.ts server/src/__tests__/teams-catalog-install-no-overrides.test.ts ui/src/lib/company-routes.test.ts ui/src/pages/TeamCard.test.tsx ui/src/pages/TeamCatalog.test.tsx ui/src/pages/useInstallTeamCatalogEntry.test.tsx` - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/teams-catalog typecheck && pnpm --filter paperclipai typecheck && pnpm --filter @paperclipai/server typecheck && pnpm --filter @paperclipai/ui typecheck` - Confirmed branch is rebased onto `public-gh/master` (`78dc3625a`) and `public-gh/master` is an ancestor of `HEAD`. - Confirmed PR diff excludes `pnpm-lock.yaml`, `.github/workflows/*`, generated screenshot images, and screenshot helper scripts. ## Risks - Medium review surface: this crosses package generation, shared contracts, server install behavior, CLI, docs, and hidden UI code. - Catalog install behavior creates agents/projects/tasks/skills and must keep company scoping, permissions, source policy, and provenance checks strict. - `pnpm-lock.yaml` is intentionally excluded per repo policy; CI/default-branch automation owns lockfile refresh. - The Team Catalog UI is included but hidden from primary navigation, so future enablement should re-check visual QA before exposure. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. > > ROADMAP checked: this aligns with reusable companies/templates and plugin-adjacent onboarding work. This PR packages work already developed on the Paperclip task branch for review. ## Model Used - OpenAI Codex, GPT-5 series coding agent in this Paperclip session; exact runtime context window was not exposed. Used shell, git, `gh`, and local test/typecheck tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots, or documented why screenshots are intentionally omitted - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
70b1a9109d |
Improve CLI API parity coverage (#6626)
## Thinking Path > - Paperclip is a control plane for AI-agent companies, with the CLI acting as a scriptable operator and agent interface to that control plane. > - The REST API surface has grown across companies, agents, issues, routines, plugins, auth, workspaces, secrets, and operational inspection commands. > - The CLI had drifted from that API surface: some commands were missing, some command shapes differed from docs/reference material, and several edge cases only failed during end-to-end local-source testing. > - The local development runbook requires these tests to be disposable and isolated from a real `~/.paperclip`, `~/.codex`, or `~/.claude` installation. > - This pull request adds broad CLI/API parity coverage, fixes the actionable bugs found during that pass, and records the reproducible test log under `doc/logs`. > - The benefit is a more complete, scriptable CLI surface with regression coverage for the command families exercised by the parity run. ## What Changed - Added or expanded CLI command coverage for access/auth, companies, agents, projects, goals, issues and subresources, routines, plugins, workspaces, activity/run/cost/dashboard inspection, assets, skills, secrets, tokens, prompt/wake flows, and local setup helpers. - Fixed CLI/API parity bugs found during the run, including context profile patching, issue interaction optional payloads, malformed tree-hold errors, environment duplicate handling, configure invalid-section exit codes, worktree pnpm invocation, token agent ID resolution, plugin tool worker lookup, and routine webhook secret cleanup. - Added missing CLI wrappers and route coverage for health/access, invite resolution URL forwarding, join status normalization, secret lifecycle commands, LLM docs routes, available-skill isolation, positive board-claim coverage, and interactive `connect` prompt-flow tests. - Added a schema-backed `/api/openapi.json` route sufficient for CLI parity and `paperclipai openapi --json` smoke coverage. - Added `doc/logs/2026-05-24-cli-api-parity-e2e-log.md` with the detailed living test/bug log and renamed the log directory from `doc/bugs` to `doc/logs`. - Added `doc/plans/2026-05-23-cli-api-parity.md` and the OpenAPI parity reference used during the pass. OpenAPI note: this PR intentionally does not try to subsume `feature/openapi-spec`. The OpenAPI implementation here is schema-backed and better than the earlier route-inventory stub, but `feature/openapi-spec` is the fuller/better OpenAPI branch because it includes exact mounted-route coverage tests and additional current route coverage. That branch should stay as its own PR and can supersede this OpenAPI route implementation. ## Verification Targeted automated checks run: - `pnpm exec vitest run server/src/__tests__/openapi-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/board-claim.test.ts` - `pnpm exec vitest run cli/src/__tests__/connect.test.ts` - `pnpm exec vitest run cli/src/__tests__/agent-lifecycle.test.ts` - `pnpm exec vitest run server/src/__tests__/plugin-database.test.ts` - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts` - `pnpm --dir cli typecheck` - `pnpm --dir server typecheck` Manual/local E2E verification: - Ran the full disposable local-source CLI/API parity pass with isolated `PAPERCLIP_HOME`, `PAPERCLIP_CONFIG`, `PAPERCLIP_CONTEXT`, `PAPERCLIP_AUTH_STORE`, `CODEX_HOME`, and `CLAUDE_HOME` under `tmp/cli-api-parity`. - Verified `DATABASE_URL` and `DATABASE_MIGRATION_URL` stayed unset for the scratch server. - Verified live health and schema-backed OpenAPI responses on non-default port `3197`. - Revoked created board/agent tokens and cleaned up temporary plugins, secrets, non-default environments, and project workspaces. - See `doc/logs/2026-05-24-cli-api-parity-e2e-log.md` for the full command-by-command reproduction log. Not run: - Full `pnpm test`, `pnpm test:run`, or `pnpm build` were not run after the entire branch because the branch is broad and the parity pass used focused test/typecheck verification plus live isolated CLI reruns. ## Risks - This is a broad PR and touches many CLI command modules, so review surface is high. The changes are grouped around one theme, but a split may be easier if maintainers prefer narrower PRs. - The OpenAPI route in this PR is not the final/best OpenAPI implementation. `feature/openapi-spec` has stronger exact-route coverage and should remain the source for the dedicated OpenAPI PR. - The living log is intentionally detailed and large. It is useful for reproducibility but adds documentation weight. - No UI changes are intended; screenshots are not applicable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent in Codex desktop. Exact served model/context-window identifier was not exposed in the local app. Work used shell/Git/GitHub CLI tooling, local source inspection, targeted test execution, and live isolated Paperclip CLI/API smoke testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Devin Foley <devin@devinfoley.com> |
||
|
|
c4bb68c14b |
Bundle artifact upload helper with Paperclip skill
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e7cdd0f8c5 |
Move artifact upload guidance into Paperclip skill
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0bd13c23a9 |
Add agent artifact upload workflow
Co-Authored-By: Paperclip <noreply@paperclip.ing> |