mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
53bcf3897f05bb279a33e6289a12d3f0cc2d16c3
312
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd7f84965c |
fix(recovery): preserve hand-back wake liveness (#10562)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service delivers issue work to assigned agents. > - Recovery can hand an issue back to its agent while the recovery run is still active. > - The hand-back wake can merge into that active run and disappear when the run exits. > - The stranded-work scan also treats the successful recovery run as proof that the handed-back issue is live. > - This pull request keeps the hand-back wake for follow-up delivery and lets the scan repair a lost wake. > - The benefit is that an assigned issue continues after recovery without manual operator action. ## Linked Issues or Issue Description No public issue exists. This is related to the wake reconciliation work in #8943. **What happened?** A recovery action could hand an assigned issue back from `blocked` to `todo`. The `issue_recovery_action_restored` wake then merged into the recovery run that made the change. The wake disappeared when that run exited. The stranded-work scan did not repair the issue because it treated the successful recovery run as current liveness. **Expected behavior** Paperclip must dispatch the hand-back wake after the recovery run exits. If that delivery is lost, the stranded-work scan must enqueue the assigned `todo` issue again. **Steps to reproduce** 1. Start a recovery run for an assigned blocked issue. 2. Resolve a recovery action with the `handed_back` outcome. 3. Move the issue to `todo` while the recovery run is still active. 4. Observe that the wake merges into the active run and no new run starts after it exits. 5. Run the stranded-work scan and observe that the successful latest run prevents repair. **Paperclip version or commit** `131d476a7e` **Deployment mode** Local dev (`pnpm dev`). The defect is in the core server and is not deployment-specific. **Agent adapter(s) involved** Not adapter-specific. This is a core heartbeat and recovery defect. ## What Changed - Added `issue_recovery_action_restored` to the wake reasons that require follow-up delivery when an issue run is active. - Made the stranded-work scan detect a resolved hand-back that occurred during or after the latest successful run. - Added focused regression tests for the heartbeat coalescing seam and the stranded-work repair shape. - Documented the hand-back liveness guarantee in execution semantics section 9.1. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` passed: 104 tests. - `pnpm --filter @paperclipai/server typecheck` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm test:run` passed the server shard (3,095 passed, 2 skipped) and UI shard (3,182 passed). One unrelated CLI test failed because the agent environment exports static AWS credentials. `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY -u AWS_SESSION_TOKEN pnpm exec vitest run cli/src/__tests__/secrets.test.ts` passed all 8 tests. - `git diff --check` passed. - All GitHub checks passed on commit `8b380e67e6`. - Greptile gave 5/5 confidence with no comments or unresolved threads. ## Risks - Low risk. The follow-up rule affects only a recovery hand-back wake that arrives while the same issue already has an active run. - The backstop adds one indexed recovery-action lookup for an assigned `todo` issue whose latest run succeeded. - The timestamp check uses the latest run start time. This includes hand-backs made by that run and later hand-backs, but excludes older resolved actions. ## Model Used - OpenAI Codex with GPT-5 (`gpt-5`), agentic reasoning, tool use, and code execution. The serving context-window size is not exposed to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7cfb655f60 |
fix(worktree): disable automatic database backups (#10520)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip creates isolated instances for linked git worktrees so development does not affect the primary instance > - Those instances inherited the source instance's automatic database-backup setting and also repaired older configs without overriding it > - As worktrees accumulated, each isolated instance could schedule its own backup stream, producing redundant backup churn for disposable database clones > - This pull request makes backup disablement an invariant of worktree config creation and repair > - The benefit is that automatic backups remain focused on the durable primary instance while isolated development instances stop accumulating redundant backup files ## Linked Issues or Issue Description No public GitHub issue exists for this bug, so the report is included here. The closest related open change is Refs #10266, which hardens where worktree config repair may write; this PR changes the backup policy applied by that repair and by worktree initialization. ### What happened? Isolated worktree instances copied `database.backup.enabled` from their source config. When the source instance enabled automatic backups (the normal default), every linked worktree also enabled a scheduled backup stream. Existing worktree configs kept that state during startup repair, so the redundant backups continued after the policy changed. ### Expected behavior Automatic database backups are disabled for isolated worktree instances created by `paperclipai worktree init` or `paperclipai worktree:make`, and legacy worktree configs are migrated to that policy during normal startup repair. The durable primary/default instance keeps its existing backup behavior. ### Steps to reproduce 1. Start from a Paperclip instance whose database backup setting is enabled. 2. Create or initialize a linked worktree with `paperclipai worktree init`. 3. Inspect the generated worktree config and environment. 4. Before this change, the config retained `database.backup.enabled: true` and the environment had no disabling override; after this change, the config is false and `PAPERCLIP_DB_BACKUP_ENABLED=false` is persisted. ### Paperclip version, deployment mode, and environment - Reproduced against `master` before commit `ea5e0a0269`. - Deployment mode: local trusted development with linked git worktrees and embedded PostgreSQL. - Environment: Node.js 22, pnpm workspace install. ## What Changed - Always generate isolated worktree configs with automatic backups disabled. - Persist `PAPERCLIP_DB_BACKUP_ENABLED=false` in generated worktree environments. - Repair existing isolated worktree configs and environments that still enable backups. - Add CLI and server regression coverage for creation and legacy repair paths. - Document the worktree-specific backup policy and primary-instance exception. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts server/src/__tests__/worktree-config.test.ts` — 52 tests passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - All repository commands above were run with inherited worktree runtime identity variables removed. ## Risks - Low operational risk: the change is limited to explicitly isolated worktree instances. - Operators who intentionally relied on automatic backups of disposable worktree databases will now need to run a manual backup or explicitly manage those files outside the scheduled worktree runtime. - No schema, migration, API, UI, lockfile, or workflow changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5 (the runtime does not expose a more specific snapshot ID or context-window value), using reasoning, tool use, local code execution, and GitHub CLI integration. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b0156238a |
docs(connections): cross-reference the identity model (identity vs. connections) (#9953)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Integrations are governed through the Apps v2 "connections" substrate documented in `doc/connections/` > - Contributors adding connectors had no written boundary between *signing a user in* (identity) and *connecting an external resource* (a connection), so app-store work could drift into merging the two > - Unwritten, that boundary risks a design where a sign-in provider's token gets reused as a resource token, or the identity service grows a connections hub — exactly the failure the separation exists to prevent > - This pull request adds an "Identity vs. connections" section to the public connections docs that names the three planes, states the standing rule verbatim, and aligns naming across surfaces > - The benefit is that connector implementers inherit the identity/connection boundary from accessible repository documentation instead of re-deriving it ## Linked Issues or Issue Description Docs-only change. No public GitHub issue. **Problem:** The `doc/connections/` docs describe how to add a connector (the First-30 matrix and connector playbook) but never state where identity ends and connections begin. Without an accessible statement of that boundary, connector work can accidentally blur sign-in tokens and resource tokens. **Change:** Adds a public reference section fixing the boundary and makes it self-contained for contributors without access to private identity-service implementation documentation. ## What Changed - `doc/connections/README.md`: adds an **Identity vs. connections** section with the P1/P2/P3 model, D7 standing rule, rationale, naming alignment, and canonical-doc cross-reference. - `doc/connections/FIRST-30-MATRIX.md`: notes that every listed provider is a plane-P2 connection and links the new section. - `doc/connections/CONNECTOR-PLAYBOOK.md`: points connector authors to the boundary and standing rule before implementation. - Removes inaccessible private-repository documentation links so every reference in the new public guidance is usable by contributors. ## Verification - `git diff --check origin/master...HEAD` - Verified Markdown table column consistency across the three changed files. - Verified no private `paperclip-id` URLs or new internal Paperclip issue references remain in the diff. - Current-head CI is fully green, including typecheck, build, server/workspace tests, serialized suites, canary, and both e2e shards. - Greptile reviewed commit `d964abab50b39a79fef394c0b1855dd1f2e4b92b` at 5/5 with no unresolved findings; its connections-doc validation returned `RESULT: PASS`. ## Risks Low risk — documentation only. No code, schema, migration, or runtime behavior changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 (`claude-opus-4-8`), 1M context, extended thinking with tool use, authored the initial documentation change. - OpenAI Codex CLI (runtime model identifier not exposed to this session), agentic reasoning with shell and GitHub tooling, performed PR preparation, rebase, review-loop diagnosis, and the public-link correction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
71a6535792 |
docs(execution-semantics): define routable blocking and watchdog restoration (#10094)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Execution semantics define when work may stop, how child work reports upward, and how watchdog recovery proves liveness was restored. > - The existing docs did not require a routable waiting path before entering `blocked`, leaving permission-denial and review workflows vulnerable to dead ends. > - Child-to-parent reporting and low-trust review delegation also needed explicit channels and ownership boundaries. > - Watchdog recovery needed bounded restoration verification rather than treating a recovery write as proof of restored progress. > - The security-sensitive defaults were reviewed and approved before publication. > - This pull request documents those contracts and synchronizes the bundled Paperclip skills that operationalize them. > - The benefit is clearer stop conditions, safer reporting defaults, and bounded watchdog recovery without sacrificing liveness. ## Linked Issues or Issue Description - No public GitHub issue exists for this execution-semantics documentation phase, so the problem and solution are described here in full. - Problem: `blocked` could be entered without a routable owner/action path, review findings could be reported on the wrong issue or treated as blockers, and watchdog recovery lacked bounded verification of restored liveness. ## What Changed - Require a routable waiting path for `blocked` and clarify that permission denial alone is not a blocker. - Define canonical child-to-parent completion/report channels, review-delegation ownership, and the sanctioned courier pattern. - Specify atomic watchdog recovery batches, fingerprint validation, bounded restoration attempts, and human escalation. - Document the low-trust preset default for report comments and align both bundled Paperclip skills. ## Scope and Sequencing - This is the contract-first P1 documentation PR. Runtime enforcement is intentionally excluded and follows in the separately owned implementation phases; these docs define the acceptance contract those phases must satisfy. - The five-file patch is approved and frozen for this PR, so review findings about absent runtime support are tracked as implementation-phase requirements rather than edits to this docs-only change. ## Verification - `git diff --check origin/master...HEAD` - Confirmed the PR diff is exactly 5 approved docs/skill files with 91 insertions and 1 deletion. - Confirmed the patch preserves the approved execution-semantics contract after replay on latest `origin/master`. ## Risks - Low implementation risk: documentation and skill guidance only; no runtime code, schema, migration, dependency, lockfile, or workflow changes. - Semantic risk is limited to readers or agents applying the clarified contracts; the security-sensitive R1a default was approved before publication. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent (exact underlying model ID and context-window size are not exposed by this runtime), with reasoning, terminal tool use, Git/GitHub operations, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced private URLs or internal issue identifiers - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the applicable docs-only checks locally and they pass - [x] Tests are not applicable because this is a docs/skill-guidance-only change - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e55d702916 |
feat(plugin-sdk): human-attributed issue comments for chat gateway plugins (#10050)
Adds the `issue.comments.create_human_attributed` capability and `ctx.issues.createComment` `actorUserId` option, with host-side active-human-member verification. LOOA-627. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b247bf7150 |
feat(runtime): opt-in sandbox file-sync lifecycle hooks (API + provider docs) (#10013)
## Thinking Path > - Paperclip is an open-source AI-agent management platform; agents run tasks inside sandboxed environments (Daytona, Kubernetes, E2B, etc.) > - The control-plane ↔ sandbox file-transfer path flows through the `environmentExecute` seam in `protocol.ts` — the only verb available to plugins — which forces a base64-over-exec chunked loop for every file move: workspace files, assets, Codex home sync > - This transport is correct and safe, but it bypasses provider-native bulk/streaming APIs (Daytona `uploadFiles`, K8s `FastUploadInterceptor` / volume mounts), leaving significant throughput on the table for large workspaces > - The right fix is an opt-in seam extension: providers with faster native transfer declare two optional verbs; providers that do not opt in stay on the existing fallback with zero code or behavior change required > - This PR adds the first layer of that extension — two optional verbs (`environmentSyncIn` / `environmentSyncOut`) in the plugin SDK, the runtime plumbing to prefer the native path for the two clean destroy-then-replace cases, and a doc for the contract > - The core correctness invariant is byte-identical fallback: if no provider opts in, execution is exactly what ships today; `assertSyncOperationsConfined` enforces host-side path confinement for providers that do opt in > - No provider advertises the verbs yet → zero production behavior change; future PRs wire up Daytona and K8s providers against this contract ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Description follows the `feature_request` issue template: **Subsystem affected:** packages/plugins — plugin system; packages/adapter-utils — adapter runtime; server/ — EnvironmentRuntimeService **Problem or motivation:** Sandbox file transfers currently always use a base64-over-exec chunked loop regardless of what the underlying provider supports. For workspaces larger than a few MB this becomes the dominant wall-clock cost of every sandbox run, and it bypasses bulk/stream APIs that providers like Daytona already expose natively. **Proposed solution:** Add two optional, opt-in plugin hooks — `onEnvironmentSyncIn` / `onEnvironmentSyncOut` — to the plugin SDK. When a provider defines both hooks and both are advertised via the existing `supportedMethods` negotiation, the runtime prefers the native path for the two clean destroy-then-replace transfer cases; all other cases fall back to the existing byte-identical base64 transport. **Alternatives considered:** An unconditional verb would require every provider to implement or stub the verb. The opt-in / `METHOD_NOT_IMPLEMENTED` pattern (already used by `environmentExecute`) preserves backward compatibility with zero provider changes required. **Roadmap alignment:** Consistent with the ✅ "Cloud / Sandbox agents" and ✅ "Plugin system" milestones; extends the plugin seam rather than adding control-plane-level logic. **Additional context:** Searched open pull requests and issues for duplicate sandbox file-sync / native-transfer work; none found. ## What Changed - **`packages/plugins/sdk`** - `protocol.ts`: two new optional `HostToWorkerMethods` — `environmentSyncIn` / `environmentSyncOut` — plus generic `SyncOperation`, `SyncFileMapping`, and `SyncOutcome` types - `define-plugin.ts`: optional `onEnvironmentSyncIn` / `onEnvironmentSyncOut` fields on `PluginDefinition`; worker advertises each verb only when its hook is defined (else `METHOD_NOT_IMPLEMENTED`, mirroring `environmentExecute`) - `worker-rpc-host.ts`: route new verbs to plugin hooks - `index.ts`: re-export new public types - **`packages/adapter-utils`** - `command-managed-runtime.ts`: expose optional `syncIn` / `syncOut` on `CommandManagedRuntimeRunner` (available only when both verbs are advertised); add `assertSyncOperationsConfined` host-side path-confinement guard - `sandbox-managed-runtime.ts`: `SandboxManagedRuntimeClient` gains optional `syncIn` / `syncOut`; orchestrator prefers native path for default-provision asset inbound and workspace-download-into-fresh-dir outbound; all other paths keep the existing base64 fallback - `sandbox-file-sync.test.ts` (new): 234-line characterization suite — native-opt-in branch, fallback branch, `assertSyncOperationsConfined` escape-path rejection, `followSymlinks` → tar `-h` - `command-managed-runtime.test.ts`: negotiation + native-sync + confinement tests - **`server/src/services/environment-runtime.ts`**: `EnvironmentRuntimeService` delegates to `syncIn` / `syncOut`, gated on advertised support - **`server/src/services/environment-execution-target.ts`**: minor typing fix alongside the new verbs - **`doc/plugins/SANDBOX_FILE_SYNC_HOOKS.md`** (new): documents the full contract — opt-in / no-op guarantee, operation ordering, provider-may-tar, atomicity, `followSymlinks`, secret modes (0600, no window), path confinement, `operationId` opacity, resource bounds, shell-quoting ## Verification ```bash # SDK suite pnpm --filter packages/plugins/sdk test # Adapter-utils suite (includes new sandbox-file-sync characterization tests) pnpm --filter packages/adapter-utils test # Expected: 255 pass / 4 skip # Type-check across affected packages pnpm --filter packages/plugins/sdk typecheck pnpm --filter packages/adapter-utils typecheck # Server changed-file spot check: cd server && npx tsc --noEmit --skipLibCheck 2>&1 | grep -E "environment-(runtime|execution-target)" | head -20 ``` Key behavioral invariant to spot-check: with no provider opting in (the current state), run any sandbox task and confirm file-transfer behavior is byte-for-byte identical to what the pre-PR code produces. The characterization tests assert this at the unit level. ## Risks - **Zero production risk today**: no provider advertises `environmentSyncIn` / `environmentSyncOut`, so the new code paths are unreachable in production; all real traffic stays on the existing base64 fallback - **Path confinement**: `assertSyncOperationsConfined` rejects any `targetPath` that escapes the declared root — this is the primary security boundary for future providers. The test suite covers escape-path rejection - **Atomicity**: the contract delegates atomicity to providers; the doc explicitly calls out that directory-level ops are not guaranteed atomic - **Secret transport**: credential assets (e.g., Codex `auth.json`, directory mappings) continue to use the existing tar path — they do not go through the new verbs in any current provider > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: Anthropic Model: `claude-sonnet-4-6` (Claude Sonnet 4.6) Context window: 200 K tokens Capabilities: extended tool use, multi-file code generation, agentic reasoning via the Paperclip agent framework ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d23fbf8ae4 |
feat(connections): add AppDefinition Wave 1 catalog (#9981)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Connections is the subsystem that defines which external apps and MCP-style integrations operators can browse, configure, and run > - The v3 schema core in #9958 added stable connection identities, auth metadata, and grant-aware contracts, but the app catalog still used the older gallery shape > - The product needs a richer, typed AppDefinition catalog so browsing and setup can render provider-specific auth and configuration requirements consistently > - This pull request moves the Wave 1 app catalog onto generated AppDefinition data and carries that shape through shared types, server lookup paths, and app connection UI > - The benefit is that follow-up runtime and wizard work can build against one catalog contract instead of local-only mock/gallery data ## Linked Issues or Issue Description Refs #9958. No public GitHub issue exists for this branch. This is the catalog layer for the Connections v3 stack after the schema-core foundation in #9958. ## What Changed - Adds generated AppDefinition data for the Wave 1 catalog and ingestion reporting. - Replaces the legacy tool app gallery exports with AppDefinition-centered shared contracts, validators, and tests. - Updates server tool-access lookup behavior to use the AppDefinition catalog. - Updates app connection UI surfaces and tests to consume AppDefinition-backed catalog data. - Documents the catalog ingestion workflow in the connector playbook. ## Verification - `pnpm run preflight:workspace-links` - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx server/src/__tests__/tool-access-service.test.ts` ## Risks - Medium: this changes the catalog contract used by shared, server, and UI app connection surfaces. - Catalog data quality matters because generated definitions now drive browse/setup display. - Follow-up runtime and wizard PRs must rebase on this branch or on master after this lands. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent with repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e00f67138 |
feat(connections): add v3 schema core (#9958)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their governed access to external systems. > - Connected Apps build on the existing Apps and MCP gateway substrate so companies can configure reusable, auditable integrations. > - The current connection record does not yet have a stable public address, explicit ownership/auth method fields, or subject-specific credential grants. > - Without that schema core, later OAuth, per-user authorization, token brokering, triggers, and connector-service phases cannot enforce tenant and subject boundaries consistently. > - This pull request adds the forward-compatible Connections v3 schema core while preserving the existing connection lifecycle and directly migrating the remote MCP transport name. > - The benefit is a company-scoped, least-privilege foundation for one-click integrations without bypassing Paperclip secrets, profiles, rules, or audit controls. ## Linked Issues or Issue Description No matching public issue was found. **Problem** Paperclip's current app connections need a durable identity and authorization substrate before Connected Apps can safely support multiple setup methods, per-user credentials, provider tenants, and managed connector services. The existing schema only models a single connection-level credential set and uses legacy transport terminology. **Proposed solution** Add a stable company-scoped connection UID, explicit ownership/auth/transport fields, a subject-aware `connection_grants` table, and multi-key credential annotations. Backfill existing connections and workspace grants in a reversible migration, then update shared/server/UI contracts to the new `mcp_remote` transport name. **Related work** - Related foundation: #9534 - Roadmap: Connected Apps (one-click integrations) ## What Changed - Added company-scoped connection `uid`, `ownership`, `authKind`, and canonical transport fields across database, shared contracts, validators, services, and UI fixtures. - Added `connection_grants` with workspace/user subject rules, provider tenant metadata, credential secret refs, revocation state, company scoping, and uniqueness constraints. - Added migration `0182_connections_v3_schema_core` to backfill stable UIDs, rename `remote_http` to `mcp_remote`, infer auth kinds, create default workspace grants, and support rollback coverage. - Added multi-key credential annotations and updated gateway/access services without changing the existing lifecycle behavior. - Updated the connection glossary, connector playbook, and security threat model for the new identity, grant, and relay boundaries. - Added explicit test UIDs to direct database fixtures so the new non-null invariant is exercised across affected server suites. ## Verification - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts server/src/__tests__/tool-gateway-service.test.ts server/src/__tests__/tool-gateway.test.ts server/src/__tests__/heartbeat-runtime-skills.test.ts server/src/__tests__/tool-oauth-legacy-backfill.test.ts server/src/__tests__/tool-access-policy-service.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts packages/db/src/connections-v3-schema-core-migration.test.ts packages/shared/src/validators/tool-access.test.ts --config vitest.config.ts` — 9 files, 218 tests passed. - Latest-head GitHub Actions: build, typecheck, general/serialized suites, backup/worktree restore coverage, both e2e shards, canary, policy, and security scans pass. - Greptile: 5/5 with zero unresolved threads. - `pnpm check:token-gates` remains red only on five pre-existing `#9627` color literals outside this change. ## Risks - **Migration risk:** UID backfill and default-grant creation touch every existing connection. The migration uses company-scoped uniqueness, deterministic legacy UIDs with ID suffixes, and seeded up/rollback coverage. - **Authorization risk:** Grant rows carry credential references. Constraints enforce workspace-vs-user subject shape, company/connection lookup indexes, one default grant per connection, and one user grant per connection/subject. Security review is requested specifically for this design. - **Compatibility risk:** `remote_http` is renamed directly to `mcp_remote`; all repository call sites and fixtures are updated in the same change. - **Future-phase risk:** Subject-bound token issuance, triggers, and connector-service relay verification remain fail-closed requirements documented for later phases; this PR does not expose those capabilities. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI coding agent. The runtime did not expose an exact underlying model ID or context-window size; capabilities used include repository inspection, code editing, shell execution, test execution, Git/GitHub CLI operations, and structured reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
631cca509d |
docs: refresh README roadmap and four pillars (#9922)
## Thinking Path > - Paperclip is the open source control plane people use to organize and govern companies of AI agents > - The README is the project’s front door, while `ROADMAP.md` is the detailed statement of shipped and planned product direction > - The existing README roadmap had fallen behind capabilities that are already available and did not explain the product’s organizing model > - That drift made current features look unfinished and left the relationship between tasks, organization, training, and infrastructure implicit > - This pull request synchronizes the roadmap preview with the detailed roadmap and introduces the four-pillar product framing > - The benefit is a more accurate, coherent first impression for users and contributors, with theme-aware visuals that render correctly on GitHub ## Linked Issues or Issue Description - **Subsystem affected:** Documentation (`README.md`, `ROADMAP.md`, and documentation assets); no runtime subsystem changes. - **Problem or motivation:** The README roadmap understated shipped capabilities, the detailed roadmap had corresponding stale statuses, and the README did not clearly explain the four product pillars that connect Paperclip’s task, organization, training, and infrastructure surfaces. - **Proposed solution:** Refresh both roadmap views from one ordered list, add the four-pillars section and light/dark graphic, and keep future capabilities explicitly marked as planned or partial. - **Alternatives considered:** Updating only `ROADMAP.md` would leave the project’s front page stale; updating only the README would create a second source of truth; a single non-theme-aware image would be less legible for half of GitHub’s color schemes. - **Roadmap alignment:** This change directly updates `ROADMAP.md` and its README preview rather than introducing an unplanned runtime feature. - **Additional context:** Searched open pull requests for related README roadmap/four-pillars work and found no duplicate. ## What Changed - Added a **The four pillars** section covering Agentic Task Manager, Org Chart for Agents, Agent Employee Training, and Agentic OS, with a theme-aware `<picture>` graphic. - Flipped four roadmap entries from ⚪ to ✅: Cloud / Sandbox agents, Artifacts & Work Products, Deep Planning, and Enforced Outcomes. - Added shipped ✅ entries for MCP Tool Gateway & Apps, per-agent Secrets Manager access, activity logging and action attribution, self-healing recovery, and agent evals/feedback; expanded the skills entry to include Skill Studio and Skills Store. - Added planned ⚪ entries for Bring-your-own-ticket-system and Connected Apps, resolving the README FAQ inconsistency around external ticket systems. - Moved Cloud deployments to 🟡 and documented the shipped multi-tenant isolation, local-to-cloud sync, and cloud-managed bootstrap scope. - Synchronized the README preview and `ROADMAP.md` ordering/statuses: 18 ✅, 9 ⚪, and 1 🟡 entry. ## Verification - `git diff --check origin/master...HEAD` - Focused Node consistency check confirms README and `ROADMAP.md` have the same 28 ordered entries/statuses and expected 18 ✅ / 9 ⚪ / 1 🟡 counts. - `file doc/assets/four-pillars-light.png doc/assets/four-pillars-dark.png` confirms both assets are valid 3840×2160 RGB PNGs. - Visual QA rendered the README through GitHub’s GFM API in Chromium at GitHub content width under both light and dark `prefers-color-scheme` modes; the correct graphic swapped per theme, all links resolved, the section placement/table were correct, and `ROADMAP.md` matched the README preview. ### Visual Evidence **Light theme**  **Dark theme**  - Prettier was not run because this repository does not expose a configured `prettier` binary; markdown whitespace was checked with `git diff --check` and the rendered QA above. ## Risks - Low risk: documentation and image assets only; no application code, API contract, schema, dependencies, lockfile, or workflow changes. - Roadmap status wording can become stale as product work evolves, so future shipped capabilities should update both the README preview and `ROADMAP.md` together. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, with reasoning, terminal/tool use, GitHub API access, and Paperclip control-plane integration. The model context-window size was not exposed by this execution environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a090c09ee5 |
feat: add decision training snapshot foundation (#9702)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - Human approvals, issue interactions, and execution decisions already capture high-value decision moments > - Those moments are currently transient and cannot be reused as stable evaluation or training examples > - Reusable examples need a server-owned, immutable snapshot so later comments or runs cannot leak into the recorded state > - Human notes need to remain editable and auditable without changing the captured state > - This pull request adds the database model, snapshot capture service, API, export format, and attention-feed enrichment for decision training > - The benefit is a durable, inspectable foundation for evaluating whether agents can reproduce good human decisions from only the context available at decision time ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (`server/`, `packages/db`, and `packages/shared`). ### Problem or motivation Paperclip has no durable dataset for converting human decisions into evaluation-ready examples. Teams need to capture pending or resolved decisions with the exact issue context, comments, runs, and repository evidence available at a cutoff, while preventing future context from leaking into the example. ### Proposed solution Store immutable, schema-versioned snapshots anchored to durable interaction, approval, or execution-decision records; keep notes separately editable with history; expose human-only CRUD, list, and JSONL export APIs. ### Alternatives considered Client-generated snapshots were rejected because they duplicate cutoff logic and cannot reliably enforce no-leakage boundaries. Automatic outcome backfill was deferred so captured examples remain faithful to what was known at capture time. ### Roadmap alignment Supports the roadmap direction of turning completed work and decision patterns into reusable organizational knowledge. ### Additional context The implementation records explicit commit-resolution confidence (`exact`, `nearest_run`, `workspace`, or `none`) so downstream evaluation can distinguish evidence quality. ## What Changed - Added the `decision_training_examples` schema and idempotent migration with company, issue, and source/author indexes. - Added shared types for decision-training records, notes history, and versioned snapshots. - Added a single server-side snapshot capture path with inclusive comment cutoffs, pre-cutoff run capture, durable decision payloads, and explicit commit-resolution confidence. - Added create, list, detail, notes-only update, delete, and JSONL export routes with human-only write authorization and activity logging that skips no-op note submissions. - Added per-user `trainingExampleId` enrichment to attention items. - Added focused embedded-Postgres tests for cutoff boundaries, post-cutoff leakage, immutable snapshots, human-only writes, duplicate prevention, notes history, attention enrichment, and export shape. - Updated UI test and Storybook attention-item factories for the new required `trainingExampleId` contract. ## Verification - `pnpm exec vitest run server/src/__tests__/decision-training.test.ts` — 10 tests passed. - `pnpm --filter @paperclipai/db typecheck` — passed, including migration numbering and safety checks. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. ## Risks - The migration adds a new table and indexes only; it does not rewrite existing rows or install resolve-time hooks. - Snapshot JSON can grow with long comment threads and run histories; v1 intentionally favors complete, inspectable examples over aggressive truncation. - Commit SHA resolution is evidence-based and records `exact`, `nearest_run`, or `none` so downstream consumers can account for confidence. - The API is additive, but future UI work must continue to treat the snapshot as immutable and use notes-only updates. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.3-codex`, with repository tool use, terminal execution, and code-editing capabilities; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
59fb27ff79 |
feat(inbox): let agents safely tidy user inboxes (#9724)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their work > - The inbox is a per-user attention view, so archiving an item must not alter the underlying issue, assignment, or status > - Agents can help responsible users tidy resolved work only when the action is company-scoped, reversible, policy-controlled, and fully attributable > - The database and authorization foundations landed in #9654 and #9658, but the end-to-end archive routes, audit details, agent workflow guidance, and operator UI still need to ship together > - Separate stacked PRs #9659 and #9661 made the complete behavior harder to review and land as one coherent capability > - This pull request consolidates the remaining server, shared-contract, documentation, skill, and UI work on top of current master > - The benefit is a single reviewable change that lets agents safely archive responsible-user inbox items and lets users control or undo that behavior ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting inbox management across shared contracts, server authorization/routes/services, shipped agent skills, and the board UI. ### Problem or motivation Agents may complete work whose issue remains in the responsible user's Mine inbox. Existing board-user archive behavior does not provide the agent-facing policy endpoints, target resolution, heartbeat-run attribution, typed denials, conservative workflow guidance, or UI needed for safe agent-managed cleanup. ### Proposed solution Allow authorized agents to archive or unarchive responsible-user inbox items under the user's open, allowlist, or disabled policy; preserve actor/agent/run attribution in issue detail and activity records; expose policy controls and agent archive attribution in the UI; and document conservative cleanup rules for agents and PR gardening. ### Alternatives considered - Reuse generic issue mutation permissions: rejected because inbox state belongs to a target user and requires user-scoped authorization. - Automatically archive every completed or closed item: rejected because completion signals can still require human review or a decision. - Keep the backend and UI as separate stacked PRs: superseded by this consolidated PR so the complete user-visible behavior can be reviewed and verified together. ### Related work - Builds on merged foundations #9654 and #9658. - Supersedes the remaining stacked changes in #9659 and #9661. - `ROADMAP.md` has no overlapping inbox archive or inbox authorization initiative. ## What Changed - Added shared inbox-agent policy types and validators plus company-scoped self-service policy routes and OpenAPI coverage. - Enabled agent archive/unarchive mutations with responsible-user targeting, policy enforcement, typed failures, attribution, idempotency, and detailed activity auditing. - Returned agent archive attribution in issue detail and documented reversible inbox cleanup semantics in the implementation spec and Paperclip skill. - Added conservative PR-gardening inbox tidy guidance that keeps GitHub access read-only and avoids archiving work that still needs human action. - Added the Profile settings policy control and Issue Properties attribution/unarchive UI with focused component coverage and narrow-pane handling. ## Verification - `pnpm exec vitest run server/src/__tests__/inbox-archive-routes.test.ts server/src/__tests__/inbox-agent-policy-routes.test.ts server/src/__tests__/authorization-service.test.ts server/src/__tests__/openapi-routes.test.ts ui/src/components/InboxAgentPolicyControl.test.tsx ui/src/components/IssueProperties.test.tsx` — 110 passed. - `pnpm --filter @paperclipai/db exec vitest run src/inbox-archive-agent-policies-migration.test.ts` — 1 passed. - `node --test .agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` — 9 passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — changed files are clean; the repository-wide command currently reports nine unrelated pre-existing `#9627` literals outside this PR's diff. ## Risks - Agent inbox mutations broaden an existing endpoint path, so authorization and target resolution must remain fail-closed; focused route and authorization tests cover allowed and denied paths. - Archive state affects only the responsible user's inbox presentation and remains reversible; it does not mutate issue status, assignment, or visibility. - The UI policy defaults to the existing open behavior, while allowlist and disabled modes can reduce agent access. - This PR intentionally builds on #9654 and #9658 and contains no new migration number or modification to an already-applied migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using GPT-5.4, medium reasoning, repository tool use, shell execution, code review, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`feat/inbox-agent-archive-complete`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
992389480a |
fix(server): restore hot-restart run adoption (#9647)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local heartbeat/runtime subsystem starts long-running local agent processes and records their run state. > - Operators sometimes need to rebuild and restart the Paperclip server while local agent processes are still alive. > - A normal restart should remain conservative, but a guarded production hot restart needs an explicit marker, startup reconciliation, and an inspectable report. > - The broader hot-restart PR is currently merge-conflicted, so this pull request lands the minimal server-side recovery path on current `master`. > - The benefit is that deploy operators can restart from a current branch without reverting production changes and without marking adopted live runs as `process_lost`. ## Linked Issues or Issue Description No public GitHub issue exists for this deploy-safety fix. Bug fix: - What happened: the current deployable `master` branch did not include the hot-restart marker CLI, startup adoption report path, or health version proof needed by guarded service restarts. - Expected behavior: a deploy operator can write a one-shot marker before restarting, the old server snapshots eligible running child processes, the new server reports adopted/finalized/lost runs, and adopted live runs are not reaped as `process_lost`. - Steps to reproduce: restart a server with running local child-process heartbeat runs without the marker/adoption path; startup orphan reaping has no adoption metadata and treats live detached children as lost. - Paperclip version/commit: fixed on top of `master` at `b606869a6`. - Deployment mode: production/local-service style deployments that rebuild and restart the primary `paperclip.service`. - Related PR: Refs #9628. This PR intentionally lands a smaller deploy-safe subset because #9628 is currently merge-conflicted. - Duplicate search: searched public PRs/issues for `hot restart` and `process_lost adoption`; #9628 is the directly related prior implementation. ## What Changed - Added `scripts/request-hot-restart.ts` to write a one-shot hot-restart intent marker under `PAPERCLIP_HOME`. - Added `server/src/services/hot-restart.ts` for intent/report path resolution, parsing, atomic writes, shutdown snapshots, and marker cleanup. - Wired server shutdown/startup so explicit hot restarts snapshot active runs, skip the normal heartbeat drain, reconcile live child processes on boot, and write `hot-restart-report.json`. - Preserved adopted run metadata so normal orphan reaping does not regress adopted live runs to `process_lost`. - Added `serverVersion` health proof alongside existing `version`, plus docs and regression coverage. ## Verification - `pnpm vitest run server/src/__tests__/health.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 2 files passed, 100 tests passed. - `pnpm --filter @paperclipai/server typecheck` - `env PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/hot-restart-cli-smoke" pnpm --filter @paperclipai/server exec tsx ../scripts/request-hot-restart.ts --server-pid 12345` - Branch ancestry checked after `git fetch origin master`: `origin/master` was `b606869a6`, and `HEAD..origin/master` was empty. ## Risks - Medium risk: process adoption depends on PID/PGID metadata and the service manager leaving child processes alive for the guarded restart. - Normal restarts remain conservative, but an incorrect marker PID intentionally falls back to graceful drain instead of adoption. - The PR is server-only and does not include the broader UI/experimental-setting work from #9628. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 via Codex coding agent in a Paperclip execution workspace; tool use and shell/code execution enabled; context window not surfaced by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4f9894df44 |
fix(server): bound accepted-interaction continuation recovery (#9656)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Heartbeat recovery keeps assigned issues moving when a run or continuation path disappears > - Accepted issue-thread interactions can create a continuation wake after an agent previously parked for review > - The recovery sweep could requeue that accepted-interaction wake while the queued-run gate cancelled it using the older pre-acceptance park summary > - That cancellation path had no bound, so recovery could repeat the same wake and cancellation indefinitely > - This pull request makes accepted-interaction evidence supersede the older park and caps repeated recovery cancellations at three attempts > - The benefit is that accepted work resumes normally, while genuine repeated failures become a visible dependency wait or escalation instead of a cancel loop ## Linked Issues or Issue Description Refs #9331 The accepted-interaction continuation recovery added by #9331 can encounter a stale continuation summary written before approval. The sweep requeues a continuation carrying the accepted interaction timestamp, but queued-run invalidation cancels it because the older summary says to wait for review. Recovery then sees the accepted interaction without a successful run and requeues again. This PR prevents that stale-summary cancellation and adds a bounded fallback if three equivalent cancellations have already occurred. ## What Changed - Let queued continuation wakes with a parseable `interactionResolvedAt` bypass a pre-acceptance waiting-for-review park summary. - Count consecutive unsuccessful continuation runs for the same issue and agent since interaction acceptance; after three review-park cancellations, convert a real dependency wait or use the existing visible escalation path. - Add focused regression coverage for the park bypass, unchanged non-interaction park behavior, below-cap requeue, cap escalation, and successful-run skip. - Document the accepted-interaction precedence and bounded requeue contract in execution semantics §9.2. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "accepted interaction continuation recovery|accepted interaction recovery after its continuation succeeds|requeues accepted interaction continuations stranded"` - `pnpm vitest run server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts -t "pre-acceptance review park|continuation summary parks executor work"` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Low risk: the park bypass only applies when the queued context contains a parseable interaction resolution timestamp. - The retry bound is scoped to unsuccessful `issue_continuation_needed` runs for the same company, issue, agent, error code, and post-acceptance time window. - No schema, migration, API, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.5`, high-reasoning coding mode with repository tool use and command execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea0e899905 |
fix(search): honor extract match limits + harden pr-gardening candidate discovery (#9652)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `/pr-gardening` skill drives a bundled agent that scans a
company's issues for those linked to open GitHub PRs, then reports on
their state; it relies on the server's company-search **extract**
endpoint to pull PR references out of issue bodies
> - Two gaps surfaced during end-to-end QA of the gardening workflow:
the extract service silently ignored a per-issue match cap, so callers
could not bound how many matches came back per issue, and the skill's
candidate-discovery scripts fell over on large repos and on issues that
referenced deleted PRs
> - Left unaddressed, the gardener either truncated its scan
unpredictably or aborted outright, so it could not reliably enumerate PR
candidates
> - This pull request honors an explicit `matchesPerIssue` limit in the
extract search API and hardens the skill's candidate discovery against
missing/unavailable PRs and oversized `gh` output
> - The benefit is a PR-gardening workflow that scans deterministically
and finishes cleanly on real-world companies
## Linked Issues or Issue Description
No pre-existing public GitHub issue — describing the bug in-PR following
the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`).
### What happened?
The company-search extract endpoint accepted a per-issue match limit but
did not apply it, returning matches capped only by the old hardcoded
constant regardless of the caller's request. Separately, the
`/pr-gardening` skill's candidate-discovery scripts crashed when a
scanned issue referenced a deleted PR (GitHub `Not Found (HTTP 404)` /
GraphQL `Could not resolve to a PullRequest`) and could exceed the
default `gh` output buffer on large result sets, aborting the whole
scan.
### Expected behavior
The extract API bounds matches per issue when a caller passes
`matchesPerIssue` (default 20, max 200), and omitting it preserves the
previous default. The gardening scripts skip PRs that are
deleted/unavailable and tolerate large `gh` responses without aborting
the scan.
### Steps to reproduce
1. Call the company-search extract endpoint with a `matchesPerIssue`
value against an issue containing many PR references — previously the
value was ignored.
2. Run the pr-gardening candidate scan against a company whose issues
reference a since-deleted PR — previously the scan threw instead of
skipping that PR.
### Paperclip version or commit
`master` at the base of this PR (branch cut from current
`origin/master`).
### Deployment mode
Local Paperclip instance / self-hosted.
## What Changed
- **Extract search honors `matchesPerIssue`**: added the
`matchesPerIssue` field to the shared search validator/types and applied
the cap in `company-search-extract` so results are bounded per issue
(`packages/shared`, `server/src/services/company-search-extract.ts`,
`doc/SPEC-implementation.md`).
- **Hardened pr-gardening candidate discovery**: `find-candidates.mjs` /
`lib.mjs` now request `matchesPerIssue=200`, treat missing/unavailable
PRs (deleted PR → `isMissingPullRequestError` / `unavailable`) as skips
instead of fatal errors, and raise the `gh` `maxBuffer` to 50 MB for
large repos.
- **Tests**: expanded `company-search-extract-{routes,service}.test.ts`
for the new limit and added coverage in `pr-gardening.test.mjs`.
## Verification
Re-run on a fresh worktree cherry-picked onto current `master`:
- `node --test
.agents/skills/pr-gardening/scripts/pr-gardening.test.mjs` → 8/8 pass
- `pnpm vitest run
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-extract-service.test.ts` → 10/10
pass
## Risks
Low risk. `matchesPerIssue` is optional and backward-compatible
(omitting it preserves prior behavior). The skill changes only add
skip/tolerance paths and a larger buffer; no schema or migration
changes.
## Model Used
Claude — Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use /
code execution via the Claude Agent SDK.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
d32ed88443 |
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies and their work. > - Its recovery subsystem detects stranded issue execution and decides whether to retry, escalate, or request operator intervention. > - The existing recovery path used a mostly generic owner ladder and generic execution contract, so transient failures could wake a manager who then performed the deliverable instead of repairing and returning the task. > - Provider quota failures also entered the same takeover path even when the correct action was to wait for capacity and retry the original assignee. > - Recovery actions already retain the source owner and evidence needed to choose a cause-specific route, render a scoped contract, and measure whether work was handed back. > - This pull request adds a cause-keyed recovery playbook, propagates its contract through every built-in adapter, and makes resolved recovery actions return work to the original owner by default. > - The benefit is bounded self-recovery that preserves task ownership, avoids needless management takeover, and makes recovery outcomes observable. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related recovery work was reviewed but is not duplicated here: #9630 restores bounded recovery continuations, #8807 changes one assignee-ranking case, and #9404 records runtime-failure transition evidence. This change instead introduces cause-specific routing and recovery contracts across the recovery lifecycle. ### What happened? When an issue became stranded, recovery generally selected an owner through the same fallback ladder and rendered the normal execution contract. That made the recovery wake look like ordinary deliverable work, even when the correct action was to retry the original agent, repair its runtime, or wait for a provider quota reset. ### Expected behavior Recovery should select a response by failure cause, tell the recipient to recover rather than complete the deliverable, suppress takeover wakes for provider quota waits, and return repaired work to its original assignee unless the recovery owner explicitly completes it. ### Actual behavior Recovery could escalate transient failures to management, omit the cause-specific next action from the wake, and leave the recovery owner assigned after the runtime problem was resolved. ### Impact The generic path creates avoidable management work, ownership churn, and budget consumption while obscuring whether recovery successfully returned work to the responsible agent. ## What Changed - Added cause-keyed routing for process loss, missing disposition, provider quota limits, Codex output inactivity, workspace validation failures, and fallback recovery causes. - Added recovery-scoped wake rendering that replaces the generic execution contract with the failure summary, original assignee, attempt count, next action, and cause-specific playbook instruction. - Propagated the structured recovery contract through all built-in adapter execution paths, including Hermes local and gateway adapters. - Added provider-quota wait monitoring so capacity failures schedule the original assignee instead of enqueueing a takeover wake. - Added hand-back behavior and `handed_back` / `owner_completed` outcome accounting when recovery actions are resolved. - Added focused routing, renderer, quota-monitor, and hand-back regression coverage plus implementation-spec documentation. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/issue-recovery-actions.test.ts` - 4 test files passed; 194 tests passed. - Targeted `pnpm --filter ... typecheck` across `@paperclipai/adapter-utils`, `@paperclipai/shared`, `@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter packages. - 13 affected workspace packages passed typecheck. - `pnpm check:token-gates` - All UI token gates passed. ## Risks - Recovery routing behavior changes for stranded work, so an incorrectly classified cause could select a different recipient than before; fallback causes retain the existing management ladder. - Provider quota detection depends on structured failure evidence and conservative text matching; unmatched failures continue through fallback recovery. - Adapter prompt plumbing changes across built-ins, covered by shared renderer tests and compile-time call signatures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae77908618 |
feat(search): add bulk extract endpoint (#9507)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents and operators need company-scoped search to discover relevant issue history safely > - The interactive search endpoint intentionally returns compact excerpts and low pagination caps for UI use > - Automation that inventories repeated references, such as pull-request URLs, needs exhaustive distinct matches without loading full issue objects into an LLM context > - Client-provided regular expressions would create an unsafe and expensive query surface, so extraction must remain literal with server-owned expansion modes > - This pull request adds a bounded agent-oriented extraction endpoint with explicit truncation > - The benefit is deterministic, compact bulk discovery across issues, comments, and documents while preserving company authorization and rate limits ## Linked Issues or Issue Description ### Subsystem affected `server/` REST API and `packages/shared/` contracts. ### Problem or motivation The existing interactive company search caps issue pagination and snippets, so automation cannot reliably enumerate every distinct literal or pull-request URL across issue descriptions, comments, and linked documents without fetching large full issue payloads. ### Proposed solution Add `GET /api/companies/:companyId/search/extract` with escaped literal matching, optional server-owned URL token expansion, issue/comment/document scopes, status/date filters, higher issue-level pagination caps, compact source references, and explicit pagination/match truncation flags. ### Alternatives considered Reusing `GET /issues?q=` would return unnecessarily large issue objects; increasing interactive-search snippet limits would make the UI API heavier; accepting arbitrary client regex would expose avoidable database cost and ReDoS risk. ### Roadmap alignment `ROADMAP.md` does not currently list a conflicting company-search or bulk-extraction initiative. GitHub searches found no directly duplicative open issue or pull request. ## What Changed - Added shared query validation and response contracts for literal and URL extraction. - Added a company-scoped extraction service that pages issues, gathers matching issue/comment/document sources, expands URL tokens, deduplicates values, and reports truncation explicitly. - Added the authenticated route using the existing company-search authorization decision and rate limiter. - Added targeted Vitest coverage for URL extraction, multi-source dedupe, date/status filters, match caps, cross-company denial, and rate limiting. - Documented the extraction surface in the implementation specification. ## Verification - `pnpm exec vitest run server/src/__tests__/company-search-extract-service.test.ts server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-rate-limit-routes.test.ts server/src/__tests__/company-search-service.test.ts` — 30 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Bulk substring search can scan large text columns. The endpoint mitigates this with a minimum literal length, bounded issue pagination, a 20-distinct-match cap per issue, explicit truncation, existing company-search rate limiting, and no client-provided regex. - URL expansion uses a fixed server-owned pattern plus an escaped literal. A security review is requested as part of PR review to confirm the pattern and abuse controls. - No database migration or existing API response shape changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact runtime model ID and context-window size were not exposed to the session. Tool-enabled code execution and repository editing were used with medium reasoning effort. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9af96461d5 |
fix(server): restore stranded recovery continuations (#9630)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work. > - Its server recovery layer classifies blocked issue graphs and restores interrupted heartbeat execution. > - A dependent issue could remain dispatch-suppressed by a cancelled blocker without producing operator-visible attention when the dependent still displayed as todo or backlog. > - Separately, a monitor-triggered run that lost its process before disposition could consume the monitor's one-shot wake without scheduling the existing bounded continuation. > - Both gaps strand useful work even though Paperclip already has the relevant blocker-attention and process-loss recovery mechanisms. > - This pull request widens the existing classification path and reuses the single process-loss retry for monitor dispatches with no future wake. > - The benefit is visible, routable recovery without weakening dependency checkout rules or introducing an unbounded retry loop. ## Linked Issues or Issue Description No matching public GitHub issue or pull request was found. ### What happened? Two server recovery cases could leave work stranded: 1. A non-terminal, agent-assigned issue with an unresolved cancelled blocker remained ineligible for checkout, but blocked-chain liveness classification only inspected issues already displaying `blocked` or `in_review`, so the existing `blocked_by_cancelled_issue` attention was not surfaced. 2. A one-shot issue monitor cleared its next check when dispatched. If that monitor-triggered run ended as `process_lost` without a tracked local child, the existing bounded retry gate rejected it and no future monitor wake remained. ### Expected behavior - Cancelled blockers continue to be unresolved dependencies, and their dependents receive blocker attention regardless of whether the dependent currently displays as backlog, todo, blocked, or in review. - A monitor-triggered run lost before disposition receives exactly one bounded continuation when no future monitor check exists; a second loss follows the normal recovery-action escalation path. ### Steps to reproduce 1. Create an agent-assigned todo issue blocked by a cancelled issue and run issue-graph liveness classification. 2. Observe that no cancelled-blocker finding appears before this change. 3. Dispatch a due issue monitor, clear its one-shot `monitorNextCheckAt`, and mark the resulting untracked run `process_lost`. 4. Observe that no retry is queued before this change. ### Environment - Paperclip commit: `3e348b96b` - Deployment: built from source / local test environment - Adapter: not adapter-specific; core server recovery - Database: embedded test database ## What Changed - Inspect non-terminal, agent-assigned issues with unresolved blocker edges during blocked-chain liveness classification. - Include cancelled dependents in the existing blocked-inbox attention query while preserving company-scoped relation checks. - Allow monitor-triggered `process_lost` runs with no future monitor wake to use the existing single bounded retry. - Mark monitor recovery retries as continuation-needed context and retain the existing second-loss escalation behavior. - Document cancelled-blocker and monitor-dispatch recovery semantics. - Add focused regressions for liveness findings, attention propagation, one retry, and second-loss escalation. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/issue-blocker-attention.test.ts server/src/__tests__/issue-liveness.test.ts` — 3 files, 126 tests passed. - `pnpm --filter @paperclipai/server typecheck` — passed. ## Risks - Low risk and server-only. The liveness scan inspects more unresolved dependency shapes, which can produce additional existing attention entries for previously invisible cancelled blockers. - Monitor recovery remains bounded by `processLossRetryCount < 1`, and the extra path only applies when the dispatch was monitor-triggered and no future monitor check exists. - No schema, migration, authorization, API-contract, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.4` through Codex CLI, with reasoning, repository tool use, command execution, and test execution capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
89ce36d7af |
feat(skills): open-by-default company skill policy and core UX (#9564)
## Thinking Path > - Paperclip uses company skills to make agent capabilities reusable across an organization. > - Skill operations currently mix capability availability with permission checks, which creates avoidable setup friction and inconsistent denial handling. > - The policy contract needs to remain open by default while allowing company-scoped restrictions for governed deployments. > - Core owns the canonical policy actions, persistence, evaluation, API behavior, safe import boundaries, and generic denial/read-only UI. > - Enterprise policy-editor implementation belongs in the separate `paperclip-ee` repository and is intentionally excluded from this PR. ### Problem or motivation Company skill operations can encounter permission dead ends even when no explicit restriction has been configured, and import-source classification can drift between policy evaluation and execution. ### Proposed solution Define eight canonical skill policy actions, default all actions to allowed, persist company-scoped restrictions, expose policy evaluation APIs, normalize import sources at the boundary, and update Skill Studio to present actionable restriction states without embedding Enterprise Edition implementation in the core repository. ### Alternatives considered Keeping capability checks distributed across routes and UI surfaces was rejected because it duplicates policy logic and makes denial behavior inconsistent. Shipping the Enterprise policy editor in this repository was rejected because `paperclip-ee` is a separate repository and must receive its own PR. ### Roadmap alignment Extends the completed **Skills Manager** roadmap area by adding coherent governance and removing workflow dead ends. ### Additional context The core API contract remains suitable for a separate Enterprise Edition editor, but this PR contains no `paperclip-ee` package or EE-specific UI integration code. ## What Changed - Added the company skill policy contract to product and implementation documentation, including the open-by-default rule, eight canonical actions, decision shape, and core/EE ownership boundary. - Added the company-scoped policy schema, migration `0170`, shared validators, policy service, REST routes, OpenAPI coverage, and focused server tests. - Hardened import policy enforcement by normalizing import sources and keeping source classification consistent between policy evaluation and execution. - Updated core Skill Studio behavior to remove generic permission dead ends and show actionable policy/platform denial states only when an operation is actually denied. - Removed the `plugin-paperclip-ee` package, Docker wiring, EE discovery/deep-link helpers, and EE-specific UI tests/stories from this PR so that implementation can be submitted separately to the EE repository. - Preserved open-by-default behavior when no explicit company restriction exists. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/skill-studio/SkillPolicySurfaces.test.tsx src/lib/skill-policy-denial.test.ts` — 20/20 passed. - `pnpm --filter @paperclipai/ui exec tsc --noEmit` — passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/worktree-config.test.ts` — 12/12 passed. - `pnpm check:token-gates` — passed with all gates clean. - `git diff --check` — passed. - `git diff --name-only origin/master | rg 'paperclip-ee|ee-skill-policy'` — no matches. ## Risks - Migration `0170` introduces company policy persistence; rollout depends on the migration applying before policy routes are exercised. - Open-by-default is an intentional behavioral policy: deployments expecting implicit denials must configure explicit restrictions. - Import normalization is security-sensitive and should retain focused review. - The separate EE editor must stay contract-compatible with the core policy API as policy actions evolve. ## Model Used - OpenAI Codex CLI, runtime model identifier and context-window size not exposed by this execution environment; reasoning, repository tool use, shell execution, and code review capabilities enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details available to this runtime) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked or described the result above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green on the latest head - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Evyatar Bluzer <bluzername@users.noreply.github.com> |
||
|
|
3db2e6bdd2 |
feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f712cedbc2 |
feat(mcp) [split 7/8]: activate Apps and gateway UI (#9562)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 7/8 and focuses on Apps/Gateways UI and feature-flagged activation > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The standalone UI foundation needs feature-flagged routes, navigation, app detail flows, gateways, Storybook scenarios, and QA configuration. - Proposed solution: Adds Apps/Gateways pages and components, navigation/route activation, remaining page integrations, Storybook stories, and QA Vite configuration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/06-ui-tools-foundation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: UXDesigner sanity pass on Apps/Gateways flows and flag-off behavior; Greptile on every PR. ## What Changed - Adds Apps/Gateways pages and components, navigation/route activation, remaining page integrations, Storybook stories, and QA Vite configuration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `pnpm check:token-gates` — all gates clean - Focused UI Vitest run with `NODE_ENV=test` — 17 files, 147 tests passed ## Risks - Navigation or flag regressions could expose incomplete experiences; activation remains controlled by existing experimental settings. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. ## UI Evidence QA captured these from the live Garden MCP split stack at 1440px and verified clean rendering:    --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a116713b9a |
feat(mcp) [split 6/8]: add Tools and Profiles UI foundation (#9561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 6/8 and focuses on UI API, shared components, and Tools/Profile surfaces > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: Operators need typed clients and administration surfaces that compile independently before navigation exposes them. - Proposed solution: Adds UI APIs, hooks, libraries, shared components, Tools/Profiles pages, and the plugin settings consumer required by the new company-scoped API. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/05-runtime-integration`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: UXDesigner sanity pass on Tools/Profile surfaces; Greptile on every PR. ## What Changed - Adds UI APIs, hooks, libraries, shared components, Tools/Profiles pages, and the plugin settings consumer required by the new company-scoped API. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `pnpm check:token-gates` — all gates clean - Focused UI Vitest run with `NODE_ENV=test` — 28 files, 183 tests passed ## Risks - Large dead-code UI additions can drift from activation routes; PR 7 supplies the registration layer and top-level parity catches omissions. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. ## UI Evidence QA captured these from the live Garden MCP split stack at 1440px and verified clean rendering:    --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
98d9360658 |
feat(adapters): confine local coding processes (#9504)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work. > - Local coding adapters currently spawn their CLI processes directly on the Paperclip host. > - CLI-native approval and sandbox flags do not provide a reliable host filesystem or network boundary. > - An agent can therefore inspect unrelated host files or fetch external material when an operator needs stronger isolation. > - The confinement must stay opt-in so existing local adapter behavior does not change unexpectedly. > - This pull request adds a shared Linux Bubblewrap spawn layer for workspace filesystem and deny/allowlist network scopes. > - The benefit is enforceable defense in depth around Codex and Claude local runs while preserving explicit provider connectivity. ## Linked Issues or Issue Description ### What happened? `codex_local` and `claude_local` processes could read arbitrary host paths and make unrestricted outbound network requests because Paperclip did not impose a spawn-level boundary. ### Expected behavior Operators can opt into a workspace-only filesystem view and either deny network egress or allow exact provider/API hosts, independently of CLI approval flags. ### Steps to reproduce 1. Run current `master` on Linux and configure a Codex or Claude local adapter. 2. Ask the agent to read a canary file outside its active workspace. 3. Ask the agent to `curl` a public host. 4. Observe that both operations succeed without a Paperclip-level confinement option. ### Environment - Paperclip commit: `c36f1a4af` / current `master` base. - Deployment mode: Linux local dev or self-hosted server. - Installation: built from source. - Adapters: Codex and Claude Code. - Database: not related. ## What Changed - Added a shared Bubblewrap process wrapper with opt-in `filesystemScope: "workspace"`, managed/extra path mounts, private `/tmp`, and Linux-only validation. - Added `networkScope: "deny" | "allowlist"`; both use a private network namespace, while allowlist mode exposes an exact-host HTTP(S) proxy over a Unix-socket bridge. - Wired Codex and Claude local CLI execution through the wrapper and forced scoped auto runs onto the CLI lane because ACP processes are not covered. - Added unit and gated Bubblewrap canaries for outside-file denial, workspace writes, direct network denial, allowlisted forwarding, and rejected destinations. - Documented both scopes, provider allowlist examples, Bubblewrap requirements, and default-off behavior. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts` — 34 passed, 4 gated Bubblewrap tests skipped by default. - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-utils build` - `pnpm --filter @paperclipai/adapter-codex-local build` - `pnpm --filter @paperclipai/adapter-claude-local build` - Attempted the gated tests with a vendored Bubblewrap binary; this container blocks unprivileged namespace setup (`setting up uid map: Permission denied` / loopback `RTM_NEWADDR: Operation not permitted`), so kernel-level execution remains for CI or a namespace-enabled Linux host. ## Risks - Bubblewrap must be installed and unprivileged user/mount/network namespaces must be enabled on the host; scoped runs fail clearly if the prerequisite is missing. - Allowlist mode depends on the coding CLI honoring standard `HTTP_PROXY` / `HTTPS_PROXY` variables; custom providers must list every required exact hostname and port. - Exact-host allowlists intentionally reject wildcards, which is safer but may require operators to enumerate multi-host provider setups. - No behavior changes unless an operator enables `filesystemScope` or `networkScope`. > This aligns with the ROADMAP direction toward safer remote and sandboxed agent environments and does not duplicate an open PR or issue found in the repository search. ## Model Used - OpenAI GPT-5.5 (`gpt-5.5`) via Codex CLI, with reasoning, repository tool use, shell execution, code editing, and test execution. The serving context-window size is not exposed to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
90f85a7d11 |
Add telemetry proposal extractor (#9544)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - It emits telemetry events to understand product usage — registered event names are gated by a generated `PAPERCLIP_EVENTS` registry so the client only enqueues known, schema-approved events > - When product teams want to instrument a new behaviour, they must first register the event name — but schema registration is a commit-and-release cycle, which creates friction in fast-moving product iterations > - A proposal lane is needed: let developers mark a `track()` call with a typed `@ts-expect-error` proposal marker so the event name can be reviewed and tracked in CI before the schema is formally registered > - The existing client had no guard against unregistered event names, so any call with an out-of-registry name (or a prototype-inherited key) would silently enter the queue, state, and network flush path > - This PR adds an `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard at the entry point of `track()` to swallow unregistered calls before any side effects, adds `scripts/extract-proposed-events.mjs` to scan source for proposal markers and emit a v2 JSON manifest with provenance and rationale, and documents the complete proposal workflow > - The benefit is that new instrumentation can be proposed and reviewed in code without touching the registered schema, and tooling can surface missing rationale before events graduate to stable ## Linked Issues or Issue Description No existing GitHub issue covers this change. This PR introduces a new feature. **Feature motivation:** Paperclip's telemetry schema is intentionally stable — registered event names are code-generated and gated. Product engineers who want to instrument a new behaviour today must land a schema change first, creating a two-step process that slows iteration. A proposal lane lets developers write the instrumentation call ahead of schema registration, protected by a compile-time `@ts-expect-error` marker that an extractor script can surface for review. This PR implements both the client-side safety gate and the extraction tooling. Refs: #9518 (closed predecessor — docs-only; this PR supersedes it with the full implementation) ## What Changed - Added `Object.hasOwn(PAPERCLIP_EVENTS, eventName)` guard at the top of `TelemetryClient.track()`: unregistered event names (including prototype-inherited keys) are now swallowed before any state, queue, or network operation - Added `scripts/extract-proposed-events.mjs`: scans TypeScript source for `@ts-expect-error -- proposed-telemetry(<issue>): <rationale>` markers; emits a v2 JSON manifest per proposed event including name, rationale, provenance (repo-relative file + line), and a `rationale_missing` flag for CI enforcement - Added `scripts/extract-proposed-events.test.mjs`: test suite covering marker parsing, multi-line markers, path validation, out-of-repo rejection, and the v2 schema output contract - Added `doc/TELEMETRY_WORKFLOW.md`: documents the proposal workflow, the canonical multi-line marker example, rationale requirements, and how to graduate a proposed event to stable schema - Updated `packages/shared/src/telemetry/README.md`: added "Proposed Events" section to the Telemetry Data Contract per the contributing guide requirement for telemetry changes ## Verification Run all of the following from the repo root: ```sh # Extractor unit tests node --test scripts/extract-proposed-events.test.mjs # Telemetry client + types tests pnpm exec vitest run --config vitest.config.ts \ src/telemetry/client.test.ts src/telemetry/client-types.test.ts \ --reporter=verbose # (run from packages/shared) # Type-check pnpm --filter @paperclipai/shared typecheck # Smoke-run the extractor in local-test mode node scripts/extract-proposed-events.mjs --ref local-test ``` All four commands pass locally. ## Risks - **Silent drop on unregistered events:** The `Object.hasOwn` guard fails closed — any event name not in `PAPERCLIP_EVENTS` is silently dropped. If the generated registry is missing an event that was previously tracked, those calls will be silently lost. Mitigation: the extractor script surfaces proposed events that need registration; the TypeScript type system already enforces `TelemetryEventName ⊆ PAPERCLIP_EVENTS` at compile time. - **Extractor is read-only:** `extract-proposed-events.mjs` reads source and emits JSON; it does not modify any files. No runtime or schema risk. - Overall risk: **low**. The guard is additive and defensive; the extractor and docs are additive only. ## Model Used - Provider: Anthropic - Model ID: `claude-sonnet-4-6` - Context window: 200 K tokens - Capabilities: tool use, extended context, code generation ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8b6a06ee25 |
[codex] Add built-in agents and Reflection Coach bundle (#9206)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators need first-party agent capabilities for repeatable company work, not just manually created one-off agents. > - Built-in agents need to behave like normal company-scoped agents while preserving approval gates, permissions, budgets, and audit trails. > - Reflection and coaching work also needs bundled instructions, skill content, and a routine so the feature can be installed and reset predictably. > - The API, database, UI, portability, and tests all need to agree on the built-in lifecycle from not provisioned through setup, approval, ready, paused, and reset. > - This pull request adds built-in agent provisioning and the Reflection Coach bundle end-to-end. > - The benefit is a safer first-party path for Paperclip-managed agents without bypassing the same governance model used for operator-created agents. ## Linked Issues or Issue Description No public GitHub issue was found for this exact built-in agent and Reflection Coach bundle work. Problem/motivation: - Paperclip did not have a first-party built-in agent lifecycle for product-owned agents. - Bundled agent resources such as default instructions, skills, and routines needed managed ownership and reset semantics. - Approval-gated companies needed built-in setup to preserve requested adapter, budget, manager, and permission state through board approval. - The board UI needed clear built-in badges, setup affordances, readiness state, and bundle status without exposing secrets. Proposed solution: - Add a company-scoped built-in agent registry, provisioning/reset/reconcile/status APIs, and Reflection Coach bundled resources. - Track bundled managed resources in the database with idempotent migration behavior. - Reuse existing agent approval, authorization, budget, and activity-log paths instead of creating a bypass. - Add UI setup, badges, gates, bundle panels, and route coverage for built-in agents. Duplicate search: - Searched GitHub PRs for `built-in agents Reflection Coach repo:paperclipai/paperclip`; only this PR was returned. - Searched GitHub issues for the same query; no public issues were returned. ## What Changed - Added built-in agent definitions, lifecycle state derivation, provisioning, reset, reconcile, status, and routine-control routes. - Added the `built_in_managed_resources` migration and schema exports for bundled instructions, skill, and routine ownership. - Added the Reflection Coach built-in bundle with default instructions, skill catalog content, routine template, default permissions, and managed-resource drift handling. - Added approval-aware provisioning behavior that preserves requested adapter config, budgets, manager assignment, and built-in permissions through hire approval. - Added authorization and mutation gates for built-in agent and skill changes, including consented Reflection Coach change paths. - Added UI surfaces for built-in agent setup, roster/detail badges, readiness gates, bundle status, routine controls, and route filtering. - Added company import/export and validator coverage for built-in managed resources and low-trust/red-team presets. - Addressed Greptile follow-ups for pending approval reconciliation, consent-gate error propagation, config-read authorization fallback, approval-path manager preservation, and non-model adapter provisioning. ## Verification Local verification: - `git diff --check public/master..HEAD` passed. - `pnpm check:token-gates` passed with all gates clean. - `pnpm exec vitest run ui/src/components/ConfigureBuiltInAgentModal.test.tsx` passed: 1 file, 4 tests. - `pnpm exec vitest run ui/src/components/EntityRow.test.tsx ui/src/pages/Agents.test.tsx ui/src/components/BuiltInAgentGate.test.tsx ui/src/components/ConfigureBuiltInAgentModal.test.tsx ui/src/components/BuiltInBundlePanel.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx ui/src/pages/Routines.test.tsx` passed: 7 files, 64 tests. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts src/__tests__/authorization-service.test.ts src/__tests__/company-skills-routes.test.ts` passed: 3 files, 91 tests. - `pnpm --filter @paperclipai/db check:migrations` passed. - `pnpm -r typecheck` passed after the rebase; `pnpm --filter ui typecheck` passed after the final UI review fix. Remote verification on latest head `1c61f693a4ec881d739022b0e75a8ca8bf8c2cd8`: - Merge state: `CLEAN`. - Greptile: `5/5`, zero unresolved Greptile threads. - PR check rollup: all checks successful, neutral, or skipped as expected. - Passing gates include Build, Typecheck + Release Registry, all server shards, all workspace shards, all serialized server suites, e2e, Canary Dry Run, policy, review, verify, Socket, Superagent, and Snyk. ## Risks - This adds a new managed-resource table and migration; the migration uses idempotent create/add/index guards and passed migration safety checks. - Built-in agent provisioning touches approval and authorization paths; tests cover pending approval preservation, stale retry rejection, consent gates, and config-read fallback behavior. - Reflection Coach creates managed instructions, skill, and routine resources; drift/reset behavior is covered by service tests and redacted API responses. - Non-model adapter setup now provisions a `needs_setup` built-in row before command/endpoint fields are complete; this matches the server lifecycle and is covered by the setup modal regression test. ## Model Used OpenAI Codex coding agent based on GPT-5. Exact hosted model ID, context-window size, and reasoning-mode labels are not exposed in this runtime; tool use, shell execution, GitHub CLI/API access, and local code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53d09d4c34 |
fix(ui): inbox/task list parity, nesting alignment, hover perf, and routine detail polish (#9317)
## Thinking Path > - Paperclip's UI is governed by the design system merged in #9134 and the component convergence in #9240 — one Card, one Badge, one nav row, one `IssueRow`, a single multiplicative radius ladder. > - With those primitives in place, the remaining rough edges were interaction and alignment details on the surfaces people use every day: the inbox, the task list, the sidebar, and the routine/task detail pages. > - Each item here was reported from live use and fixed against a running instance, then verified by measurement (pixel alignment, frame timing) rather than by eye alone. > - The result is that the inbox and task lists now behave as one system (same hover, keyboard nav, tree-guide, and archive language), list nesting reads correctly, list hover is smooth, and the routine detail page scrolls and aligns like the rest of the app. ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template. - **Problem**: after the design-system foundation (#9134) and component convergence (#9240), the inbox/task lists still had interaction and alignment gaps — hover lag on long lists, keyboard navigation that only partly matched between the two lists, workspace/parent nesting whose guides and chevrons didn't line up, an inbox that sat offset from the task list, and a routine detail page with an odd double-scroll and a bespoke sub-nav. - **Proposed behavior**: the inbox and task lists share one interaction contract (hover, keyboard nav, collapse, archive), list nesting aligns to the status column with clean chevrons, list hover is CSS-only (no per-hover re-render), and the routine detail page uses a fixed header/sub-nav with a single scrolling content region and a sub-nav that matches the primary nav. ## What Changed - **List hover performance**: hover is now painted purely by CSS `:hover` and records the hovered row in a ref, instead of writing list-selection state on every `mouseenter` (which re-rendered 100–300 non-memoized rows per hover). Keyboard nav reads the ref so it still continues from the hovered row; the keyboard band clears on the first real mouse move so hover and keyboard selection never show two bands at once. The row's `transition-colors` fade was removed so the highlight snaps (no comet-tail). Measured on a 100-row scrub: worst frame **333 ms → 33 ms**, long frames (>50 ms) **12 → 0**, avg **41 → 60 fps**. - **Inbox ↔ task-list parity**: keyboard navigation works on every inbox tab (archive/read keys stay scoped to the archivable tab); group headers and parent tasks collapse/expand with the arrow keys in both lists; the task list gains the same j/k / arrows / Enter selection model as the inbox; hover selection bands match; the `g` then `i` go-to-inbox chord works app-wide. - **List nesting & alignment**: the workspace group-header chevron lines up exactly with the task chevrons below it (both lists); the parent→child connector line drops from under the parent's **status icon** rather than its chevron and breaks with a 14px gap around a nested row's own chevron; and inbox rows line up with the task list (read rows no longer reserve a mark-read column). - **Inbox archive affordance**: moved from a bare `x` left of the status icon to an `Archive` icon + label button on the right (before the timestamp), revealed on row hover; the left slot now carries only the unread dot; swipe-to-archive is unchanged. - **Sidebar**: when no agent has a live run, the AGENTS section shows 3 recent agents (was 5) plus "See all agents"; the working-agents view is unchanged. - **Routine detail page**: the layout is bounded to the main scroll area so the header (Run now + automation toggle) and the sub-nav stay fixed and only the section content scrolls (was a page-level scroll competing with a `sticky` sub-nav). The sub-nav items adopt the primary nav's rhythm — row padding/height, inset rounded pill, type scale, and 16px icons — and its background matches the main nav. - **Task detail**: dropped the redundant `🔵` prefix the breadcrumb added for live/in-progress tasks; the status glyph already conveys that state. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2255/2255 (assertions updated in lockstep where behavior changed) - `pnpm --filter @paperclipai/ui build` → exit 0 - Storybook visual regression: run locally throughout (the baseline-manifest archive is still unpublished, so CI cannot run this suite — pre-existing condition from #9134). Every visible delta was reviewed against a live instance and, where it was intentional (nesting alignment, routine sub-nav restyle, unread-row shift), the affected snapshots were re-baselined locally. - Manual/measured: inbox + task lists (grouped and nested, light + dark), hover-scrub frame timing, keyboard navigation, and the routine detail scroll/alignment were exercised on a running instance. ## Risks - Behavior-and-alignment changes concentrated in `IssueRow` / `IssuesList` / `Inbox` (the shared task-row surfaces). The riskiest area — the hover/keyboard-selection model — is covered by unit tests (updated in lockstep) and was measured and driven live. - The routine detail scroll change restructures that page's layout container; verified the page's own scroll stays fixed while only the section content scrolls. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing); snapshot coverage is local-only until that lands. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright measurement/screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c90e66bdd6 |
feat(ui): design-system component convergence — Card/Badge adoption, multiplicative radius ladder, unified list surfaces (#9240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is governed by a design system (`DESIGN.md` + the token layer in `ui/src/index.css`, merged in #9134) whose first principle is "one way to say each thing" — one Card, one Badge, one nav row > - After the token extraction landed, ~35 files still hand-rolled card containers, ~45 files hand-rolled pill spans, the sidebar agents section duplicated the nav-row chrome, and the inbox and tasks lists rendered the same task rows two subtly different ways > - Each divergence is a place where a future design change (radius, hover language, status vocabulary) silently misses surfaces, defeating the "edit tokens + run checks" model the design system exists for > - This pull request converges those surfaces onto the shared primitives, codifies the radius scale as the modern multiplicative shadcn ladder, and unifies the row/hover/tree-guide language across the inbox and tasks lists — every visible delta was human-reviewed screen-by-screen against a live instance across nine feedback rounds > - The benefit is that the app's look is now steerable from single knobs (one `--radius` anchor, one Card, one Badge, one row component), and the visual regression suite covers the result (514 snapshots including a new AgentDetail page story) ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template: - **Problem**: after the design-token foundation (#9134), component-level drift remained — hand-rolled cards/pills, duplicated sidebar row chrome, and two different renderings of task rows (inbox vs tasks list) meant design changes had to be applied per-surface and frequently missed spots (e.g. status glyphs rendered 16px in the inbox but 20px in the tasks list because a slot override silently beat the component default). - **Proposed behavior**: all card-shaped containers render via `Card`, all label pills via `Badge`, sidebar rows via `SidebarNavItem`, and both task-list surfaces via one `IssueRow` configuration; the radius scale is a single multiplicative ladder anchored at `--radius: 0.5rem`. - **Alternatives considered**: converting interactive `<button>`/`<Link>` cards to `Card` divs (rejected — breaks semantics; documented inline with `design-allow` comments instead); keeping the legacy additive radius ladder (rejected in favor of the standard shadcn multiplicative mapping). ## What Changed - `Card` adoption across ~35 files (settings pages, auth/board flows, dashboards, list containers, KPI tiles); non-adoptable sites (interactive cards, `<li>` rows, class-string props, chart tooltip) carry documented `design-allow(card-pattern)` comments - `Card` gains an `interactive` prop — one quiet hover affordance for clickable cards (cursor, border darken, shadow lift, focus ring), applied to skills tiles, artifact cards, and the company selector; cards carry no resting shadow - `Badge` adoption for 113 hand-rolled pill spans across ~45 files; `PropertyChip` wraps `Badge` internally; `StatusBadge`, external-object chips, and match chips stay bespoke by documented decision (WCAG-tuned status mechanics) - Radius ladder becomes the multiplicative shadcn mapping (`sm/md/lg/xl/2xl/3xl/4xl = 0.6/0.8/1.0/1.4/1.8/2.2/2.6 × --radius`, anchor `0.5rem`); every card surface unifies on `rounded-lg`; the orphaned 8px literal token is deleted - Sidebar: agent rows render via `SidebarNavItem` (new additive props: `iconNode`, `active`, `trailing`, `liveAccessory`); live dots use `--status-agent-running`; one row rhythm and inset rounded pill highlight; right-aligned trailing badges; every labeled section is collapsible - Inbox + tasks lists unified: md status glyphs, `accent/50` rounded row hovers, vertical tree guides under parent rows (opaque underlay so dark-mode translucent borders don't stack), no horizontal dividers under expanded parents; the swipe-to-archive reveal layer shows only mid-swipe; board toggle uses the `SquareKanban` glyph - Kanban: every column carries a status-hued tint; lanes default expanded (including empty); compact mode collapses empty lanes to labeled rails (fixes a clipped, label-less empty-column state) - Storybook: new AgentDetail page story (realistic fixtures, light+dark) joins the visual suite; suite captures with `reducedMotion: 'reduce'` and the ux-lab reasoning ticker honors `prefers-reduced-motion`; a stale lexical alias in `storybook/main.ts` is fixed (build was broken since the lexical 0.46 bump) - Keyboard navigation, from live review of the unified lists: inbox navigation keys work on every tab (archive keys stay scoped to the archivable tab); keyboard-driven scrolling no longer hands the selection to whatever row lands under the stationary cursor (hover selects only after real pointer movement); the tasks list view gains the same j/k / arrows / Enter selection model as the inbox; and the `g` then `i` go-to-inbox chord works app-wide instead of only on the issue detail page - Token gates restored to 3/3 CLEAN (tokenized a post-#9134 regression in the recovery card); decisions recorded in `doc/design/DECISION-SHEET.md` and `doc/design/COMPONENT-INVENTORY.md` (investigation verdicts: FileTree vs WorkspaceFileBrowser and the four entity pickers stay separate — evidence included) ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2106/2106 (assertions updated in lockstep where they documented superseded decisions; new tests for the global go-to-inbox chord) - `pnpm --filter @paperclipai/ui build` → exit 0 - `pnpm build-storybook` → succeeds (also fixes the lexical-alias break on master) - Visual regression: 514-snapshot Playwright suite green against the updated baseline (zero diffs from the keyboard-navigation round — those changes are purely behavioral). Note: baselines live outside git per the suite design and the baseline-manifest archive is not yet published, so CI cannot run this suite — it was run locally throughout; every visible delta was reviewed screen-by-screen in a live instance across nine review rounds. Review evidence (before/after triplets) intentionally kept out of the repo for size; available on request. - Manual: exercised dashboard, tasks (list + board), inbox, agents, skills, costs, settings, and artifact surfaces in light and dark themes ## Risks - Wide but shallow visual surface: most changes are class-string substitutions with behavior preserved (props, handlers, roles, test ids). The riskiest areas — dnd-kit card refs (React 19 ref-as-prop), inbox swipe-to-archive, and sidebar overlays — are covered by existing unit tests (all green) and were manually exercised. - Intentional visual deltas (rounded cards, tinted kanban columns, md status glyphs, unified hovers) are design decisions recorded in `doc/design/DECISION-SHEET.md`; each maps to a re-baselined snapshot set locally. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing condition from #9134); until that lands, snapshot coverage is local-only. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending first review) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3919713bc |
[codex] Document Storybook visual baseline platform lock (#9216)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Storybook visual baselines protect UI surfaces from unintended visual drift. > - Pixel-perfect screenshot baselines are sensitive to OS, font rasterization, and browser environment. > - The suite already stores external baseline artifacts and has an opt-in CI path. > - Local runs on non-matching platforms can report false-positive diffs unless the platform lock is explicit. > - This pull request documents the Linux/Ubuntu baseline constraint and makes the local static server port explicit. > - The benefit is clearer visual-review guidance and more predictable Playwright web server startup. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? The Storybook visual baseline suite requires a matching Linux capture environment for pixel-exact comparisons, but the docs did not clearly warn local users that non-Linux environments can produce false-positive diffs. The Playwright web server command also relied on the static server's default port instead of passing the configured port explicitly. ### Expected behavior Developers should see clear Linux/Ubuntu baseline guidance before running the visual suite locally, and Playwright should start the Storybook static server on the same explicit port that the test config expects. ### Steps to reproduce 1. Review the Storybook visual docs before this PR. 2. Run or inspect the Storybook visual Playwright config. 3. Notice the missing platform guidance and implicit static server port coupling. ### Paperclip version or commit Reproducible on `master` before this branch. ### Deployment mode Local dev (pnpm dev) / built from source. ## What Changed - Documents the Linux/Ubuntu-only baseline limitation in the developer docs and visual-suite README. - Adds `--port` parsing and validation to the Storybook static server helper. - Adds regression coverage for `--port` followed by another flag. - Passes the Playwright web server port explicitly from the Storybook visual config. ## Verification - Passed: `node --check scripts/serve-storybook-static.mjs` - Passed: `node --test scripts/__tests__/serve-storybook-static.test.mjs` - Passed: `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` - Greptile: 5/5 with no unresolved review threads after commit `94a649755a2ae7c4a34a3e8a1f16ec4d26d738fd`. - Not run: full `pnpm test:storybook-visual`, because it builds Storybook and runs the browser visual suite; this PR only changes docs plus server port plumbing. ## Risks Low risk. The server still defaults to port 6106 when no explicit port is provided, and invalid port values now fail fast with a clear error before the Playwright server waits for an unreachable URL. ## Model Used OpenAI GPT-5 Codex coding agent with local command execution and repository editing tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c07e650cd7 |
feat(ui): single-source design tokens, visual regression suite, and theme retune (#9134)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be821a4f7e |
Fix DB backup health alerts (#9147)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators depend on `/api/health` and OpenAPI status surfaces to know whether the local control plane is healthy. > - Database backups are a safety-critical background process, but backup failures were not represented in health responses. > - That gap means an instance can look healthy while backup state is stale, failing, or unavailable. > - This pull request adds backup-health evaluation and exposes it through the health route, server startup wiring, and OpenAPI contract. > - The benefit is earlier operator visibility when automatic backups stop protecting instance data. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug report: **Pre-submission checklist** - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip (or can reproduce on `master`). - [x] I have confirmed the error originates in Paperclip itself — not in my agent adapter, API provider, or local configuration. **What happened?** Automatic database backup health was not included in the app health response, so backup failures or stale backups could be missed while `/api/health` still looked otherwise usable. **Expected behavior** The health endpoint should include backup-health details that let operators identify disabled, stale, failing, or healthy backup states. **Steps to reproduce** 1. Configure a Paperclip instance with automatic database backups. 2. Force backup status into a stale or failing state. 3. Call `/api/health` and inspect whether backup state is represented. **Paperclip version or commit** `master` at the PR base. **Deployment mode** Local dev (`pnpm dev`) and self-hosted server deployments. **Installation method** Built from source (`pnpm dev` / `pnpm build`). **Agent adapter(s) involved** - [x] Not adapter-specific (core bug) **Database mode** Embedded development Postgres and external Postgres backup paths. **Access context** Board/operator health checks. **Relevant logs or output** Covered by the added `server/src/__tests__/health.test.ts` cases. **Relevant config (if applicable)** Not applicable. **Additional context** This surfaces backup status only; it does not change backup execution scheduling. **Privacy checklist** - [x] I have reviewed all pasted output for PII (usernames, file paths, API keys, tokens, company names) and redacted where necessary. ## What Changed - Added a database backup health service that classifies backup recency, status, and failure conditions. - Wired backup health into app/server startup and the health route response. - Documented the backup-health behavior in development docs and OpenAPI output. - Added focused health route tests for healthy, stale, disabled, and failing backup states. ## Verification - `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/vitest run server/src/__tests__/health.test.ts` ## Risks Low-to-medium risk. This changes health response content and may affect external health consumers that parse fields strictly. It should not alter backup execution itself. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5.5 coding agent with repository tool use and local shell execution. Context window was not surfaced by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
329652a2dd |
refactor(db): add migration authoring checklist (#9122)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip stores agent work in a PostgreSQL database that evolves via numbered sequential migrations > - Migrations that run large unbounded table scans can block the server's `listen()` call during upgrades, causing multi-minute startup stalls on large databases > - Migration 0126 ran a full sequential scan backfill over `issue_comments` for derived attribution columns — O(n²) due to unindexed `LIMIT`/`OFFSET` batching, pegging CPU for ~5 minutes on a 3–4 M row table > - The safe fix (landed in #9108) replaced 0126 with a new forward-only migration using a partial index + keyset pagination backfill > - But the root cause is the absence of author-time guidance: contributors have no documented rules for writing bounded, indexed migration backfills before they land > - This PR adds a migration authoring checklist to `doc/DATABASE.md` so future contributors have those rules at hand before opening a PR > - The benefit is a durable, discoverable guide that prevents the same class of startup-blocking slowness before it reaches production ## Linked Issues or Issue Description This PR is a documentation follow-on to #9108, which landed the fast 0132 migration fix. It adds author-time guidance that captures the root-cause lesson from that incident. No separate public issue exists for the doc addition; the motivation is described above. Refs #9108 ## What Changed - `doc/DATABASE.md`: Added a **Migration authoring checklist** section with rules for indexed, bounded backfill batches — keyset pagination over `LIMIT`/`OFFSET`, mandatory partial index, idempotent guards, and split-phase schema-vs-data changes. The `check:migrations` CI gate is referenced as the enforcement backstop. ## Verification - `git diff --check -- doc/DATABASE.md` passes (no whitespace errors). - No executable code changed; the checklist is an additive documentation section. ## Risks Low. The change is additive text in `doc/DATABASE.md`. No schema, migration, or code changes. No behavioral diff. ## Model Used Claude claude-sonnet-4-6 (Anthropic, 200 K context, tool use) — used to author the migration authoring checklist and coordinate the PR workflow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with \`Fixes: #\` / \`Closes #\` / \`Refs #\` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub \`#NNN\` / \`github.com/paperclipai/paperclip\` URLs) - [x] My branch name describes the change (e.g. \`docs/...\`, \`fix/...\`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5cdf5103c9 |
docs(spec): humans/permissions granularity is V1, not OOS (#6744)
## Thinking Path > - `doc/SPEC-implementation.md` §5.2 (Out of Scope V1) conflates two distinct concerns into a single bullet: _"Multi-board governance or role-based human permission granularity"_. > - Role-based human permission granularity has been V1 for a while — the `humans-and-permissions` plan and the `principal_permission_grants` table shipped, the `PERMISSION_KEYS` set covers `users:invite`, `users:manage_permissions`, `tasks:assign`, `tasks:assign_scope`, `tasks:manage_active_checkouts`, `tasks:view_all`, `agents:view_all`, `joins:approve`, `agents:create`, `environments:manage`. > - The remaining OOS item from that bullet is _multi-board governance_ — running multiple board UIs against one company. That's a deployment-topology concern, separate from permission granularity, and stays V1 OOS. > - This PR splits the bullet so the OOS list reflects reality and contributors don't read the SPEC and conclude that per-user permission scoping is unplanned. ## What Changed Single-file docs edit in `doc/SPEC-implementation.md` §5.2: - Drop the conflated "Multi-board governance or role-based human permission granularity" bullet. - Keep "Multi-board governance (multiple board UIs for a single company)" as OOS — that part is still out of scope. - Add a short paragraph below the OOS list pointing readers to the `humans-and-permissions` plan, `principal_permission_grants`, and the existing `tasks:view_all` + `agents:view_all` opt-out scoping primitives so anyone reading §5.2 finds the V1 surface immediately. ## Verification - N/A — docs-only change, no code/schema/test impact. ## Risks - None. Single bullet rewording. ## Model Used Claude (Anthropic). Model ID: `claude-opus-4-7`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work — it documents already-shipped V1 work that the SPEC was lagging on - [x] I have run tests locally and they pass — N/A docs-only - [x] I have added or updated tests where applicable — N/A - [x] If this change affects the UI, I have included before/after screenshots — N/A docs-only - [x] I have updated relevant documentation to reflect my changes — this PR IS the doc update - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Mike <ms@moar.tools> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
574b4e71db |
[codex] Bundle UI webfonts with the app (#9020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The browser UI depends on the Inter font family for its intended visual baseline > - The app previously referenced remote Google Fonts stylesheets at runtime > - That meant self-hosted, offline, or privacy-sensitive deployments could lose the intended typography or depend on an external request > - This pull request bundles the required Inter variable font files with the UI and serves them from the app > - A normal UI test now covers the static assets and CSS wiring without adding a package script or build step > - The benefit is a more reliable, self-contained UI that does not rely on third-party webfont hosting ## Linked Issues or Issue Description No public GitHub issue was found for this exact gap. **Subsystem affected** ui/ — React + Vite board UI **Problem or motivation** Paperclip's UI should ship the webfont assets it references so production and self-hosted deployments render consistently without reaching out to Google Fonts at runtime. **Proposed solution** Bundle the Inter variable font files under the UI public assets, load them with local `@font-face` declarations, document the bundled assets, and cover the source assets/CSS wiring with a normal UI Vitest test. **Alternatives considered** Keeping the remote stylesheet dependency is simpler, but leaves deployments dependent on external font hosting. Using system fonts only would avoid the asset footprint, but changes the intended UI typography. **Roadmap alignment** This is a focused UI reliability/polish fix, not a roadmap-level core feature. **Additional context** Searched public GitHub issues and PRs for `webfonts repo:paperclipai/paperclip`; no duplicates or closely related open items were found. ## What Changed - Added bundled Inter variable font assets and their notice under `ui/public/fonts/`. - Replaced remote Google Fonts imports with local `@font-face` declarations using relative public-asset URLs that remain subpath-safe from built CSS. - Documented the local font asset expectation in development and UI spec docs. - Removed the follow-up font asset checker scripts and package/build wiring after review feedback clarified they are not required for building the UI. - Added `ui/src/lib/ui-font-assets.test.ts` to verify the shipped WOFF2 files, notice text, and CSS font references through the normal UI test suite. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/ui-font-assets.test.ts --config vitest.config.ts` - Passed. - Confirms the bundled font files exist, are WOFF2 files, have notice coverage, and are referenced by `ui/src/index.css`. - `pnpm --filter @paperclipai/ui build` - Passed. - Confirmed the UI still builds after removing the checker from `ui/package.json`. - Confirmed `ui/dist/fonts/` contains `InterVariable.woff2`, `InterVariable-Italic.woff2`, and `NOTICE.md` after the build. - The UI build emitted existing warnings about `::highlight(...)`, a dynamic/static import overlap for `MarkdownEditor.tsx`, unresolved relative public font URLs left for runtime resolution, and large chunks, but completed successfully. ## Risks Low risk. This adds static font assets and swaps the font source from a remote stylesheet to same-origin files. The main tradeoff is a larger repository/UI asset footprint from the bundled `.woff2` files. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/GitHub workflow with reasoning support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad961227f5 |
feat(secrets): add user-specific runtime secrets (#8825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs often need provider credentials, API tokens, and other environment-bound secrets. > - Company-level secrets work for shared credentials, but they do not model values that should differ by human operator. > - Without a user-scoped model, a run can dispatch without knowing whether the responsible human has supplied the needed value. > - Paperclip also needs run attribution to make those user-scoped runtime checks deterministic and auditable. > - This pull request adds user-specific secret definitions, per-user values, environment bindings, responsible-user attribution, and runtime resolution gates. > - The benefit is that teams can define the secret once, let each user provide their own value, and block runs before dispatch when required user secrets or active definitions are unavailable. ## Linked Issues or Issue Description Refs #224 Refs #6057 This PR implements user-specific secret support as a core secret-management capability rather than a one-off adapter setting. It is related to existing public work on company secrets UI and runtime secret refs, but is distinct because the value is owned by the responsible user and resolved at run dispatch time. Related PR search before opening found existing secrets work such as #1550, #8256, #8614, #8634, and #8647; none of those add the full user-secret definition/value/runtime gate covered here. ## What Changed - Added user-secret definitions and per-user "My secrets" values, keeping stored values out of access metadata. - Added `user_secret_ref` environment bindings and UI affordances to pick them alongside existing secret refs. - Added responsible-user runtime resolution so user-secret refs resolve against the human responsible for the run. - Added pre-dispatch missing-secret gates so runs fail before adapter dispatch when required user values are absent or definitions are inactive. - Added low-trust allowlist hardening for user-secret runtime access. - Added issue, routine, run, and agent API key responsible-user attribution and fail-closed dispatch behavior when attribution cannot be resolved. - Added denial-copy mapping so responsible-user authorization failures surface as actionable run outcomes instead of opaque setup failures. - Added OpenAPI documentation for the user-secret routes. - Rebases cleanly on current `master`; migrations were renumbered incrementally as `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant` after upstream `0126`/`0127` migrations. - Removed previously committed local design screenshots so the PR contains code/docs/tests only. ## Verification - PASS: PR head `2527febd106bcf3ca264ca0da7fca491084192d6` is based on `paperclipai/paperclip:master`. - PASS: `git diff --check` - PASS: `git diff --name-only public/master...HEAD | rg '^(pnpm-lock\\.yaml|\\.github/workflows/|screenshots/)' || true` produced no files. - PASS: migration journal audit confirmed unique indexes through `130` with tail entries `0126_issue_comment_derived_attribution`, `0127_environment_custom_images_instance_scoped`, `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant`. - PASS: `pnpm --filter @paperclipai/ui typecheck` - PASS: `pnpm --filter @paperclipai/server typecheck` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-active-run-output-watchdog.test.ts src/__tests__/heartbeat-stale-queue-invalidation.test.ts src/__tests__/heartbeat-workspace-finalize-branch.test.ts src/__tests__/issue-monitor-scheduler.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-comment-wake-batching.test.ts src/__tests__/heartbeat-retry-scheduling.test.ts src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts src/__tests__/heartbeat-plugin-environment.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/low-trust-red-team-routes.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` (55 tests) - PASS: `pnpm vitest run server/src/__tests__/secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` (89 tests after final Greptile cleanup fixes) - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-issue-liveness-escalation.test.ts` (17 tests after the final rebase CI fix) - PASS: focused server Vitest batches covering heartbeat recovery, project env, plugin env, routines, low-trust, pipelines, monitors, watchdog, and stale queue paths. - PASS: GitHub checks are green on `2527febd106bcf3ca264ca0da7fca491084192d6`, including Typecheck + Release Registry, Build, General tests, serialized server suites, e2e, Canary Dry Run, verify, security checks, and Greptile Review. - PASS: Greptile Review completed successfully on `2527febd106bcf3ca264ca0da7fca491084192d6` with Confidence Score 5/5, and GraphQL review-thread audit returned zero unresolved non-outdated threads. ## Risks - Runtime behavior now depends on a run having a correct responsible user; missing or incorrect responsibility assignment can block runs before adapter dispatch. - `user_secret_ref` bindings intentionally expose metadata without values, but UI/API callers may need to handle the new binding kind explicitly. - External secret providers and IAM policies are not automatically provisioned by this PR; operators still need to configure provider-side access for non-local vaults. - The PR is broad across db/shared/server/UI/runtime paths, so release validation should include both API and UI secret workflows before merge. - The migration renumbering is intentionally incremental after upstream migrations; the branch migrations use guarded column/table/index/constraint creation so users who tested the older draft numbering should not hit duplicate DDL for the existing objects. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), Codex local adapter with shell/tool use and code execution. Context window and internal reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
936687ca55 |
fix(workspace): restore clean branch drift on finalize (#8914)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs can execute inside reusable, runtime-created git worktree execution workspaces. > - Those managed worktrees record the expected branch so later dispatches do not accidentally run an agent in the wrong checkout. > - Successful run finalization already checked branch coherence, but it treated every unrecorded branch switch as fatal. > - A common publishing flow can briefly switch a clean worktree to a PR/publish branch that points at the same commit as the recorded issue branch, leaving no divergent work to protect. > - This pull request keeps the strict finalization guard for unsafe drift, but lets finalization restore the recorded branch when same-commit repair is provably safe. > - The benefit is fewer false failed runs after harmless branch switches while preserving hard failures for divergent or dirty worktrees. ## Linked Issues or Issue Description No public issue exists for this exact finalization failure. Related public worktree-recovery context: #3087 and #3056, but those address different worktree realization/reuse recovery paths rather than successful-run finalization branch repair. Bug report details: **What happened?** When an adapter run succeeded after switching a managed git worktree from its recorded issue branch to a publish/PR branch, finalization failed with a managed worktree branch mismatch even when the publish branch and recorded branch pointed at the same commit and the worktree was clean. **Expected behavior** Finalization should restore the recorded branch only when it can prove the worktree is clean, registered, and the recorded branch points at the current `HEAD`. If the actual branch has different commits or unsafe state, finalization should continue to fail with bounded validation evidence. **Steps to reproduce** 1. Create a runtime-managed `git_worktree` execution workspace for an issue run. 2. During the adapter run, create and check out a new publish branch without committing new changes. 3. Return adapter success and let heartbeat finalization run. 4. Before this change, finalization records a failed branch check and fails the run even though the branches point at the same commit. 5. With this change, finalization records the repair operation, restores the recorded branch, and records a successful finalize row. 6. Repeat with a commit on the publish branch; finalization still fails because the branch heads differ. **Paperclip version or commit** Reproduced against `master` at `bac7307ec`; fixed by this PR at `64ec605cf`. **Deployment mode** Local dev / built from source. **Agent adapter(s) involved** Not adapter-specific. This is core heartbeat/workspace finalization behavior. **Database mode** Embedded test Postgres in the focused server test. **Access context** Agent run finalization. **Node.js version** `v25.6.1` **Operating system** `Darwin 24.6.0 arm64` **Relevant logs or output** The new focused test intentionally exercises both outcomes: ```text Test Files 1 passed (1) Tests 3 passed (3) ``` **Relevant config** Runtime-created `git_worktree` execution workspace. **Additional context** The unsafe divergent branch case still fails with `workspace_validation_failed` and `git_worktree_branch_incoherence` evidence. **Privacy checklist** Reviewed; this description avoids internal task links, local workspace paths, credentials, and instance-specific URLs. ## What Changed - Reused the existing guarded branch-coherence repair helper during heartbeat finalization when the final branch inspection finds clean same-commit branch drift. - Recorded repair metadata in the `workspace_finalize` operation so reviewers/operators can audit whether finalization repaired branch drift. - Preserved failure behavior for divergent branch heads and surfaced the bounded workspace validation evidence from the repair helper. - Added focused server coverage for safe finalization repair and unsafe divergent branch failure. - Updated execution semantics docs to describe the narrower finalization rule. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks Low to medium risk. The change affects successful-run finalization for runtime-created git worktree execution workspaces. The repair path is constrained to clean, registered, same-commit branch drift, and the focused test confirms divergent branch heads still fail instead of being restored silently. ## Model Used OpenAI Codex, GPT-5-based coding agent. Exact hosted model ID was not exposed in the runtime; tool use and local shell execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2eba718bef |
Fix sandbox bridge credentials and stalled review recovery (#8844)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The local adapter and heartbeat recovery systems decide whether an agent has a real control-plane mutation path. > - Sandboxed local adapters split execution between the trusted host process and the sandbox shell/tool surface. > - A host-side adapter can still reach Paperclip while the sandbox shell surface cannot, which leaves agents thinking no endpoint or credentials are configured even though the host can still post comments. > - Execution-policy review stages can also remain pending after a reviewer run finishes without recording a decision. > - This pull request makes the sandbox bridge available to the actual shell mutation surface and adds bounded recovery for terminal-but-still-pending review participants. > - The benefit is that agents get a real reachable Paperclip API path where they need it, and stalled review stages become visible recovery work instead of silently drifting. ## Linked Issues or Issue Description No exact public GitHub issue matched this combined failure. I searched for exact and related terms including `cannot reach the Paperclip control plane`, `execution_review_participant_recovery`, `sandbox callback bridge`, `review participant in_review`, and `control plane sandbox`. Related public issues: - Refs #8482 for `in_review` liveness invariant recovery. - Refs #863 for prior agent API-key reachability confusion. - Refs #248 for the broader sandboxed agent execution model. Bug summary: - What happened: a sandboxed local-adapter run could have host-side Paperclip access while the sandbox Bash/tool surface lacked a reachable API endpoint or usable run credentials. Separately, a reviewer run could finish while its execution-review stage remained pending, leaving the source issue in `in_review` with no decision and no live participant run. - Expected behavior: the mutation surface that agents actually use should receive a run-scoped Paperclip bridge, and pending review participants should get one bounded normal-model recovery wake before moving to explicit blocked/source-scoped recovery. - Steps to reproduce: run a sandbox-backed local adapter that needs Bash/curl/tooling to call Paperclip from inside the sandbox, or finish an execution-policy reviewer run without submitting the pending review decision. - Deployment mode: local/authenticated private development instance with sandbox-backed local adapters. ## What Changed - Changed sandbox callback bridge startup so bridge credentials are passed through the sandbox runner environment instead of embedded in the visible `nohup env ...` command string. - Added adapter-utils coverage proving the sandbox shell can call Paperclip through the bridge, forwards the host run JWT with `X-Paperclip-Run-Id`, and does not leak host or bridge tokens into stdout/stderr, runner command text, or runtime files. - Added one bounded execution-review participant recovery path for terminal reviewer runs whose `executionState` remains pending. - Escalated exhausted or non-invokable review participant recovery to blocked/source-scoped recovery with dedicated evidence, activity, and next-action text. - Documented the mutation-surface reachability contract in `doc/execution-semantics.md` and updated the Paperclip skill authentication guidance for sandbox bridge env vars. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism --maxWorkers=1` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - `curl -fsS $PAPERCLIP_API_URL/api/health` returned `status: ok` on the local instance. ## Risks - Medium behavioral risk: more `in_review` issues with terminal-but-pending reviewer runs will now be retried once and then blocked explicitly instead of remaining quiet. - Low sandbox bridge risk: credential delivery moved from command text to the runner environment, which is less leaky but depends on sandbox providers honoring the env payload for startup commands. - No database migration is included. - Full repo build and CI were not run locally before opening the PR; targeted server/adapter tests and typechecks passed. ## Model Used OpenAI GPT-5 via the Codex local agent, with repository tool use and shell-based code execution. The runtime did not expose a precise context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d68c34f2cc |
Fix managed workspace branch coherence recovery (#8826)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed issue workspaces are part of the control-plane runtime boundary: the server records which git worktree and branch an agent run is allowed to use. > - Existing reuse checks validated the worktree path and cleanliness, but did not fully validate that the actual checked-out branch still matched the recorded execution workspace branch. > - That gap let an agent run switch a managed worktree onto a publishing branch without updating the execution workspace record, then later reuse or finalize the workspace as though it were coherent. > - The runtime needs a bounded repair path for provably safe mismatches and a hard validation failure for dirty, divergent, or unrecorded branch transitions. > - This pull request adds branch coherence to managed git worktree validation, records explicit recovery evidence, and prevents finalize success when a run silently changes branches. > - The benefit is that branch drift becomes either safely repaired or visibly recoverable instead of silently corrupting managed workspace state. ## Linked Issues or Issue Description No public GitHub issue exists for this bug. Bug report: - What happened: a managed agent workspace could be recorded for one branch while the underlying git worktree was actually checked out on another branch. Reuse and finalization could still treat the workspace as healthy. - Expected behavior: managed git worktrees should verify the actual branch against the recorded execution workspace branch. Safe same-HEAD clean mismatches may be repaired, while dirty, divergent, or unrecorded branch transitions should fail into explicit workspace validation recovery. - Reproduction outline: create a runtime-managed issue worktree, switch its checkout to another branch without updating the execution workspace record, then attempt reuse or run finalization. - Deployment mode: local/self-hosted Paperclip server using managed git workspaces. - Related public work: Refs #7644 and #7579. Related but not duplicate: #8275 and #5851. ## What Changed - Added managed git worktree branch inspection, formatted validation evidence, and safe same-HEAD repair logic to the workspace runtime service. - Validated recorded managed workspace branch state before reuse and during heartbeat setup. - Added finalization-time branch guards so runs that silently switch branches fail with `workspace_validation_failed` instead of recording a successful finalize. - Added recovery fingerprints and evidence for `git_worktree_branch_incoherence`, including manual-repair next actions for unsafe branch drift. - Documented branch coherence as part of runtime-created git worktree workspace coherence. - Added focused tests for safe branch repair, dirty/divergent recovery evidence, heartbeat setup validation, and finalize failure/success paths. ## Verification - `pnpm install --frozen-lockfile` - `git diff --check origin/master...HEAD` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-recovery-actions.test.ts server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` Notes: - An initial full `pnpm test:run` attempt hit a transient `socket hang up` in one `plugin-routes-authz` case. The exact case passed when rerun directly, the full `plugin-routes-authz` file passed, and the subsequent full `pnpm test:run` passed. - `pnpm build` still emits existing Vite CSS pseudo-element and chunk-size warnings unrelated to this change. ## Risks - This intentionally changes behavior for managed runs that switch branches without recording the transition: they now fail during workspace validation/finalization instead of silently proceeding. - The automatic repair path is intentionally narrow. It only repairs clean branch mismatches when both branches point at the same commit; dirty or divergent worktrees require manual recovery. - Recovery fingerprints now include workspace-validation evidence, so duplicate recovery-action grouping is more precise for branch-incoherence failures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 Codex CLI/API coding agent, with shell/git/test execution and reasoning mode enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
a8f0ebaa80 |
Refresh run config before reusing workspaces (#8797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs are assembled by the heartbeat service from agent config, project workspaces, environment config, secret bindings, skills, and runtime session state. > - The heartbeat service intentionally reuses adapter sessions, execution workspaces, and sandbox leases when that preserves useful state. > - Reuse becomes incorrect when the effective next-run config changes after a saved session, workspace, or lease was created. > - Stale reuse can make a later run appear pinned to old agent, environment, secret, instruction, or workspace settings. > - This pull request records non-sensitive fingerprints for the effective session, workspace, and lease config at run boundaries. > - When those fingerprints drift, Paperclip refreshes persisted runtime config or starts fresh execution instead of reusing stale state. > - The benefit is predictable next-run config freshness without storing raw secret values, full env maps, provider credentials, or private path details. ## Linked Issues or Issue Description - Refs #8058 - Related PRs checked during dedup search: #4968, #4155, #84, #8480. These cover nearby workspace/session routing or model-config freshness areas, but do not duplicate this effective run config fingerprinting path. ## What Changed - Added effective run config fingerprinting for session, workspace, and lease reuse decisions, with canonicalization that ignores generated runtime noise and redacts sensitive values. - Updated heartbeat reuse logic to compare stored and next-run fingerprints, reset stale saved sessions, refresh persisted workspace config snapshots, replace stale reused workspaces when required, and avoid stale sandbox lease reuse. - Included plain environment value drift via value hashes, without storing the raw env values. - Root-bound instruction content hashing so legacy direct absolute instruction paths are represented but not read for config fingerprints. - Batched secret/version metadata lookups for environment lease fingerprinting. - Added workspace operation/run result freshness metadata so operators can inspect non-sensitive decision categories. - Surfaced config freshness labels and next-run copy in the UI and docs. - Added focused coverage for fingerprint redaction, session reset decisions, workspace refresh/replace behavior, environment lease drift, and persisted workspace restoration. ## Verification - `git diff --check` - Sensitive-data scan before push: - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "(AWS_ACCESS_KEY_ID|AWS_SECRET_ACCESS_KEY|ghp_[A-Za-z0-9_]{20,}|github_pat_[A-Za-z0-9_]{20,}|sk-[A-Za-z0-9]{20,}|-----BEGIN (RSA |OPENSSH |EC |DSA )?PRIVATE KEY-----|AKIA[0-9A-Z]{16})"` - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}"` - `pnpm exec vitest run server/src/__tests__/effective-run-config-fingerprints.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/environment-runtime.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm --filter @paperclipai/db clean` - `pnpm test:run` - `pnpm build` - UI screenshots from Cutter: - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-01.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-02.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-03.png ## Risks - Medium: overly broad fingerprints could start fresh sessions, workspaces, or sandbox leases more often than necessary. - Medium: missing a config category would allow stale reuse to persist for that category. - Medium: legacy direct absolute instruction paths are no longer content-hashed unless they are paired with an absolute managed instructions root. - Low data risk: fingerprint metadata stores hashes and category names, not raw secrets, raw env values, provider credentials, or private path details. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex CLI / Codex coding agent, tool-enabled with shell, Git, GitHub CLI, local test execution, and code editing. The exact deployed model variant and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
fd2f82ac5b |
[codex] Add built-in Hermes adapters (#8543)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters are the boundary between the control plane and the runtimes that actually do work. > - Hermes support needs to be available as first-class local and gateway adapters while still preserving the adapter-manager override path for external packages. > - The adapter work touches runtime execution, UI adapter metadata, onboarding prompts, scoped credentials, release packaging, and smoke coverage, so the handoff needs concrete verification rather than only unit tests. > - This pull request adds built-in Hermes local and Hermes gateway support, keeps external adapter overrides compatible, and documents/tests the gateway flow end to end. > - The benefit is that operators can hire Hermes-backed agents without a manual plugin install, while self-hosted installs can still override/shadow the built-ins through Adapter manager packages. ## Linked Issues or Issue Description No public GitHub issue exists for this exact Hermes built-in adapter, gateway onboarding, and release-source work. Problem description: - Hermes local and gateway adapters need a public, reviewable source path in the monorepo so package artifacts and built-in adapter behavior match the application source. - Operators need built-in `hermes_local` and `hermes_gateway` adapter choices without losing the ability to install external Hermes packages as overrides. - Gateway onboarding needs secure defaults for API server URLs, API keys, and generated agent setup text. - Hermes-originated task bridge credentials need narrower API-key scope configuration. - Related public PRs found during duplicate search include #3027, #2363, #7544, #7950, #8095, and #8543. ## What Changed - Added the unified Hermes adapter package with local and gateway server/UI/CLI exports, config schemas, transcript parsing, model detection, and package metadata. - Registered `hermes_local` and `hermes_gateway` as built-in adapters across shared constants, server registries, CLI packaging, and UI adapter registries. - Kept the external adapter override path compatible so installed Hermes packages can shadow built-ins and restore the built-in parser when disabled. - Added Hermes gateway onboarding docs, board-operator docs, Docker smoke assets, and shell smoke harnesses for join/e2e validation. - Added scoped task-bridge API-key support, authorization checks, issue-origin handling, and tests for Hermes-created Paperclip tasks. - Hardened gateway transport and redaction behavior for API keys, headers, session data, and smoke diagnostics. - Updated release packaging/bootstrap checks for the Hermes packages while leaving `pnpm-lock.yaml` out of the PR per repository policy. ## Verification Targeted local verification recorded before PR handoff: - `pnpm --filter @paperclipai/hermes-paperclip-adapter exec vitest run src/gateway/server/execute.test.ts` — 14/14 passed. - `pnpm test:hermes-gateway-smoke` — 6/6 passed. - Hermes package typecheck/build checks passed. - Focused server/UI adapter tests passed — 31/31. - Release helper Node tests passed — 18/18. - `git diff --check origin/master..HEAD` passed. Fresh Docker E2E smoke evidence: - Ran `pnpm smoke:hermes-gateway-e2e` on 2026-06-26 with a fresh state directory and fresh Docker container against a live Paperclip dev server. - Hermes direct execution reached `completed`. - Hermes stop/cancel path reached `cancelled`. - Hermes gateway created a Paperclip task, Paperclip ran the Hermes agent, and the task reached `done` with the expected marker response. - Temporary board auth keys, token files, smoke state, and Docker containers were cleaned up after the run. PR checks on head `b5eae40ce`: - GitHub Actions passed: `policy`, `review`, `Typecheck + Release Registry`, all general test shards, all serialized server shards, `Build`, `Canary Dry Run`, `e2e`, and aggregate `verify`. - External checks passed: Snyk and Socket Project Report. - External Socket Pull Request Alerts remained pending after the first-party CI matrix completed. ## Risks - Medium risk: this spans adapter registration, package publishing, gateway execution, onboarding docs, API-key scoping, and UI adapter metadata. - Migration risk is low: the scope-config migration adds a nullable column and does not rewrite existing keys. - Gateway execution depends on operator-provided Hermes API configuration; the smoke covers the Docker gateway path but real deployments may differ by network/auth setup. - Direct Greptile review on the latest expanded diff is file-count limited, although the commitperclip review gate passed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool use enabled in a local repository workspace. Context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Commitperclip review gate is green; direct Greptile review is file-count limited on the latest expanded diff - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed65d08d57 |
[codex] Gate skill mutations with skills:create permission (#8616)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents and board users operate inside a company-scoped control plane where permissions decide which mutating actions they can perform > - Company skills are part of the reusable agent-company setup surface, but skill mutation had been coupled to broader agent-creation authority > - That coupling meant importing or managing skills required a permission that also implies hiring power, which is broader than the operation needs > - Paperclip already has a grant-based permission vocabulary, so skill mutation should be authorized through a dedicated `skills:create` capability while preserving existing default behavior for trusted agents > - This pull request adds the skill creation permission contract, enforces it on company skill mutations, exposes it in agent permission management, and documents the changed CLI/API expectations > - The benefit is a narrower, auditable permission path for skill import/create/update/delete flows without forcing agents to receive broader agent-creation authority ## Linked Issues or Issue Description No public issue is linked. Problem: company skill mutation APIs were effectively tied to broader agent creation authority. This PR splits skill mutation authorization onto the public `skills:create` permission while keeping existing default skill creation behavior for agents unless explicitly disabled. Related public PR found during duplicate search: #5330. That PR uses an older `canManageSkills` shape; this PR implements the `skills:create` grant path instead. ## What Changed - Added `skills:create` to shared permission constants and agent permission types/validators as `canCreateSkills`. - Backfilled default human/member role grants for `skills:create`. - Updated company skill mutation routes to require board/user or agent access to `skills:create`, while preserving legacy/default agent behavior through `canCreateSkills` unless explicitly disabled. - Updated agent permission update handling, UI permission controls, duplicate-agent payloads, plugin SDK fixtures, and agent detail API surfaces for `canCreateSkills`. - Added regression coverage for skill route authorization, permission schema/default behavior, invite grants, omitted permission updates, and duplicate-agent payloads. - Updated CLI and Paperclip skill documentation for the new skill creation permission. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-permissions-service.test.ts server/src/__tests__/agent-permissions-routes.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/invite-join-grants.test.ts ui/src/lib/duplicate-agent-payload.test.ts` — 5 files, 90 tests passed. - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/server typecheck && pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm test:run ...changed files...` was attempted first, but the stable wrapper rejects explicit file arguments; direct Vitest was used for the same targeted files. ## Risks - Moderate authorization risk: this changes the gate for company skill mutations, so the tests cover board grant checks, agent explicit grant checks, legacy default allowance, and explicit denial. - Migration/backfill risk is low: the migration only grants `skills:create` to existing human roles that already need broad management capability. - UI/API compatibility risk is low: `canCreateSkills` remains default-on for full agent permissions, and the update validator preserves omitted values so unrelated permission edits do not re-enable disabled skill creation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with terminal/tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1951c80237 |
feat(cli): surface plugin install target host + add plugin target (#8575)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The CLI (`paperclipai plugin ...`) installs and manages plugins against a Paperclip server resolved from `--api-base` / `PAPERCLIP_API_URL` / the active profile / an inferred default > - During local plugin development you can have more than one Paperclip running (a released host plus a branch build on another port), and nothing told you *which* instance a command actually talked to > - So a plugin that depends on a route or response field only present on a feature branch could be silently installed/tested against a stale host, returning `API route not found`, and look broken when the real problem was the test target > - This pull request makes the install target explicit: it probes `GET /api/health` and prints the resolved API URL + server status/version/mode/exposure before installing, and adds a `plugin target` command plus docs for running and verifying against a branch service > - The benefit is that local plugin authors can confirm they are exercising the runtime they intend to, instead of debugging phantom plugin bugs caused by hitting the wrong server ## Linked Issues or Issue Description No public GitHub issue exists, so the underlying problem is described inline following the feature-request template. **Problem or motivation** Local plugin development assumes a single Paperclip on `http://127.0.0.1:3100`. When a plugin depends on server code that only exists on a feature branch (a new scoped route, a new response field, a new managed-resource capability), installing it into a long-lived host still on older code makes the route/field missing there. The plugin falls back or errors and *looks* broken, when the real cause is that it was tested against the wrong runtime. The CLI already let you point at any server, but it never surfaced which server you ended up on — so the mistake was invisible. **Proposed solution** Make the install target explicit. Before `plugin install` runs, probe `GET /api/health` and print the resolved API URL plus server status/version/deploymentMode/exposure, so the developer can confirm which Paperclip they are installing into. Add a standalone `plugin target` command to inspect the target without installing, a `--no-verify-target` escape hatch, and docs covering how to run a branch service on its own port and verify a branch route end-to-end. **Alternatives considered** - Do nothing and rely on the existing `--api-base` / `PAPERCLIP_API_URL` resolution — rejected because the gap was never the inability to point at a branch server, it was the lack of feedback about which server was actually hit. - Fail the install when the target looks stale — rejected as too aggressive; the probe is advisory and degrades gracefully when health details are not exposed or the server is unreachable. ## Dedup Search - [x] I searched the open and recently closed GitHub PRs for similar or duplicate PRs — this is not a duplicate ## What Changed - Add `probeTargetDiagnostics` / `formatTargetDiagnostics` helpers (`cli/src/commands/client/plugin.ts`) that read `GET /api/health` and report the resolved API URL plus server `status` / `version` / `deploymentMode` / `deploymentExposure`. - `plugin install` now prints these target diagnostics before installing, so you can confirm which instance you are installing into. Skippable with `--no-verify-target`. - `plugin install --json` keeps its original flat `PluginRecord` shape (top-level `id` / `pluginKey` / `version` / `status` are unchanged); when the target was probed it gains an additional top-level `target` field. Existing automation that reads the plugin fields keeps working. - Add a standalone `paperclipai plugin target` command to inspect the install target without installing anything. - Update `doc/plugins/LOCAL_PLUGIN_DEVELOPMENT.md`: how the CLI resolves its target, how to run a branch service on its own port and point the CLI at it explicitly, an end-to-end check that the branch route is actually served, and a troubleshooting entry for the stale-target symptom. - Unit tests for the diagnostics helpers (reachable + unreachable probe, and both render paths). ## Verification - `npx vitest run cli/src/__tests__/plugin-init.test.ts` — 10/10 pass (covers `probeTargetDiagnostics` success/failure and `formatTargetDiagnostics` rendering). - CLI typecheck (`tsc --noEmit` in `cli/`) — clean. - Manual: with a server running, `paperclipai plugin target` prints `Target Paperclip: <url>` and the health line; `plugin install` prints the same block before installing and `--no-verify-target` skips it. ## Risks Low risk. The probe is read-only (`GET /api/health`) and runs before install; if the server does not expose details it degrades to `ok (no details exposed)`, and an unreachable target prints a remediation hint rather than failing the command. The `--json` output keeps its original flat shape, so existing scripts are unaffected. No server or schema changes. ## Model Used Claude Opus 4.7 (`claude-opus-4-7`), extended thinking + tool use, via Claude Code. |
||
|
|
2dbaf4a7fa |
External object references across issue surfaces (#8512)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents, issues, approvals, comments, and work products. > - The involved subsystem is issue context: markdown links, issue properties, related work, lists, filters, inbox/sidebar status, and plugin-provided external context. > - The gap is that URLs to external systems currently remain mostly plain links, so humans and agents must manually open them to understand status, identity, and liveness. > - This matters because external work objects such as GitHub issues and pull requests are part of the operational state of a Paperclip company. > - The implementation keeps core provider-neutral: shared contracts, storage, sync, routes, and UI surfaces live in core while providers can contribute detection and status resolution. > - This pull request adds the external object reference foundation, GitHub provider support, issue-surface rendering, filters, sidebar/list/inbox signals, and test/story coverage. > - The benefit is that linked external work becomes inspectable Paperclip context without hardcoding every provider directly into the UI. ## Linked Issues or Issue Description No public GitHub issue exists for this work. Feature request: - Problem: URLs in Paperclip issues, comments, documents, and related surfaces do not expose provider status or object identity inline. - Proposed behavior: detect supported external object URLs, persist normalized references, refresh provider status, and render concise status-aware links across issue surfaces. - Users affected: board users, agents, and maintainers who triage issues containing external work links. - Acceptance: external object references are company-scoped, provider-extensible, visible in key issue surfaces, filterable where relevant, and covered by focused shared/server/UI tests. Related PR search: - No open duplicate PRs found for `external object references`. - Closed related prior attempt: #4556. ## What Changed - Added shared external-object contracts, validators, status/liveness helpers, and plugin protocol declarations. - Added database schema and additive migrations for external objects, source mentions, and display metadata. - Added server services/routes for detecting, syncing, summarizing, refreshing, and resolving external objects across issues, documents, comments, projects, and plugins. - Added a GitHub external-object provider plus plugin SDK authoring docs. - Wired UI presentation across markdown links, comments, issue chat, documents, properties, related work, issue rows, filters, inbox/sidebar badges, and Storybook stories. - Rebasing cleanup: moved the branch onto current `master`, repaired stale worktree provision config, hardened environment-sensitive tests/mocks, and removed committed screenshot artifacts from the PR branch to keep the reviewable file set below tool limits. ## Verification - `pnpm exec vitest run packages/shared/src/external-objects.test.ts server/src/__tests__/external-object-routes.test.ts server/src/__tests__/external-objects-service.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/lib/external-objects.test.ts` passed after rebasing: 5 files, 56 tests. - Historical branch verification before this PR creation included `pnpm test:run`, `pnpm -r typecheck`, and `pnpm build`; this PR body does not claim those were rerun after the final rebase. ## Risks - Medium: this adds a new cross-surface sync path on issue/document/comment writes. The implementation uses safe sync wrappers so external-object failures warn instead of blocking core mutations. - Medium: the migrations introduce new tables and indexes. They are additive and company-scoped. - Medium: provider-specific URL parsing can miss or misclassify edge cases. Shared canonicalization tests and provider tests cover current GitHub shapes. - Low: UI badge/filter behavior could add visual noise for object-heavy issues; component tests and Storybook stories cover the intended surfaces. > Roadmap checked: `ROADMAP.md` references the plugin system as the current extension path and does not list a duplicate core feature. Related long-range docs discuss external references, work products, preview URLs, and plugin extension points; this PR implements the scoped external-object reference foundation. ## Model Used OpenAI Codex, GPT-5 coding-agent runtime, with shell and GitHub CLI tool use. Reasoning mode: medium. Exact deployed runtime model ID and context window were not exposed in the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fce3b439af |
fix: warn operators that experimental features may break (#8382)
## Thinking Path > - Paperclip is the control plane operators use to manage AI-agent companies. > - Board operators rely on the settings UI and CLI docs to understand which product surfaces are stable to depend on. > - Experimental features already existed in the product, but the operator-facing contract around them was too soft and too fragmented. > - That created a risk that users would enable experiments without being told clearly that they can break, change, or disappear. > - The docs and the in-product settings page both needed the same explicit warning language so the contract is visible at the moment of decision. > - This pull request adds that warning to the board-operator guide, CLI references, and the experimental settings page. > - The benefit is clearer operator expectations without changing the underlying feature flags or rollout behavior. ## Linked Issues or Issue Description No public GitHub issue exists for this docs/polish gap. Problem description: - Board operators could enable experimental features without a clear operator-facing statement that those features are opt-in and come without compatibility guarantees. - The docs site, repo CLI reference, and in-product experimental settings page did not present one consistent warning contract. - This PR closes that gap by documenting the risk explicitly where operators discover and enable those settings. Related public search: - Searched public issues/PRs for related work with `gh search issues --repo paperclipai/paperclip 'experimental features warning'` and `gh search prs --repo paperclipai/paperclip 'experimental features warning'`. - Reviewed open PR #6165 during that search and found it unrelated; it changes experimental auth/routing flags rather than documenting experimental-feature risk. ## What Changed - Added a new board-operator guide at `docs/guides/board-operator/experimental-features.md` that defines the Paperclip contract for experimental features. - Registered that guide in `docs/docs.json` so it appears in the public docs navigation. - Added matching caveat language next to `instance settings:experimental` in `docs/cli/control-plane-commands.md`. - Added the same caveat to `doc/CLI.md` so the repo CLI reference does not drift from the published docs. - Added a single page-level warning banner to `ui/src/pages/InstanceExperimentalSettings.tsx` stating that experimental features are opt-in, carry no compatibility guarantees, and may change, break, or be removed. - Added a targeted UI test in `ui/src/pages/InstanceExperimentalSettings.test.tsx` that asserts exactly one page-level warning renders with the new risk language. ## Verification - `jq empty docs/docs.json` - `git diff --check` - `cd ui && pnpm vitest run src/pages/InstanceExperimentalSettings.test.tsx` - Manual review of the warning contract across: - `docs/guides/board-operator/experimental-features.md` - `docs/cli/control-plane-commands.md` - `doc/CLI.md` - `ui/src/pages/InstanceExperimentalSettings.tsx` UI note: - This is a copy-level warning addition rather than a layout rework. I did not attach before/after screenshots in this PR body. ## Risks - Low risk: this changes operator-facing documentation and warning copy, not feature-flag behavior. - The main failure mode is wording drift across docs and UI in future edits, which is why this PR adds the same contract to all relevant operator-facing surfaces. > I checked `ROADMAP.md` before opening this PR. This is docs/UI polish around an existing experimental surface, not overlapping roadmap-level core feature work. ## Model Used - OpenAI Codex Local using `gpt-5.4` with high reasoning and tool use for coordination, review, docs changes, and PR preparation. - Anthropic Claude Local using `claude-opus-4-8` with high reasoning and tool use for the in-product warning and targeted UI test. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
277a9a43d6 |
fix(recovery): convert review-parked continuations into dependency waits (#8371)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The productivity-recovery subsystem watches for "stranded" assigned issues — claims whose live run disappeared — and repairs, resumes, or visibly blocks them > - When an executor decomposes an umbrella issue into sub-tasks, it parks its own continuation as "waiting on review/approval" (error code `issue_continuation_waiting_on_review`) — a deliberate pause, not a lost run > - Recovery's staleness gate mistook that deliberate park for a disappeared run: it retried once, then escalated the issue to `blocked` with a recovery action and an operator-facing failure notice — even though nothing had failed and there was nothing for a human to do > - The user is left staring at an inscrutable, over-technical "stranded" error on a task they did nothing wrong with, with no idea what action to take > - This pull request teaches recovery to recognize a review-parked continuation and, when the issue has a real waiting target (open sub-tasks or unresolved blockers), convert it into a first-class dependency wait: `blocked`-by-children, original assignee kept, plus a plain-language comment saying it will resume automatically > - The benefit is that post-decomposition umbrellas sit on a real waiting path and self-resume through the normal blockers-resolved flow, while genuine strands (no waiting target) still escalate exactly as before ## Linked Issues or Issue Description Refs #6503 ## What Changed - `server/src/services/recovery/service.ts`: add `resolveContinuationWaitingOnReview`. When a continuation was cancelled with `issue_continuation_waiting_on_review` and the issue has a real waiting target — open (non-terminal) sub-tasks or existing unresolved blockers — recovery sets the issue `blocked` by those issues, keeps the original assignee, posts a plain-language `system` comment, and logs the activity. Wired into `reconcileStrandedAssignedIssues` ahead of the escalation path, with a new `waitingOnReviewResolved` counter on the result. - With no waiting target, the code falls through to the existing escalation, preserving genuine stranded-run detection. - `server/src/__tests__/heartbeat-process-recovery.test.ts`: two new tests — (1) a review-parked continuation converts into a dependency wait on its open sub-tasks (done children excluded, no recovery issue opened, plain-language comment, raw error code never leaks), and (2) it still escalates when no open dependency remains. - `doc/execution-semantics.md`: document the "Deliberate wait is not a lost run" recovery rule and the requirement that a post-decomposition umbrella hold a first-class waiting path rather than relying on `parentId` rollup. ## Verification - `cd server && npx tsc --noEmit` — passes against current `master`. - New tests in `server/src/__tests__/heartbeat-process-recovery.test.ts` (describe: "heartbeat orphaned process recovery"): - "converts a continuation parked for review into a dependency wait on its open sub-tasks" - "still escalates a continuation parked for review when no open dependency remains" - Run with the repo's vitest setup, e.g. `pnpm vitest run server/src/__tests__/heartbeat-process-recovery.test.ts` (requires the embedded-postgres test harness). ## Risks Low. The change adds a single guarded pre-check ahead of the existing escalation path; behavior is unchanged when the cancellation error code is not `issue_continuation_waiting_on_review` or when the issue has no open sub-task / unresolved blocker to wait on. No schema or migration changes. ## Model Used Claude (Anthropic), Opus-class model, via the Claude Code agent harness — extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — server-only) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a71c4b6782 |
[codex] feat(watchdog): add task watchdog control plane (#8339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task lifecycle and recovery subsystems decide when agent work is still productive, stalled, or ready for review. > - Existing recovery paths can observe stopped or incomplete work, but there was no first-class per-task watchdog model with scoped review permissions. > - Watchdog follow-ups also need strict boundaries so recovery/status-only runs cannot mutate approvals or perform deliverable work. > - This pull request adds the task watchdog data model, API/service layer, scheduler/review flow, adapter wake context, UI configuration surfaces, and docs. > - The branch has been rebased onto current `paperclipai/paperclip` `master`; the watchdog migration is now ordered after master's latest migrations as `0104_issue_watchdogs`. > - The benefit is a more explicit task-review loop that preserves Paperclip's single-assignee and governance invariants while making stalled work easier to route. ## Linked Issues or Issue Description No linked GitHub issue. Paperclip task: [PAP-11275](/PAP/issues/PAP-11275). ## Problem or motivation Task recovery needs a first-class watchdog path that can inspect stopped work and create scoped follow-ups without bypassing normal task ownership. Board/UI users need a way to configure watchdogs on tasks and see watchdog-related live work. Recovery/status-only runs must remain limited to status reporting and must not create approvals, link approvals, or submit approval comments. ## Proposed solution Add a task-watchdog data model, scheduler/classifier, scoped mutation guard, adapter wake context, API/UI configuration surfaces, and documentation so watchdog agents can review stopped task subtrees under explicit boundaries. ## Alternatives considered Reuse the existing recovery-action flow only. That would keep stopped-work detection implicit, make per-task watchdog assignment harder to expose in the UI, and would not provide a durable scoped-review issue for stalled task trees. ## Roadmap alignment This is Paperclip control-plane lifecycle infrastructure for task execution and recovery. I checked `ROADMAP.md`; this PR does not duplicate an existing planned core item. ## What Changed - Added issue watchdog schema, migration, shared contracts, validators, CRUD API, and service support. - Added task watchdog scheduler/classifier behavior, scoped mutation enforcement, adapter wake context, and default watchdog mandate guidance. - Added UI surfaces for configuring watchdogs on new/existing tasks, viewing watchdog activity, and exposing the experimental setting. - Added docs for the user-facing task watchdog workflow and implementation semantics. - Gated new-task watchdog setup behind `enableTaskWatchdogs` and blocked cheap status-only recovery runs from approval mutations. - Rebased onto current `master` and renumbered the idempotent watchdog migration from the branch-local `0102_issue_watchdogs` slot to `0104_issue_watchdogs`. - Addressed Greptile feedback by loading watchdog classifier input with a recursive subtree query and centralizing the watchdog origin-kind constant. - Added and updated focused server/UI tests for watchdog routes, scheduler/classifier behavior, scope boundaries, live task visibility, settings, and new issue dialog behavior. ## Verification - `pnpm vitest run server/src/__tests__/task-watchdogs-scheduler.test.ts server/src/__tests__/task-watchdogs-classifier.test.ts` - `pnpm vitest run server/src/__tests__/approval-routes-idempotency.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts` - `pnpm vitest run ui/src/components/NewIssueDialog.test.tsx` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` - Verified the PR diff does not include `pnpm-lock.yaml` or `.github/workflows`. ## Risks - Medium risk: this introduces a new task lifecycle surface touching DB schema, server routes/services, adapter wake context, and UI task configuration. - Watchdog scheduling behavior depends on the new experimental setting and runtime context checks behaving consistently across local and production agents. - The watchdog migration is idempotent (`IF NOT EXISTS` / duplicate-object guards) so users who tried the previous branch-local migration number should not get duplicate-object failures. - CI and the second Greptile pass are pending after the latest review-fix push. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in the Paperclip workspace. Exact runtime model id and context window were not exposed to the agent; tool use and local command execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots — N/A per Paperclip task instruction: do not add screenshots/images to this PR unless they are specifically part of the work. - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7069053a1f |
[codex] Add ask issue work mode (#8334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue work mode controls how a task starts and how the conversation composer frames the operator's intent. > - Paperclip already supports standard agent execution and planning mode, but there is no lightweight mode for asking a question without immediately implying execution or plan drafting. > - That gap makes low-commitment clarification workflows look like normal task execution. > - This pull request adds an explicit Ask mode and threads it through shared contracts, server heartbeat context, and the issue composer UI. > - The benefit is that operators can create or switch a task into a question-oriented mode while preserving existing agent and planning flows. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Inline feature request follows the repository feature request template. ### Subsystem affected Cross-cutting: `packages/shared`, `server/`, and `ui/`. ### Problem or motivation Issue conversations currently distinguish standard agent work from planning work, but question-first conversations do not have a clear public mode in the shared contract or UI. Operators who want to ask an agent a focused question have to use standard mode, which can imply normal task execution, or planning mode, which asks for a plan rather than an answer. ### Proposed solution Add Ask as a first-class issue work mode. It should be selectable from issue creation and issue chat, cycle alongside Standard and Planning from the keyboard shortcut/menu, appear distinctly in composer styling, and be included in heartbeat context so agents know to answer directly instead of executing or drafting a plan. ### Alternatives considered - Keep using standard mode for questions: rejected because it does not communicate answer-only intent to the agent or the UI. - Reuse planning mode for questions: rejected because planning mode asks for a plan and is semantically different from asking a question. - Add only local UI copy: rejected because the mode needs to be represented in the shared contract and server heartbeat context to be reliable. ### Roadmap alignment This is a focused issue-workflow improvement. `ROADMAP.md` was checked and no duplicate planned core work was found. ### Additional context Related public searches performed before opening this PR: - GitHub PR search for `"ask mode" repo:paperclipai/paperclip` - GitHub issue search for `"ask mode" repo:paperclipai/paperclip` - GitHub PR search for `"work mode" "ask" repo:paperclipai/paperclip` No duplicate PR was found. ## What Changed - Added `ask` to the shared issue work-mode contract and validation coverage. - Included issue work mode in heartbeat context summaries so agents can see standard, planning, and ask state. - Added Ask mode metadata, styling, composer tone handling, and selection/cycling behavior in the issue chat/new issue UI. - Updated focused tests for shared validators, heartbeat context, and affected UI work-mode flows. ## Verification - `NODE_ENV=test pnpm exec vitest run ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts` - `NODE_ENV=test pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/issues-service.test.ts ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts ui/src/pages/IssueDetail.test.tsx` The broader targeted command passed 8 test files / 245 tests. Visual reference for Standard/Planning/Ask composer states: https://gist.github.com/cryppadotta/714d8590bac55500a65e7e16de5bb4b8 It emitted an expected warning from an existing server test fixture about a missing run-log fixture while verifying derived issue comment metadata. ## Risks Low to moderate risk. This adds a new enum value that crosses shared, server, and UI contracts. Existing standard and planning modes are preserved, but any downstream code assuming only two non-terminal work modes may need to handle `ask`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent in Paperclip CodexCoder mode, with shell, git, GitHub connector, and local test execution tools. Context window and exact hosted model snapshot are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`, `feat/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc95699fde |
fix(server): enforce agent secret binding sync across lifecycle flows (#8307)
## Thinking Path > - Paperclip is the control plane people use to create, configure, and run AI agents for work. > - This change sits in the server-side agent lifecycle and secret-binding subsystem, where adapter config `env` entries can reference company secrets. > - An incident (while trying to configure a Novita sandbox) showed that an agent can reach a broken runtime state if `adapterConfig.env` contains `secret_ref` entries but the matching `company_secret_bindings` rows are missing. > - The immediate run-path guard and error-surfacing work made the failure diagnosable, but they did not fully prevent new broken agents from being created. > - The risk came from create and approval flows being responsible for remembering to sync bindings at each call site, which is easy to miss as new flows are added. > - This pull request moves the invariant into `agentService` create/update/activate paths, keeps the existing hire-flow fix, and adds regression coverage for create, update, and legacy pending-approval recovery. > - The benefit is that agent secret binding integrity is enforced closer to the data mutation point, so future callers inherit the protection automatically. ## Linked Issues or Issue Description Refs #8309 ### What happened? A Paperclip agent could persist `adapterConfig.env` `secret_ref` entries without matching agent-scoped `company_secret_bindings` rows. When that happened, the config UI could still look configured, but the real run path failed pre-dispatch because the secret was not actually bound to that agent. ### Expected behavior Every normal agent create, config-update, and pending-approval activation flow should leave the agent with secret bindings that match its persisted secret-ref env config. ### Steps to reproduce 1. Create or activate an agent through a flow that persists `adapterConfig.env` secret refs without synchronizing `company_secret_bindings`. 2. Observe that the config state can still appear populated. 3. Start a run for that agent. 4. Observe that pre-dispatch binding validation fails because the secret reference exists but the agent binding does not. ### Deployment mode Local dev (`pnpm dev`) ### Installation method Built from source (`pnpm dev` / `pnpm build`) ### Agent adapter(s) involved - Claude Code - Not adapter-specific (core bug) ### Database mode Embedded PGlite / embedded local dev database flow ### Access context Board (human operator) created or approved the agent; agent runtime later consumed the config. ### Additional context This PR focuses on preventing new broken states from normal service flows and on backfilling the covered legacy pending-approval activation path. ## What Changed - Kept the existing branch-local hire-flow fix that synchronized bindings for route and approval paths. - Moved the binding integrity invariant into `agentService.create()`, `agentService.update()` when `adapterConfig` changes, and `agentService.activatePendingApproval()`. - Added `server/src/__tests__/agents-service-secret-bindings.test.ts` covering create-time sync, update-time resync, and backfill for legacy pending-approval agents. - Removed now-redundant route-layer and approval-layer binding sync calls once the service layer became authoritative. - Simplified the affected unit tests so route/approval tests no longer assert service-owned binding writes directly. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/agents-service-secret-bindings.test.ts server/src/__tests__/approvals-service.test.ts server/src/__tests__/agent-skills-routes.test.ts` ## Risks - Low to medium risk. - This changes where secret-binding synchronization is enforced, so any unexpected caller that relied on upper-layer manual sync behavior could behave differently. - Agent create/update/activation flows now perform binding synchronization consistently, which adds binding-table writes at those mutation points. - This PR does not retroactively scan and heal every already-broken historical agent row; it prevents and backfills through the covered service flows. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex / GPT-5 Codex class model via `codex_local` - Session model family: GPT-5 Codex - Tool-assisted coding with shell, git, HTTP, and local test execution - Reasoning mode: medium interactive tool-use workflow ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5320a44088 |
Guard codex_local agents from shared OpenAI key (#8272)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The `codex_local` adapter runs local Codex CLI processes and builds their environment from persisted agent config plus host process env. > - A host-level `OPENAI_API_KEY` or shared Codex auth home can silently make new agents spend through shared credentials. > - Existing agents can be repaired manually, but new and updated agents need a persistent guard at the agent configuration boundary. > - This pull request isolates new and updated `codex_local` agents with per-agent `CODEX_HOME` and an empty `OPENAI_API_KEY` override. > - The benefit is that future agent creation or adapter updates cannot silently fall back to shared OpenAI credentials. ## Linked Issues or Issue Description Paperclip work item: [ZOL-5477](/ZOL/issues/ZOL-5477). No matching GitHub issue exists, so the bug is described inline following `.github/ISSUE_TEMPLATE/bug_report.yml`. **Pre-submission checklist** - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest `master` commit for this PR branch. - [x] I have confirmed the error originates in Paperclip's `codex_local` adapter configuration boundary, not in a provider outage. **What happened?** New or updated `codex_local` agents could inherit a host-level `OPENAI_API_KEY` or use a shared Codex home when their adapter config did not explicitly isolate those values. That made it possible for future agents or manual adapter edits to silently fall back to shared OpenAI credentials. **Expected behavior** Creating, hiring, or updating a `codex_local` agent should either persist isolated per-agent configuration or reject unsafe shared Codex home configuration with a clear 422 response. The guard must not print secret values. **Steps to reproduce** 1. Create or update a `codex_local` agent without an explicit `adapterConfig.env.OPENAI_API_KEY` override. 2. Run it on a host where the Paperclip server process has `OPENAI_API_KEY` set. 3. Observe that the adapter process can inherit the host key unless Paperclip persists a blocking empty override. 4. Set `adapterConfig.env.CODEX_HOME` to a shared path such as `~/.codex` or the company-level `codex-home`. 5. Observe that the old code allowed the shared auth home instead of returning a validation error. **Paperclip version or commit** - Reproduced by inspection against `master` before this PR. **Deployment mode** - Local dev / self-hosted server with `codex_local` agents. **Installation method** - Built from source. **Agent adapter(s) involved** - Codex. **Database mode** - Not database-related. **Access context** - Board and agent configuration paths. **Relevant logs or output** - No secret-bearing logs included. **Relevant config** - Unsafe shape: missing `adapterConfig.env.OPENAI_API_KEY`, or shared `adapterConfig.env.CODEX_HOME`. - Fixed shape: per-agent `CODEX_HOME` plus empty `OPENAI_API_KEY` override. **Additional context** Related PR search for `codex_local OPENAI_API_KEY CODEX_HOME` found: - #3681 `fix: preserve managed Codex auth and repo-root env loading` - #5621 `fix: copy worktree codex auth locally` Those are adjacent auth-handling changes, but they do not add the agent create/update guard implemented here. **Privacy checklist** - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added a `codex_local` config guard in agent create, hire, and update routes. - The guard assigns `adapterConfig.env.CODEX_HOME` to `companies/<companyId>/agents/<agentId>/codex-home` when missing. - The guard persists `adapterConfig.env.OPENAI_API_KEY = ""` when missing, preventing host env inheritance. - Shared `CODEX_HOME` values for the company codex-home, host `$CODEX_HOME`, or `~/.codex` now fail with a 422 error. - Added route tests for create, hire, update, and rejected shared host Codex home. - Updated `codex_local` and development docs to describe the per-agent home contract. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/agent-skills-routes.test.ts` - `pnpm typecheck` - `git diff --check upstream/master...HEAD` - `gh pr list --repo paperclipai/paperclip --state all --search "codex_local OPENAI_API_KEY CODEX_HOME" --limit 20 --json number,title,state,url` - `rg -n "codex|OPENAI_API_KEY|CODEX_HOME|adapter" ROADMAP.md` returned no roadmap overlap. ## Risks - Existing legacy `codex_local` agents with shared `CODEX_HOME` will get a clear 422 when their adapter config is updated until the shared path is replaced. This is intentional because silent fallback is the bug being guarded. - Low migration risk: no database migration and no secret values are printed or persisted beyond the empty override. ## Model Used - OpenAI GPT-5.5 Codex, Codex coding-agent session with repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Paperclip - Issue: [ZOL-5477](/ZOL/issues/ZOL-5477) - Owner: Разработчик (`6625498c-66c9-429f-b578-4463ddc3ba16`) - Status: waiting reviewer - Next action: merge after approval and green CI --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |