mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
efcce9cc8e17e2b30f6f2edf30e00a77b0bf19bb
1179
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |
||
|
|
0e21a27301 |
fix: forward onSpawn to hermes and process adapters for PID persistence (#8722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter layer (hermes-local, process adapters) delegates agent execution to child processes via `runChildProcess()` > - `runChildProcess()` accepts an `onSpawn` callback to report child PID and process group info, but the hermes and process adapters were not forwarding `ctx.onSpawn` to this call > - Without PID persistence, the orphan reaper cannot distinguish live runs from abandoned processes, causing false-positive reaps and 5-minute timeout errors for active runs > - This pull request adds `onSpawn: ctx.onSpawn` to both adapter call sites and declares the option in the `runChildProcess` wrapper type > - The benefit is that the orphan reaper can now correctly track live child processes, eliminating false-positive reaps ## Linked Issues or Issue Description Fixes #8723 Fixes false-positive orphan reaps in hermes-local and process adapters by forwarding the `onSpawn` callback to `runChildProcess()`. All other adapters (claude-local, codex-local, cursor-local, gemini-local, grok-local, opencode-local, pi-local) already forward `ctx.onSpawn` — these two were the only ones missing it. ## What Changed - `server/src/adapters/utils.ts`: Added `onSpawn?` to the `runChildProcess()` options type so callers can forward the callback - `server/src/adapters/process/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` - `packages/adapters/hermes/src/server/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` ## Verification - `pnpm -r typecheck` passes across all packages - Confirmed all other adapters already forward `ctx.onSpawn` (12 grep matches across 9 adapter files) - The 3-line diff is additive only — no existing behavior is changed, only a previously-ignored callback is now forwarded ## Risks Low risk. This is a 3-line additive change. The `onSpawn` parameter is optional (`?`) so existing callers are unaffected. The callback is already well-established across all other adapters. ## Model Used Hermes Agent (by Nous Research) — xiaomi/mimo-v2.5-pro via OpenRouter, with tool use (file editing, git, GitHub API). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass (typecheck passes) - [x] I have added or updated tests where applicable (N/A — type-level fix only, no behavioral change) - [x] I have updated relevant documentation to reflect my changes (N/A — internal fix) - [x] I have considered and documented any risks above --------- Co-authored-by: Zephyr <zephyr@motoyuki.dev> |
||
|
|
c8253e3641 |
fix(adapters): inject execution contract once per fresh heartbeat (#9469)
## Thinking Path > - Paperclip coordinates AI-agent work through repeated heartbeat runs. > - Adapter prompts combine a default heartbeat template with scoped wake context. > - Fresh heartbeats received the same execution contract from both layers, wasting prompt tokens and obscuring which layer owns the contract. > - Resume deltas and template-less adapters do not share that composition path, so removing the wake-payload copy unconditionally would drop required guidance. > - Empty comment batches also emitted instructions and metadata that only matter when comments exist. > - This pull request makes execution-contract inclusion explicit by prompt path, preserves OpenClaw gateway behavior, and suppresses no-op comment boilerplate. > - The benefit is one contract per heartbeat path and roughly 300 fewer prompt tokens on a fresh zero-comment wake. ## Linked Issues or Issue Description - Fixes #9221 - Refs #9200 - Refs #7634 ## What Changed - Stop emitting the execution-contract paragraph from fresh scoped wake payloads because the default heartbeat template already contains the full contract. - Keep the contract in resume deltas, and add `includeExecutionContract` for adapters that do not render the default heartbeat template. - Opt `openclaw-gateway` into wake-payload contract rendering so template-less gateway runs retain the guidance. - Omit comment-batch acknowledgement/fetch guidance and empty `pending comments` / `latest comment id` metadata when a fresh wake has no pending comments. - Add regression and acceptance coverage proving composed fresh prompts contain `Execution contract` exactly once while resume and template-less paths retain it. Measured effect: the fresh zero-comment wake block drops from 1,840 to 855 characters (about 300 tokens saved per fresh heartbeat; about 220 on comment wakes), and the composed fresh prompt contains `Execution contract` once instead of twice. ## Verification - `npx vitest run packages/adapter-utils/src/server-utils.test.ts` — 63 passed - `npx vitest run server/src/__tests__/codex-local-execute.test.ts` — 13 passed - `npx vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/openclaw-gateway-adapter.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts` — 27 passed - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` — passed ## Risks - Low risk: prompt text and adapter composition only; no database or API migration. - The main compatibility risk is a template-less adapter losing the contract. The explicit option and OpenClaw gateway regression coverage protect the known template-less path. - External adapters that call `renderPaperclipWakePrompt` directly can opt into `includeExecutionContract: true` when they do not render the default template. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, reasoning mode with tool use and code execution; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5dff52631d |
fix(heartbeat): throttle redundant issue re-wakes (#9470)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents and their work > - Heartbeat admission decides when an agent should start another adapter session for an issue > - After process-loss recovery, assignment pollers and reconcilers can repeatedly request another wake while the issue remains `in_progress` > - When the preceding runs succeeded without issue-visible progress, those event-free wakes provide no new information but still pay the full cost of an adapter session > - Existing liveness evidence is too broad for this case because workspace tool calls can make a run look active without moving the issue > - This pull request adds an issue-scoped admission throttle for consecutive no-progress re-wakes while preserving every wake that carries new information or recovery intent > - The benefit is bounded recovery cost without delaying comments, operator actions, failures, or other meaningful events ## Linked Issues or Issue Description No public GitHub issue exists for this bug. **What happened?** After a process died, external wake drivers could re-wake the same agent for the same `in_progress` issue every few seconds. Each succeeded run that produced no issue-visible progress could be followed by another full adapter session despite no new issue input. In the observed recovery smoke, one recovery consumed 25 sessions and 2.4× the direct-run cost. **Expected behavior** Repeated event-free re-wakes should back off after consecutive successful runs produce no issue-visible progress. Any new information, explicit operator intent, or failed-run recovery should continue immediately. **Steps to reproduce** 1. Start an issue heartbeat and simulate process loss while the issue remains `in_progress`. 2. Allow assignment/reconciliation drivers to request repeated event-free wakes for the same agent and issue. 3. Complete each follow-up run successfully without adding a comment, issue mutation, document, work product, interaction, or continuation. 4. Observe repeated adapter sessions starting every few seconds without new issue input. **Environment** - Version: reproduced on `master` before this change - Deployment: local development, built from source - Adapter scope: core bug; not adapter-specific - Database: reproduced and tested with embedded Postgres ## What Changed - Add a pure issue re-wake throttle that detects consecutive succeeded runs without issue-visible progress and applies a 120-second exponential cooldown capped at 30 minutes. - Gate event-free `enqueueWakeup` requests and return the explicit skip reason `issue_rewake_throttled` while the cooldown is active. - Always bypass throttling for comment wakes, new issue activity, explicit resumes, `forceFreshSession`, event-shaped reasons, and post-failure recovery. - Add focused pure unit coverage and database-backed heartbeat admission coverage for throttle and bypass behavior. ## Verification - `cd server && pnpm vitest run src/__tests__/issue-rewake-throttle.test.ts` — 12 passed. - `cd server && pnpm vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6 passed with embedded Postgres. - `cd server && pnpm run typecheck` — passed. - Neighbor suites previously verified: `heartbeat-dependency-scheduling`, `heartbeat-process-recovery`, `run-continuations`, `heartbeat-issue-liveness-escalation`, `recovery-stale-issue-lock-sweep`, and `heartbeat-comment-wake-batching` — 131 tests passed. ## Risks - A progress classifier that is too narrow could defer a legitimate event-free poll; the cooldown is bounded and new issue activity bypasses it immediately. - A progress classifier that is too broad could allow the original heartbeat storm; tests intentionally distinguish issue-visible mutations from workspace-only activity. - Low compatibility risk: no schema, API contract, or migration changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The runtime does not expose the exact underlying model ID or context-window size; reasoning, terminal tool use, code inspection, GitHub CLI access, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e7e84e3fe |
fix(adapters): propagate ACP-lane usage and cost into spend telemetry (#9471)
## Thinking Path > - Paperclip is the open source control plane for running and governing AI-agent companies. > - Adapter executions feed token usage, billing identity, and run cost into the control plane's spend telemetry. > - The default ACP execution lane for local Claude and Codex adapters did not propagate per-turn usage or cost, so paid runs could be recorded with zero spend and no tokens. > - Claude CLI result events could also undercount output tokens by reading only the main-loop usage block instead of the complete per-model ledger. > - The shared executor needs to distinguish per-run usage from session-cumulative usage so the server does not apply the wrong delta heuristic. > - This pull request captures ACP usage and cumulative-cost deltas, resolves adapter billing identity, uses Claude's complete model-usage ledger, and preserves per-run usage in server normalization. > - The benefit is accurate token and cost accounting across the default paid Claude and Codex execution paths. ## Linked Issues or Issue Description ### What happened? Paid `claude_local` and `codex_local` runs using the default ACP engine can complete successfully while the control plane records zero or null cost and missing token usage. Claude CLI result parsing can additionally undercount output tokens when subagent or sidechain usage is present. ### Steps to reproduce 1. Run a paid Claude or Codex local adapter through the ACP engine. 2. Complete a turn that reports usage and cumulative cost through ACP status/events. 3. Inspect the execution result and normalized run telemetry. ### Expected behavior The execution result contains per-turn token usage, a per-run USD cost delta, and the correct billing identity. Server normalization records those per-run values without applying a session-cumulative delta a second time. ### Actual behavior before this change ACP execution results returned no usage and `costUsd: null` with unknown billing. The server therefore recorded zero spend and no tokens for paid runs. Claude CLI parsing could use an incomplete usage block. ## What Changed - Capture ACP usage from runtime status and `usage_update` events, reporting it as `usageBasis: per_run`. - Convert agent-reported cumulative ACP cost into a per-turn delta, including counter-reset and no-report safeguards. - Add a shared billing-identity resolver and map Claude and Codex authentication/provider modes to control-plane billing types. - Prefer Claude result-event `modelUsage` totals so subagent and sidechain tokens are included. - Skip the server's session-cumulative usage delta when an adapter explicitly reports per-run usage. - Add regression coverage for usage capture, event fallback, cost resets, stale reports, billing identities, model-usage totals, and server spend normalization. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/acp.test.ts server/src/__tests__/costs-service.test.ts server/src/__tests__/monthly-spend-service.test.ts` — 6 files, 126 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `pnpm --filter @paperclipai/adapter-claude-local typecheck` — passed. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - A broader Claude-local suite has a pre-existing rate-limit classification failure in `test.probe.test.ts`; it also fails on clean `master` and is unrelated to this change. ## Risks - Cost reporting depends on the agent's cumulative counter semantics; reset handling falls back to the post-turn amount and is covered by regression tests. - Incorrect billing-mode inference could misclassify spend; provider/auth mappings mirror each adapter's existing CLI behavior and have focused tests. - The new `usageBasis` contract changes server normalization only when adapters explicitly opt into `per_run`; existing adapters retain prior behavior. - No database migration, workflow, lockfile, or UI changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation commit: Anthropic Claude Fable 5, tool-enabled coding workflow (exact context window and runtime configuration were not recorded in the commit metadata). - PR preparation and verification: OpenAI Codex, tool-enabled coding agent (runtime model ID and context window are not exposed to this session). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4a40c0cb13 |
feat(routines): gate scheduled runs on external activity (#9436)
## Thinking Path > - Paperclip is the open source app people use to manage AI-agent companies and their recurring work > - Scheduled routines provide native cron-driven execution for recurring agent tasks > - Watcher-style routines currently dispatch a model run even when the control plane has been quiet since their last useful run > - Existing pause, catch-up, and concurrency policies do not distinguish external work from a routine's own bookkeeping > - This pull request adds a generic activity gate that checks company-scoped activity provenance before scheduled dispatch > - The benefit is backward-compatible zero-token quiet skips while real human, agent, or delegated-child activity still wakes the routine ## Linked Issues or Issue Description - Refs #8534 ## What Changed - Added `activity_gate_policy` and `activity_gate_scope` routine columns with backward-compatible `always` / `company` defaults. - Added a company-bounded `evaluateActivityGate()` predicate that uses the last dispatched run as its open window, excludes the routine's own execution runs and scheduler bookkeeping, ignores pure-read actions, and supports company/project scope. - Integrated the predicate into scheduled ticks after pause/worktree eligibility checks; quiet ticks create visible skipped run-history rows with reason `no_external_activity` and gate-window diagnostics without advancing the activity window. - Kept webhook, manual, and API dispatch paths ungated; catch-up schedules evaluate the gate once per scheduler tick. - Added migration-default, provenance predicate, project-scope, quiet-window, scheduler, and webhook-bypass coverage. ## Verification - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts` — 51 tests passed - Embedded Postgres `EXPLAIN` for the company-scope gate scan: ```text Limit (cost=24.56..24.58 rows=1 width=24) -> Incremental Sort (cost=24.56..24.60 rows=2 width=24) Sort Key: activity.created_at, activity.id Presorted Key: activity.created_at -> Nested Loop Anti Join (cost=0.44..24.55 rows=1 width=24) Join Filter: (own_run.id = activity.run_id) -> Index Scan using activity_log_company_created_idx on activity_log activity (cost=0.15..8.19 rows=1 width=40) Index Cond: ((company_id = '00000000-0000-0000-0000-000000000001'::uuid) AND (created_at > (now() - '01:00:00'::interval)) AND (created_at <= now())) ``` ## Risks - The migration adds two non-null text columns, but constant defaults preserve all existing routine behavior and avoid a backfill step. - Project scope resolves activity through issue/run/routine provenance; tests cover in-project and cross-project issue activity, while every top-level and correlated query remains company-bounded. - This is the scheduler/schema foundation. Public API validation and documentation for configuring the new fields are intentionally handled in the next scoped follow-up. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex using `gpt-5.4` with medium reasoning, repository/tool access, terminal code execution, and test execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR extends the existing Scheduled Routines roadmap item - [x] I have searched GitHub for duplicate or related PRs and linked the related efficiency request above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing configuration is exposed in this scoped foundation PR) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4e12bfb89 |
fix(workspaces): persist readiness state and validate ports (#9408)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces can run managed services that must report reliable lifecycle and readiness state > - Service startup previously waited for readiness before committing the starting row, making concurrent control actions see stale state > - Fixed service ports also needed clearer configuration and ownership diagnostics to avoid cross-workspace collisions > - This pull request persists startup state before readiness, validates port ownership, and exposes configurable service ports in the workspace UI > - The benefit is dependable service controls and actionable diagnostics when workspace runtimes start slowly or compete for ports ## Linked Issues or Issue Description ### What happened? Slow-starting workspace services could remain invisible to concurrent stop/restart controls until readiness completed, and fixed-port conflicts lacked enough ownership context for safe repair. ### Expected behavior A starting service is persisted immediately, control operations can observe it, configured ports are editable, and conflicts identify the owning process/workspace. ### Steps to reproduce 1. Configure a workspace service that delays binding its HTTP port. 2. Start the service and immediately request another control action. 3. Observe stale persisted state before this change. 4. Configure two workspaces for the same fixed port and observe limited conflict diagnostics. ### Paperclip version or commit `origin/master` at `02e2dd271` ### Deployment mode Local dev; built from source; not adapter-specific; database-backed workspace runtime state. ## What Changed - Commit the `starting` runtime-service row before waiting for readiness and transition it after the probe completes. - Add port-owner inspection and cross-workspace conflict details to local service supervision. - Preserve configurable runtime service ports through workspace configuration updates. - Surface service-port editing and validation in the execution workspace details UI. - Add server and UI regression coverage for slow readiness, concurrent controls, port persistence, and conflict diagnostics. ## Verification - `vitest --project @paperclipai/server src/__tests__/workspace-runtime.test.ts src/__tests__/execution-workspaces-service.test.ts` — 118 tests passed. - `vitest --project @paperclipai/ui src/pages/ExecutionWorkspaceDetail.service-ports.test.ts` — 4 tests passed. - `node scripts/check-token-gates.mjs` — all token gates clean. ## Risks - Moderate risk: changes touch workspace service lifecycle persistence and local process/port inspection. - No schema migration is required; tests exercise slow readiness, concurrent control, persisted ports, and cross-workspace conflicts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.3 Codex, reasoning with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e0f1905222 |
[codex] Quiet packaged version fallback diagnostics (#9207)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server startup path reports the product version from package metadata and, in source checkouts, Git metadata. > - Packaged installs can run from `node_modules`, where Git metadata is normally unavailable and that absence is expected. > - The fallback path was still attempting Git metadata probing in packaged contexts, which could print scary diagnostic noise during onboarding even though the package version fallback was working. > - This pull request makes the packaged path skip Git probing only when the package does not look like a source checkout, and keeps fallback diagnostics opt-in. > - The benefit is a quieter first-run experience without weakening source-checkout version detection or debug diagnostics. ## Linked Issues or Issue Description No public GitHub issue exists. ### Bug Report #### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in an agent adapter, API provider, or local configuration. #### What happened? When Paperclip starts from a packaged install, server version resolution can fall back from Git metadata to package metadata. That expected fallback path could emit scary Git diagnostic noise during onboarding even though startup could continue normally. #### Expected behavior Packaged Paperclip startup should use package metadata quietly when Git metadata is unavailable. Source checkouts should still use Git-derived versions, and operators who explicitly opt into version-resolution diagnostics should still receive useful Git failure details. #### Steps to reproduce 1. Run Paperclip from a packaged install where the server package is under `node_modules` and does not include package-local Git metadata. 2. Start the server in an environment where `git describe` cannot resolve repository metadata for that package. 3. Observe that version fallback can produce Git diagnostic noise during startup even though the package version fallback is expected. #### Paperclip version or commit Reproduced against the pre-fix server version resolution behavior on `master`-derived builds. #### Deployment mode Self-hosted server / packaged local install. #### Installation method npm / pnpm package install. #### Agent adapter(s) involved Not adapter-specific; this is core server startup/version behavior. #### Database mode Not database-related. #### Access context Unclear / not applicable. #### Relevant logs or output Git fallback diagnostics from `git describe` could appear during packaged startup. The exact path and Git output depend on the operator environment. #### Additional context The fix keeps diagnostics available behind `PAPERCLIP_DEBUG_VERSION_RESOLUTION=1` and preserves source-checkout Git version detection, including source paths that happen to contain a `node_modules` segment. #### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Skip Git metadata probing for packaged installs under `node_modules` only when no package-local Git metadata is present. - Preserve Git-derived version detection for source or linked workspace checkouts, even when their path contains a `node_modules` segment. - Keep fallback diagnostics behind the existing debug/diagnostic opt-in path. - Include useful Git failure details such as stderr/stdout/stack/cause when diagnostics are enabled. - Add version tests covering packaged fallback behavior, source-checkout detection, richer diagnostics, and quiet default output. ## Verification - `pnpm vitest run server/src/__tests__/version.test.ts` passed after the Greptile follow-up changes. - `pnpm --filter @paperclipai/server typecheck` passed. - `git diff --check` passed. - Greptile completed with confidence score 5/5 and no blocking issues on the latest reviewed commit. ## Risks Low risk. The change is scoped to version fallback behavior. Source-checkout Git version detection remains covered, while packaged `node_modules` contexts intentionally rely on package metadata instead of Git probing. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent with shell/tool use in the Paperclip workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local> |
||
|
|
36ec79c196 |
feat: add attention queue and Decisions surface (#9380)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies, where operators need a reliable way to find and act on work awaiting their input. > - The attention and issue-thread interaction subsystems expose those decision points across server APIs and the board UI. > - The previous navigation and interaction presentation left these actions fragmented and did not offer a controlled rollout for the Decisions surface. > - This branch adds the attention feed, richer interaction cards, grouping, dismiss/snooze behavior, and a gated Decisions sidebar entry. > - It also keeps experimental settings and API contracts synchronized, with an idempotent migration for the new dismissal state. > - This pull request delivers the complete, tested attention/Decisions experience as one reviewable unit. ## Linked Issues or Issue Description - Adds an operator-focused attention queue and Decisions experience: grouped decision cards, semantic interaction actions, dismiss/snooze handling, resilient interaction states, and an experimental flag to control the Decisions navigation entry. ## Feature Context ### Problem or Motivation Operators currently have to hunt across approvals, interactions, failed runs, and budget alerts to find decisions that need their action. ### Proposed Solution Provide a gated Decisions attention queue that groups actionable items, supports direct resolution, and preserves operator control through dismiss and snooze actions. ### Alternatives Considered Keep separate, source-specific views only; this leaves cross-cutting operator decisions fragmented and harder to prioritize. ### Roadmap Alignment This improves the V1 control-plane operator workflow by making pending governed actions discoverable in one company-scoped surface. ## What Changed - Added server attention-feed services, routes, interaction handling, dismiss/snooze support, and an idempotent `0145` inbox-dismissal migration. - Added shared attention, inbox-dismissal, and experimental-settings contracts. - Added Decisions/attention UI, interaction-card states, sidebar badge/navigation integration, grouping, keyboard support, and Storybook coverage. - Added tests for attention behavior, thread interactions, settings normalization, dismissals, and API behavior. - Removed generated screenshots from the final PR diff and rebased the branch onto current `master`. ## Verification - `pnpm check:token-gates` — passed. - `pnpm exec vitest run packages/shared/src/issue-thread-interactions.test.ts server/src/__tests__/attention-service.test.ts server/src/__tests__/inbox-dismissals.test.ts server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issue-thread-interaction-routes.test.ts ui/src/lib/attention.test.ts ui/src/components/AttentionQueueRow.test.tsx ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx` — passed: 158 tests across 9 focused files. - GitHub Actions for `ad636f560`: build and typecheck/release-registry have passed; remaining general-server and Greptile checks are in progress. ## Risks - Moderate: this is a cross-layer attention/interaction feature with a new migration and navigation behavior. - The `enableDecisions` experimental setting defaults to off, limiting rollout impact. - Existing dismissal data is backfilled to `dismiss`; the migration is idempotent and uses guarded constraint creation. > ROADMAP.md was checked; no duplicate planned core feature was identified. Related open pull requests were searched before opening this PR. ## Model Used - OpenAI GPT-5.5 via Codex CLI, with tool use and local code execution. Context-window size unavailable in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public PR branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; focused tests pass and the remaining unrelated AWS test failure is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
17dde9d3f2 |
fix(sandbox): keep custom-image snapshots applied to config tests, probes, and saves (#9385)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox environments can capture reusable custom images (provider snapshots) so agents boot with pre-installed tools and CLI logins > - The custom-image runtime fingerprint check included provider secret-ref paths (e.g. the Daytona `apiKey`), while capture-time fingerprinting excluded them, so any config carrying a credential never matched its captured snapshot > - As a result, agent config tests and environment probes silently booted the provider base image instead of the snapshot, test sandboxes were deleted before operators could inspect them, and any environment save orphaned the snapshot without warning > - The UI compounded the confusion by displaying an internal template id that matches nothing in the provider dashboard > - This pull request aligns runtime fingerprints with capture-time exclusions, re-stamps fingerprints on saves that cannot affect the snapshot (warning when they can), archives test/probe sandboxes instead of deleting them, and surfaces the provider snapshot ref in the UI > - The benefit is that custom images actually apply to config tests and probes, survive unrelated config edits, and are debuggable against the provider dashboard ## Linked Issues or Issue Description No public GitHub issue exists for this; describing it in-PR per the bug template. Related: Refs #9329 (saved-environment probe company context — this branch carries an equivalent fix), Refs #8794 (introduced reusable sandbox custom images). **What happened?** With a Daytona environment whose provider config stores the API key as a secret reference and an active captured custom-image snapshot: - Agent config tests and environment probes booted the provider base image (`daytonaio/sandbox:0.8.0`) instead of the captured snapshot, so CLI upgrades/logins baked into the snapshot were missing and the probe reported "login required" and an outdated CLI. - The environment card showed an internal template id (e.g. `b5be03e1-ca5…`) that does not correspond to any snapshot name in the provider dashboard, making the active image impossible to correlate. - Test/probe sandboxes were deleted immediately after the run, so the sandbox a test used could not be inspected afterwards. - Saving the environment config (even fields unrelated to the image) changed the stored fingerprint, silently detaching the snapshot with no warning. **Expected behavior** Config tests and probes boot the captured snapshot when one is active; the UI shows the provider-facing snapshot/template ref; test sandboxes stay inspectable for a short window; unrelated config edits keep the snapshot linked, and edits that genuinely invalidate it produce an explicit warning. **Steps to reproduce** 1. Configure a sandbox environment on Daytona with the API key stored as a company secret reference. 2. Capture a custom image snapshot from the environment page and mark it active (e.g. after installing/logging into a CLI in the setup sandbox). 3. Run the agent config test or an environment probe: the sandbox boots the base image, not the snapshot, and the sandbox is deleted immediately after the test. 4. Save the environment config with an unrelated field change: the snapshot silently stops applying. **Paperclip version or commit** `master` at the merge-base of this branch. **Deployment mode** Self-hosted local instance (macOS, pnpm dev server) with the Daytona sandbox provider plugin. ## What Changed - Runtime custom-image fingerprint checks now exclude provider secret-ref paths, matching capture-time exclusions, so configs carrying credentials match their captured snapshots (`environment-custom-image-runtime.ts`). - Agent config tests and saved-environment probes force fresh, non-reused sandboxes and pass company context so lease-backed probes can resolve company secrets and boot the real snapshot (`environment-probe.ts`, `routes/agents.ts`, `routes/environments.ts`). - Test/probe sandboxes are released by archiving (stop + 60-minute provider-side auto-delete) instead of immediate deletion, so operators can inspect the exact sandbox a test used (Daytona plugin). - On environment PATCH save, changes that cannot affect the captured snapshot re-stamp the template's source fingerprint so the snapshot stays linked; boot-source or provider-identity changes (new manifest field `templateIdentityPaths`) mark the template detached and the save response reports it (`environment-custom-images.ts`, shared plugin types/validators). - The custom-image overview exposes `activeTemplateMatchesConfig`; the environments UI shows the provider snapshot/template ref (internal id moved to a tooltip), warns via toast when a save detaches the snapshot, and shows a persistent "Not in use" warning when the active template no longer matches the saved config (`CompanyEnvironments.tsx`, `api/environments.ts`). ## Verification - `pnpm vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-probe.test.ts server/src/__tests__/environment-routes.test.ts server/src/__tests__/agent-test-environment-routes.test.ts` — server coverage for fingerprint exclusions, re-stamp/detach on save, probe company context, and fresh-sandbox test behavior. - `pnpm vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts` — archive-on-release and snapshot ref handling. - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx` — snapshot ref display, detach toast, and "Not in use" warning. - Manually verified end-to-end on a live self-hosted instance against real Daytona: config test boots the captured snapshot (CLI login and version persist), the test sandbox remains visible in the provider dashboard as archived, and saving unrelated fields keeps the snapshot applied. ## Risks - Fingerprint exclusion widening: a provider credential rotation alone no longer detaches a captured snapshot; that is the intended behavior (the snapshot content does not depend on the credential), and provider-identity fields (e.g. Daytona `apiUrl`) still detach via `templateIdentityPaths`. - Archived test sandboxes consume provider-side resources for up to their auto-delete window instead of being freed immediately; bounded (60 minutes) and only for test/probe sandboxes. - New optional manifest field `templateIdentityPaths` is backward-compatible; providers that omit it keep current matching behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
70ce005bef |
Ensure worktree execution starts only after activation (#9374)
## Thinking Path > - Paperclip is the open-source control plane people use to manage AI agents and their work. > - Its scheduler, routines, and heartbeat services decide when agents automatically begin work. > - Experimental per-worktree execution is useful for isolated development, but enabling it previously allowed automatic services to consider an existing backlog. > - A worktree activation must therefore create a durable eligibility boundary rather than merely toggle execution on. > - This pull request records an activation cutoff and applies it consistently to automatic routine and heartbeat dispatch. > - The result is that an enabled worktree executes only work created after its own activation, while non-worktree behavior remains unchanged. ## Linked Issues or Issue Description **Problem type:** Bug / safety regression **Summary:** Enabling experimental run execution in an existing worktree could start automatic scheduler, routine, watchdog, and heartbeat activity for work created before that worktree was explicitly armed. **Expected behavior:** A worktree that has execution enabled only considers automatically dispatched work created on or after its activation timestamp. Ambiguous activation state fails closed. Non-worktree instances keep their existing behavior. **Related public work:** Refs #8275 (runtime worktree policy gating); this PR adds an activation-time boundary for automatic execution rather than changing the general runtime policy. ## What Changed - Persist a worktree execution activation timestamp and originating instance ID; stamp them only when the experimental toggle changes from disabled to enabled. - Resolve activation state fail-closed when the cutoff is missing, invalid, disabled, or belongs to another instance. - Gate automatic routine scheduling, webhooks, watchdog activity, and heartbeat selection at the activation cutoff; manual runs remain available. - Share the canonical worktree truthy-environment helper across routine dispatch and agent inbox filtering. - Add cutoff and truthy-runtime regression coverage, plus experimental-settings UI states that explain armed and suppressed execution. ## Verification - `pnpm exec vitest run server/src/__tests__/routines-service.test.ts server/src/__tests__/instance-settings-service.test.ts` — passes: 2 files, 60 tests. - `pnpm --filter @paperclipai/server typecheck` — passes. - Existing CI completed successfully before the follow-up review fixes; this branch was rebased onto the latest `origin/master` before retesting. ## Risks - **Behavioral:** Automatic worktree execution is intentionally more restrictive; pre-existing work is suppressed until newly created after activation. - **Operational:** A malformed or cross-instance activation record fails closed, requiring an operator to disable and re-enable the experimental toggle on the intended worktree. - **Compatibility:** The worktree environment now accepts all canonical truthy values (`1`, `true`, `yes`, and `on`) consistently; non-worktree instances are unaffected. - **Branch metadata:** This existing execution-workspace branch predates the current naming rule and cannot be renamed under this task's workspace contract; the code and PR title do not include internal ticket references. > `ROADMAP.md` was checked; this targeted execution-safety fix does not duplicate planned core work. ## Model Used - Anthropic Claude Code — assisted with the original implementation; exact model identifier and context window were not recorded in the repository metadata. - OpenAI Codex CLI — assisted with PR preparation and review fixes; exact model identifier and context window are not exposed in this execution environment. Used with terminal tooling, code editing, and targeted test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
23f34491e2 |
Fix apiCompression corrupting and dropping Better Auth responses for gzip clients (#9381)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its server fronts every API route — including Better Auth sign-in — with Express middleware, and #9190 added an `apiCompression` middleware that gzips JSON responses over 1KB > - That middleware buffers `res.write()` chunks with `String(chunk)`, but Better Auth (via better-call) streams `Uint8Array` chunks and commits headers with `writeHead()` before streaming > - `String(Uint8Array)` serializes the body to comma-separated decimal bytes (~3.4x inflation), and once the inflated body crossed the 1KB threshold, `setHeader()` threw `ERR_HTTP_HEADERS_SENT` and the catch handler destroyed the socket > - Every real browser sends `Accept-Encoding: gzip`, so sign-in returned zero bytes (`net::ERR_EMPTY_RESPONSE` / "Failed to fetch"), while curl without `Accept-Encoding` worked — making the bug easy to misdiagnose as a client or network issue > - This pull request makes the middleware byte-safe for `Uint8Array` chunks, passes through responses whose headers are already committed, and falls back to the uncompressed body instead of destroying the connection when compression fails > - The benefit is that browser sign-in (and any other streamed binary-chunk response) works again for gzip-accepting clients, with regression tests locking in all three behaviors ## Linked Issues or Issue Description Refs #9190 (introduced the `apiCompression` middleware). No public GitHub issue exists; bug description: - **What happened:** Sign-in from any real browser failed with `net::ERR_EMPTY_RESPONSE` / "Failed to fetch". The server logged `ERR_HTTP_HEADERS_SENT` from the compression middleware and destroyed the response socket, so zero bytes reached the client. - **Expected:** `/api/auth/*` responses are delivered intact regardless of the client's `Accept-Encoding`. - **Steps to reproduce:** Run the server with API compression active, open the web UI in a browser (which sends `Accept-Encoding: gzip`), and attempt email/password sign-in. The auth response body exceeds ~300 bytes, so after the ~3.4x stringification inflation it crosses the 1024-byte compression threshold and the response is destroyed. `curl` without `Accept-Encoding` succeeds against the same server. - **Scope:** Any route that streams `Uint8Array` chunks and/or commits headers via `writeHead()` before writing — in practice all Better Auth routes served through better-call. ## What Changed - `server/src/middleware/api-compression.ts`: - Buffer `res.write()` chunks with a `toBodyBuffer()` helper that converts `Uint8Array`/`ArrayBuffer` views via `Buffer.from()` instead of `String()`, so binary chunks are preserved byte-for-byte. - Pass responses through untouched once headers are already sent (`writeHead()`-style streaming), since compression headers can no longer be set at that point. - On any compression failure, write the original uncompressed body instead of calling `res.destroy()`, so clients get a valid (just uncompressed) response rather than a dropped connection. - `server/src/__tests__/api-compression.test.ts`: three new regression tests — small `writeHead`+`Uint8Array` responses are delivered byte-for-byte, large ones no longer drop the connection, and `Uint8Array` JSON bodies gzip without corruption (includes `/api/auth-bridge` and `/api/uint8-json` test routes mirroring better-call's streaming pattern). ## Verification - `cd server && pnpm vitest run src/__tests__/api-compression.test.ts` — 10/10 passing (7 pre-existing + 3 new regression tests). - Manual: with the fix, browser sign-in against a dev instance succeeds for gzip-accepting clients; before the fix the same request returned `net::ERR_EMPTY_RESPONSE`. ## Risks - Low risk. The middleware still compresses large text/JSON responses exactly as before; the changes only affect paths that previously produced corrupted or destroyed responses. - Behavioral shift: responses whose headers were already committed are now delivered uncompressed instead of being (incorrectly) buffered — this is strictly less surprising than the previous corrupted output. - Failure-path shift: a compression error now yields an uncompressed 200 response instead of a dropped connection. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, running via Claude Code / Paperclip agent harness with tool use (shell, file edit, test execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1f07690184 |
fix(ui): keep issue threads from jumping to latest comment (#9354)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue and item detail pages use a shared issue chat thread to show comments, runs, activity, and interactions. > - That thread still defaulted to landing on the latest comment when messages first loaded. > - On long issue/item pages, that default can yank the operator away from the top of the page before they choose to inspect the newest message. > - Deep links to comment hashes can create the same kind of initial viewport jump when they are used as generic navigation targets. > - This pull request makes initial latest-comment and initial thread-hash scrolling opt-in instead of default behavior. > - The benefit is stable initial page position across issue-thread surfaces while keeping the explicit Jump to latest control available. ## Linked Issues or Issue Description No exact public GitHub issue was found for this bug. Bug description: - What happened: opening a page with a shared issue conversation thread could automatically move the viewport toward the newest comment/thread target. - Expected behavior: ordinary page loads should keep the initial viewport stable unless the user explicitly clicks Jump to latest. - Steps to reproduce: open an issue or item detail page with a long conversation thread and observe whether the page jumps to the newest thread entry on initial load. - Paperclip version/commit: reproduced while working on the current `master` branch lineage. - Deployment mode: local trusted/dev UI. Related public thread/comment UX work: Refs #3916, Refs #7972, Refs #8800. ## What Changed - Changed `IssueChatThread` so initial latest-comment scrolling defaults to off. - Added a separate opt-in for initial thread-hash scrolling, also defaulting to off. - Preserved stale deleted-comment hash cleanup without scrolling the page. - Updated regression coverage so default initial load stays put, comment hashes do not scroll by default, and manual Jump to latest still scrolls. ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/pages/IssueDetail.test.tsx -t "loads from the pending state into issue detail without changing hook order"` passed on the clean PR branch. - `pnpm --dir ui exec vitest run src/components/IssueChatThread.test.tsx` was attempted on the clean PR branch, but the file fails before changed assertions with the existing `TypeError: act is not a function` test-harness issue across 58 tests; 14 tests passed. - Static check: no `autoScrollToLatestOnInitialLoad={true}` or `autoScrollToHashOnInitialLoad={true}` call sites remain in `ui/src`. ## Risks Low risk. This only changes initial scroll defaults in the shared issue thread. The main behavioral shift is that direct comment/thread hashes no longer auto-scroll on first load unless a caller explicitly opts in; the Jump to latest button and post-submit scroll behavior are unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding-agent runtime; exact context window not exposed in this environment; tool-enabled repository inspection, editing, testing, git, GitHub CLI, and Paperclip API usage. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a4993a72a6 |
Fix live run streaming text readability (#9330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The issue thread UI renders live agent output from adapter run logs and transcript parsing. > - Some adapter streams emit many small or repeated token chunks, and live UI updates can expose partial words, duplicated slices, or transient markdown placeholders. > - That makes active run updates look like gibberish even when the underlying agent output is valid. > - The fix needs to preserve raw logs while making the live thread view stable, readable, and ordered. > - This pull request adds monotonic run-log sequencing, safer live transcript dedupe/order handling, markdown placeholder hiding, and readable live text stabilization. > - The benefit is a live issue thread that updates smoothly without showing confusing partial parser artifacts. ## Linked Issues or Issue Description No public GitHub issue exists yet, so this PR includes the bug details inline. ### What happened? Live run updates in the issue thread can show confusing repeated or partial text while an adapter is streaming. The visible text appears to lose parsing boundaries during active updates, especially with ACP-style token deltas, so the live output can briefly render duplicated chunks, incomplete words, or HTML-comment placeholders. ### Expected behavior Live text should remain readable while preserving the underlying run output for raw inspection. ### Steps to reproduce 1. Start a live agent run whose adapter emits small stdout token deltas. 2. Watch the issue thread while the run is still active. 3. Observe transient duplicated chunks, incomplete words, or markdown placeholder artifacts in the live rendered text. ### Paperclip version or commit Reproduced against current `master` before this PR branch. ### Deployment mode Local dev issue-thread UI with live local adapter runs. ### Additional context GitHub PR search for `live run streaming text markdown transcript` found one broad merged PR, `#252` (“Dotta updates - sorry it's so large”), but no targeted duplicate for this live streaming readability bug. ## What Changed - Added per-run monotonic sequence numbers to persisted and live run-log chunks. - Dedupe and order live transcript chunks by sequence before falling back to timestamp ordering. - Hide markdown HTML comment placeholder text from rendered markdown output. - Smooth live issue-thread text updates so partial additions reveal at readable word boundaries and sliding-window removals do not produce gibberish. - Added coverage for run-log ordering/deduping, markdown comment hiding, live issue-thread stabilization, and Greptile-reviewed edge cases where overlap rewrites could synthesize text or no-boundary additions could stay hidden. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/MarkdownBody.test.tsx src/components/transcript/useLiveRunTranscripts.test.tsx` passed before the review fix: 3 files, 86 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts` passed after the review fix: 1 file, 30 tests. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/issue-chat-messages.test.ts src/components/transcript/useLiveRunTranscripts.test.tsx` passed after the final Greptile overlap fix: 2 files, 40 tests. - `pnpm check:token-gates` passed. - Local PII/secret scan of touched files found only expected code/test words such as `secret`, `token`, and redaction-related strings; no literal credentials found. - `pnpm -r typecheck` passed after restoring declared dependencies with `CI=1 pnpm install --frozen-lockfile` and running with a short `TMPDIR` because `tsx` IPC sockets fail under the long sandbox temp path. - `pnpm build` passed with existing Vite CSS/font/chunk warnings. - GitHub PR checks passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813`: build, typecheck/release registry, server and workspace test shards, serialized server suites, e2e, canary dry run, policy, review, Socket, Superagent, Snyk, and verify. - Greptile review passed on head `4c052dfe86aecb5feb73504e6b48843f68fce813` with confidence score 5/5 and no blocking issues found. - `pnpm test:run` failed in unrelated server workspace tests on this macOS local environment: - `server/src/__tests__/heartbeat-workspace-branch-containment.test.ts`: two assertions compare `/tmp/...` with `/private/tmp/...`. - `server/src/__tests__/heartbeat-worktree-suppression.test.ts`: expected one heartbeat run but observed two, followed by cleanup fallout in the full run. - Isolated rerun of those two server suites reproduced the same three failures. ## Risks - Low product risk for the UI changes: the readable smoothing only affects active live-run display stabilization, not stored comments or raw run logs. - Moderate verification risk: local full Vitest did not pass because of unrelated server workspace tests. Targeted tests for this change, typecheck, token gates, and build passed. - Run-log sequence fields are optional for compatibility with older log rows that do not include `seq`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-using local workspace execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a02fe8d575 |
Update Codex adapter GPT-5.6 defaults (#9352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Codex local is the adapter subsystem that exposes OpenAI Codex CLI model choices to agents and issue overrides. > - OpenAI has GPT-5.6 Codex-capable models that should appear in Paperclip's built-in Codex model list and refresh behavior. > - Paperclip's server model listing falls back to the adapter metadata and merges OpenAI refresh results with known Codex defaults. > - This pull request updates the Codex default model metadata to include GPT-5.6 options and adds regression coverage for fallback and refresh paths. > - The benefit is that operators can select the new Codex models without relying on manual model IDs, and refresh behavior keeps known GPT-5.6 options visible. ## Linked Issues or Issue Description Refs #9322. Refs #9342. Refs #9346. ### Agent or provider Codex CLI (OpenAI). ### Why this adapter is useful OpenAI's GPT-5.6 Codex-capable models should be available in Paperclip's Codex adapter defaults and model refresh path. ### How the agent is invoked `codex` ## What Changed - Changed the `codex_local` default model metadata from `gpt-5.5` to `gpt-5.6`. - Added `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` to the built-in Codex adapter model list. - Updated adapter and server model-listing tests to cover GPT-5.6 fallback and refresh behavior. - Aligned Codex Fast mode support and helper text with the new `gpt-5.6` default, while preserving GPT-5.5, GPT-5.4, and manual model ID support. ## Verification - `git diff --check origin/master...HEAD` - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts server/src/__tests__/adapter-models.test.ts server/src/__tests__/adapter-model-refresh-routes.test.ts` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` ## Risks Medium risk because changing `DEFAULT_CODEX_LOCAL_MODEL` from `gpt-5.5` to `gpt-5.6` changes the adapter's default model selection for new blank configurations. The model-list additions are otherwise low risk and covered by adapter/server metadata tests. This PR intentionally overlaps related PRs #9342 and #9346, so reviewers may prefer to close or fold it into one of those branches. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, GitHub CLI, and repository editing tool use. Exact served model ID and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f5f461729 |
Avoid startup crash when Reflection Coach assets are missing (#9351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server has built-in agent definitions that are loaded during startup and used to provision optional operational agents such as Reflection Coach > - Reflection Coach stores richer stock instructions, a routine description, and a bundled skill as markdown assets outside the TypeScript module body > - A deployed server can fail before it is healthy if one of those copied markdown assets is absent from `server/dist` > - A recent build fix preserves those assets during normal server builds, but runtime should still degrade gracefully if a packaged asset is missing or unreadable > - This pull request adds resilient loading for built-in Reflection Coach assets and keeps a minimal compiled fallback available > - The benefit is that a missing optional built-in agent file no longer turns into a process-wide startup crash ## Linked Issues or Issue Description Bug fix. No public GitHub issue found for this exact startup crash. - Related public PR: #9339 - What happened: the server could throw `ENOENT` while importing the built-in agent service if `server/dist/built-ins/agents/reflection-coach/AGENTS.md` was missing from a deployed build. - Expected behavior: the server should keep starting, log that the built-in asset was missing, and use safe fallback text for the optional built-in agent resource. - Steps to reproduce: build the server, remove the compiled Reflection Coach `AGENTS.md` asset from `server/dist`, then import/start the server path that loads built-in agent definitions. - Paperclip version/commit: reproduced against a deployed build containing the Reflection Coach built-in agent assets; fixed against current `master` after #9339. - Deployment mode: Node server deployment using compiled `server/dist` output. ## What Changed - Added built-in agent text loading that checks the compiled asset path first, then source/package fallback paths, then a minimal compiled-in fallback string. - Added fallback text for Reflection Coach instructions, routine description, and bundled skill content so startup does not depend on optional markdown assets being present. - Added regression coverage for readable candidate selection and missing-file fallback behavior. ## Verification - `pnpm -w exec vitest run server/src/__tests__/built-in-agents.test.ts` — 1 test file passed, 24 tests passed. - `pnpm --filter @paperclipai/server build` — server TypeScript build completed and copied `src/built-ins` into `dist/built-ins`. - Manual smoke: temporarily moved `server/dist/built-ins/agents/reflection-coach/AGENTS.md`, imported `server/dist/services/built-in-agents.js` through the repo-pinned `tsx` runtime, and confirmed Reflection Coach definitions still loaded with output `reflection-coach:3732`; the asset was restored afterward. ## Risks Low risk. The normal path still uses the full packaged markdown assets. The fallback path is only used when those files are missing or unreadable, and it logs a warning so packaging drift remains visible. ## Model Used OpenAI GPT-5 Codex coding agent, with repository tool access and shell-based verification. Exact context window was not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc81eefb60 |
Make plan-approval continuations durable after failed wakes (#9331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents often work from reviewed plans that are approved through issue-thread interactions. > - Accepting a plan is not just a UI decision; it must reliably resume the assignee so approved work continues. > - A failed continuation wake could leave an approved plan stranded in review with no durable retry or visible recovery path. > - This pull request makes approved plan continuations retryable, recoverable, and visible when resume fails. > - The benefit is that operators can trust plan approval to either resume the agent or produce an explicit actionable failure instead of silent limbo. ## Linked Issues or Issue Description No public GitHub issue exists. Inline bug report follows the repository bug template. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest released version of Paperclip or can reproduce on `master`. - [x] I have confirmed the error originates in Paperclip itself, not in my agent adapter, API provider, or local configuration. ### What happened? When a plan-confirmation interaction was accepted, the assignee continuation wake could fail before useful agent execution. In that case the issue could remain in review even though the plan had been approved, because the failed wake was fire-and-forget and there was no durable retry or recovery path for accepted continuations. ### Expected behavior Accepted plan continuations should either wake the assignee successfully, retry bounded infrastructure failures, recover dropped wakes, or surface an explicit failure state that operators can act on. ### Steps to reproduce 1. Create an issue with an assignee and a plan confirmation that wakes the assignee on accept. 2. Accept the confirmation. 3. Simulate a pre-flight continuation failure, such as process loss before agent start or workspace validation failure. 4. Observe that the approved issue can remain in review without an active assignee wake or visible retry/failure state. ### Paperclip version or commit Reproducible on `master` before this PR's retry/recovery changes. ### Deployment mode Local dev (`pnpm dev`) and server-side recovery paths. ### Installation method Built from source (`pnpm install`, `pnpm dev`, test runner). ### Agent adapter(s) involved Not adapter-specific; this is a core continuation/recovery bug. The tests cover local-agent failure shapes without relying on a provider-specific API. ### Database mode Embedded Postgres test database for verification. The affected logic is database-backed and applies to normal Postgres deployments as well. ### Access context Board accepts the interaction; agent execution resumes through the assignee wake path. ### Relevant logs or output No sensitive logs are needed. The regression tests simulate the failed wake and recovery states directly. ### Relevant config No special config is required beyond an assignee with wake-on-demand enabled. ### Additional context This PR also prevents a stale workspace-validation payload from quarantining another issue's active workspace and prevents unrelated successful runs from masking a continuation that never resumed. ### Privacy checklist - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, and company names, and redacted where necessary. ## What Changed - Added bounded infrastructure retries for failed accepted-interaction continuation wakes. - Extended stranded issue recovery so dropped accepted-plan continuation wakes are requeued. - Recorded and rendered explicit resume-failure state on accepted confirmation cards. - Added clean-workspace fallback for workspace-validation failures while preventing cross-issue workspace quarantine. - Tightened recovery so unrelated successful runs do not mask an accepted continuation that never resumed. - Added focused server/UI coverage for retry scheduling, recovery, visible failure state, and interaction card rendering. ## Verification - `pnpm vitest run server/src/__tests__/heartbeat-retry-scheduling.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts` — 106 tests passed. - Earlier branch verification also covered the issue-thread interaction card tests for the visible resume-failure UI. - GitHub CI is green on the replacement PR head, and Greptile reports 5/5 with no blocking issues. ## Risks - Medium behavioral risk: this changes recovery behavior for accepted continuation interactions and workspace-validation retries. - Mitigation: retries are bounded, scoped to same-company issue context, and workspace quarantine now requires ownership by the issue being retried. - Existing stored confirmation results remain compatible because the new resume-failure field is optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled terminal workflow. The runtime does not expose an exact context-window value to the agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d166069bc4 |
Preserve built-in agent assets in server builds (#9339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server package ships compiled runtime code plus static runtime assets. > - Built-in agent definitions live under `server/src/built-ins` and runtime code resolves them relative to compiled server files. > - The server build already copied onboarding assets into `dist`, but it did not copy built-in agent assets alongside the compiled code. > - Packaged server builds could therefore miss built-in agent definitions even though source-based development runs worked. > - This pull request extends the server build copy step to preserve built-in agent assets in `dist/built-ins`. > - The benefit is packaged server builds keep the same built-in agent runtime assets available as source-based development runs. ## Linked Issues or Issue Description No public GitHub issue found. This PR describes the underlying bug inline using the bug report template fields. ### What happened? `@paperclipai/server` build output copied `server/src/onboarding-assets` into `server/dist/onboarding-assets`, but did not copy `server/src/built-ins` into `server/dist/built-ins`. Runtime code for built-in agents resolves those assets relative to the compiled server files, so packaged builds could omit built-in agent markdown assets that are present during source-based development. ### Expected behavior Packaged server builds should include built-in agent assets under `server/dist/built-ins`, matching the runtime location expected by the compiled server code. ### Steps to reproduce 1. Check out current `master` before this PR. 2. Run `pnpm --filter @paperclipai/server build`. 3. Check for `server/dist/built-ins/agents/reflection-coach/AGENTS.md`. 4. Observe that the built-in agent asset is missing from the server build output. ### Paperclip version or commit Reproduces on current `master` before this PR. The fix is verified on commit `2b89984ccb7857f06359bf65c48222f110c7aeff`. ### Deployment mode Build/package artifact behavior. This can affect any deployment mode that runs from the built server package rather than directly from source. ### Installation method Built from source with `pnpm --filter @paperclipai/server build`. ### Agent adapter(s) involved Not adapter-specific. This is a core server packaging bug for built-in agent assets. ### Database mode Not database-related. ### Access context Not applicable. This happens during package build output generation. Related search: - Searched public PRs/issues for `built-ins build copy repo:paperclipai/paperclip`. - Found no directly related open issue. One old closed Hermes adapter PR was not directly related. ## What Changed - Updated the `@paperclipai/server` build script to create `dist/built-ins`. - Added the copy step from `server/src/built-ins` into `server/dist/built-ins` alongside the existing onboarding asset copy. - Added a focused server package build-script test that asserts both onboarding and built-in static runtime asset directories are copied into `dist`. ## Verification - `pnpm exec vitest run server/src/__tests__/server-package-build-script.test.ts` - `pnpm --filter @paperclipai/server build` - `test -f server/dist/built-ins/agents/reflection-coach/AGENTS.md` ## Risks Low risk. This changes only the package build asset copy step and adds focused test coverage. The main risk is build-script portability, but it follows the existing `mkdir -p` and `cp -R` pattern already used for onboarding assets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent in a tool-enabled Paperclip heartbeat. Exact model ID and context-window size are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5c85ae64a0 |
Cases: experimental first-class case object (#9198)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board currently uses issues for execution, but longer-lived content work needs a separate object that can survive beyond a single task thread. > - The Cases subsystem adds an experimental, company-scoped record for content artifacts and their supporting metadata. > - The backend needs durable storage, API routes, revision history, issue linkage, and company-boundary enforcement before the UI can depend on Cases. > - The UI needs an opt-in navigation surface, list/detail views, reference chips, and issue-page context so operators can inspect Cases without making them the default workflow. > - The agent-facing skills need a contract for creating and updating Cases so automated content workflows can dogfood the feature. > - This pull request ships that experimental end-to-end path behind the `enableCases` flag. > - The benefit is a first-class place to collect content work, references, attachments, revisions, and related execution threads without polluting the core issue model. ## Linked Issues or Issue Description No public GitHub issue exists for this experimental feature. Feature request fields: ### Problem Content-oriented work such as release notes, announcements, docs, and campaigns can span many execution issues, which makes the final artifact hard to find and reason about after the execution thread moves on. ### Proposed solution Add an experimental Cases object that is company-scoped, linked to issues, queryable through the API, inspectable in the board UI, and writable by agent workflows through documented conventions. ### Alternatives considered Continue encoding content artifacts directly in issues or documents only. That keeps the data model smaller, but it does not give operators a stable artifact-centric view or a clean way to link related execution history. ### Roadmap alignment Checked `ROADMAP.md`; this PR does not duplicate an existing planned core roadmap item. ## What Changed - Added the `cases` data model, migration, schema exports, and experimental `enableCases` instance setting. - Added company-scoped Cases API routes for list/detail/update, issue links, revisions, children, activity events, annotations, attachments, and idempotent agent-oriented upserts. - Scoped case and issue lookup helpers before access checks so inaccessible cross-company identifiers resolve as not found rather than leaking existence. - Fixed case PATCH timestamp handling so non-status updates cannot overwrite `completedAt` from a stale pre-transaction row snapshot. - Moved Cases list type/status/project filters into the server request before the server-side limit is applied, including multi-select filters and no-project filtering. - Added backend route coverage for creation, updates, idempotency, issue linking, attribution, company-boundary enforcement, OpenAPI registration, list filtering, timestamp patch behavior, and inaccessible lookup regressions. - Added the experimental Cases UI surface: sidebar entry, gated routes, list filters/grouping, detail overview, activity, revisions, children, attachments, and issue-page case rail. - Added case reference rendering and company-prefixed case href generation so case links resolve directly inside the active company route. - Added Paperclip skill documentation for agent workflows that create or update Cases. - Wired release-content skills to emit Cases for dogfooding. - Rebased onto current `master` and renumbered the Cases migrations to `0143`/`0144` after the latest upstream migration sequence. ## Verification - Current PR head: `ecc13be0d`. - Rebased on current `master` (`606aa4f266`) and pushed to the existing PR branch. - `git diff --check origin/master...HEAD` — passed before the first update push; subsequent committed diffs were also checked with `git diff --check` before commit. - Guardrails checked: no `pnpm-lock.yaml` changes, no `.github/workflows` changes, and changed-file count is below the Greptile 100-file limit. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts src/__tests__/instance-settings-service.test.ts src/__tests__/openapi-routes.test.ts` — passed, 3 files / 26 tests before review-fix commits. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/cases-routes.test.ts` — passed after each server-side Greptile fix, latest 1 file / 15 tests. - `pnpm --filter @paperclipai/server typecheck` — passed after the timestamp and lookup fixes. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx src/pages/CaseDetail.test.tsx src/pages/CompanySkills.test.tsx src/App.cases-routing.test.tsx` — passed, 4 files / 30 tests before review-fix commits. - `pnpm --filter @paperclipai/ui exec vitest run src/pages/Cases.test.tsx` — passed after the list-filter fix, 1 file / 12 tests. - `pnpm --filter @paperclipai/ui typecheck` — passed after the list-filter fix. - `pnpm check:token-gates` — passed after UI changes. - Remote PR checks on head `ecc13be0d` are green: Paperclip CI, build, typecheck, test matrix, e2e, Canary Dry Run, policy, commit review, Superagent Security Scan, Socket, Snyk, and Greptile passed; Storybook visual regression is skipped and security-review is neutral. - Greptile Review: 5/5 confidence, zero unresolved Greptile threads. ## Risks - Medium feature risk because this introduces a new experimental domain object across database, server, shared contracts, skills, and UI. - The feature is gated behind `enableCases`, which limits default operator exposure while the model is exercised. - Case links now prefer company-prefixed hrefs; the unprefixed redirect remains for externally entered URLs. - Cases list filtering now sends multi-select filters to the server before limiting; the UI still applies the same local filters as a second pass for ancestor/context rows. - Migrations were renumbered on top of current master; the SQL uses guarded `IF NOT EXISTS` / `ADD COLUMN IF NOT EXISTS` patterns where relevant for safer replay. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex in the Paperclip local coding environment was used for this PR curation, rebase verification, review-fix implementation, push, and PR description update. The runtime exposes tool use and shell execution; context-window size is not exposed by this Paperclip adapter. Several implementation commits also include AI co-author trailers recorded in git history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
606aa4f266 |
feat(search): filters, sorting, operators & command-palette parity (#9327)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company search is the primary way operators find issues, comments, documents, artifacts, agents, and projects across a busy company > - Search previously supported only a bare text query: no way to narrow by status/assignee/project/label/date, no sort control, no typed operators, and weak relevance/snippets meant hunting through noise > - As companies accumulate tens of thousands of items, unfiltered single-sort search stops scaling for day-to-day operator workflows > - This pull request adds a full filtering model (filter bar, chips, mobile sheet, URL state), sort modes, typed query operators (`status:`, `assignee:`, `type:`, …) with command-palette parity, relevance/snippet/deep-link improvements, zero-results recovery, and the supporting shared validators, backend service work, and DB indexes > - The benefit is that operators can go from a vague query to the exact item in a couple of keystrokes, on desktop and mobile, with shareable filtered-search URLs ## Linked Issues or Issue Description No existing public GitHub issue; describing the underlying feature request inline (per feature_request template): - **Problem:** Company search accepted only a plain text query. Users could not filter results by status, assignee, project, label, or recency; could not change result ordering; and got no guidance when filters emptied the result set. - **Desired solution:** Structured search filters (UI controls + typed query operators + URL parameters), selectable sort modes, better relevance and snippets with exact deep links, and parity between the search page and the command palette. - **Alternatives considered:** Client-side filtering of unfiltered results (does not scale past the fetch limit); a separate "advanced search" page (splits the surface and duplicates state handling). Related (not duplicate) PRs found while searching: #4848 (issue search query planning), #8235 (search rate limiting). ## What Changed - **Shared contract:** new search filter/sort/count/zero-results types and validators in `packages/shared` (`validators/search.ts`, types index). - **Backend:** `server/src/services/company-search.ts` supports issue filters, sort modes, per-filter option counts, snippets, artifact visibility, and zero-results loosen suggestions; single-statement match replaces per-scope scans and predicates are trigram-index compatible (~3.7s → ~350ms on a live 14.8k-hit corpus). - **DB:** migration `0142_company_search_sort_indexes.sql` adds the supporting indexes. - **Search page (`ui/src/pages/Search.tsx`):** filter bar, removable chips, mobile filter sheet with result-count preview, sort menu, URL round-tripping, zero-results recovery UI. - **Query operators (`ui/src/lib/search-query-parser.ts`):** typed operators parsed into filters, operator autocomplete, filter pills. - **Command palette:** operator-aware parsing and full-search handoff. - **Stale-operator fix (latest commit):** typed operator filters are no longer folded into persistent URL-filter state, so deleting a token (e.g. removing `status:blocked` from the input) actually removes the filter from subsequent requests; filter-control edits materialize control state and strip typed tokens so a removed chip cannot resurrect from the input. ## Verification - `cd ui && npx vitest run src/pages/Search.test.tsx` — 19 tests including two new red→green regressions for the stale-operator paths (both fail on the previous commit, pass now). - `cd ui && npx vitest run src/components/CommandPalette.test.tsx` and `cd server && npx vitest run src/services/company-search-service.test.ts` — operator parity and backend filter/sort/count coverage. - `cd ui && npx tsc --noEmit` — clean. - Manual: open `/search`, type `auth status:blocked`, confirm the status filter applies; delete `status:blocked`, confirm results are unfiltered again; drive the same filters from the filter bar/chips/mobile sheet and confirm the URL round-trips (reload/back/forward preserves state). - Full end-to-end QA pass (9/9 acceptance checks) against the wireframes on desktop (1280px) and mobile (390px) with a live API and browser automation. ## Risks - Additive migration (indexes only, no data rewrites) — safe to roll forward; index creation cost is paid once at migrate time. - Search request shape gains optional parameters only; old clients keep working. - Behavioral shift: filter-control edits now strip typed operator tokens from the query text (their values persist as filter state) — deliberate, so removed filters stay removed. - Ranking changes alter result ordering for existing queries; covered by service tests and the QA pass. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic, extended thinking + tool use) — stale-operator-filter fix, regression tests, PR preparation. - GPT-5 Codex (`codex_local` adapter) and Claude Opus 4.6 (`claude-opus-4-6`) — earlier implementation phases (backend contract, filter UI, operators, ranking) under agent orchestration. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — pre-existing branch name retained to avoid closing/reopening the PR - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing docs affected) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending re-run on latest commit) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending re-review of the stale-filter fix) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cec0fc249a |
[codex] Parallelize release verify workflow (#9168)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Releases publish the same app and package set that operators install, so release verification should keep full release-strength coverage. > - The release workflow currently verifies stable and canary releases with one serial job that typechecks, runs all tests, and builds. > - The PR workflow already proves the test surface can be split into grouped general suites and serialized shards without changing coverage. > - This pull request extracts the release verify work into a reusable workflow and fans out the independent lanes. > - The benefit is faster stable and canary release verification while preserving the existing publish and preview gates. ## Linked Issues or Issue Description No public GitHub issue exists for this CI improvement. **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Release verification spends most of its wall time in a single serial test step even though the same stable test surface is already partitioned for PR CI. Stable dispatches and master-push canaries therefore wait on one long runner after setup, typecheck, tests, and build run sequentially. **Proposed solution** Add a reusable release verification workflow with parallel typecheck, grouped general tests, serialized test shards, and build lanes. Have both stable and canary release verification call it with the ref they need to verify. **Alternatives considered** Keeping the serial `pnpm test:run` job preserves the old shape but keeps stable and canary releases waiting on one long runner. Skipping verification when a source SHA already has green CI would be faster, but adds stale-check and lookup risk beyond this change. **Roadmap alignment** No overlapping item found in `ROADMAP.md`; this is release CI maintenance. **Additional context** The new workflow keeps the release-strength full `pnpm -r typecheck`, uses the existing stable test grouping/sharding entry points, and leaves publish/preview jobs unchanged. ## What Changed - Added `.github/workflows/release-verify.yml` as a `workflow_call` workflow accepting a `ref` input. - Split release verification into parallel `typecheck`, `general_tests`, `serialized_tests`, and `build` jobs with 20-minute lane timeouts. - Mirrored the PR workflow's stable test partition: `general-server` shards 1-3, `general-workspaces-a`, `general-workspaces-b`, and four serialized shards. - Replaced `release.yml` `verify_canary` and `verify_stable` job bodies with calls to the reusable workflow while leaving publish and preview jobs unchanged. - Added a Node test that guards the release workflow delegation and split verify surface. ## Verification - `actionlint 1.7.12 .github/workflows/release.yml .github/workflows/release-verify.yml` - `node ./scripts/release-package-map.mjs check` - `node --test ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check` ## Risks - Release verification now starts more jobs per release event, increasing total runner setup/install minutes. This matches the existing PR CI tradeoff and should reduce release wall time substantially. - The called workflow checks out the requested ref shallowly. That is intentional for verify lanes; publish and preview jobs still retain their existing full-history checkouts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in local tool-use mode with shell execution, repository editing, GitHub connector access, and medium reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f3ca4d24bc |
fix: repair dirty/foreign-branch execution worktrees (#9297)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces are the bridge between Paperclip's control plane and a local agent's checked-out repository state. > - When a workspace is restored after a failed or interrupted run, the recorded branch can disagree with the branch currently checked out on disk. > - A clean branch mismatch can be reconciled safely, but a dirty mismatch needs a lossless path that does not discard uncommitted agent work. > - This pull request adds a quarantine-and-restore path that saves dirty work to a rescue branch, restores the recorded branch, and exposes the repair from the board UI and run page. > - The benefit is that operators can recover wedged execution workspaces without losing work or moving another live branch unexpectedly. ## Linked Issues or Issue Description No public GitHub issue exists for this workspace-recovery failure, so this PR includes the bug report inline. **What happened** A git worktree-backed execution workspace could become wedged when Paperclip expected one branch but found a different checked-out branch with dirty tracked or untracked files. The existing safe repair path refused the restore, leaving the source task blocked with no lossless one-click recovery path. **Expected behavior** Paperclip should preserve dirty work before restoring the recorded workspace branch. If another live workspace claims the checked-out branch, or an attached runtime service is active, the repair should refuse with clear operator-facing evidence instead of risking work loss or file contention. **Steps to reproduce** Create a git worktree execution workspace whose persisted branch name differs from the checked-out branch, add dirty tracked or untracked files in that worktree, then trigger workspace validation or use the branch reconcile endpoint. Before this change, the dirty mismatch remained blocked because Paperclip had no quarantine restore mode. **Paperclip version or commit** Observed on the pre-fix workspace-recovery implementation. Verified on this PR head after rebasing onto current `master`. **Deployment mode** Local trusted development/worktree deployments using git worktree execution workspaces and optional workspace runtime services. ## What Changed - Added dirty-worktree quarantine repair that creates a rescue branch, commits dirty tracked and untracked files there, restores the recorded branch, writes audit comments/activity, and preserves the live foreign branch ref. - Added `quarantine_restore` branch reconcile API support, recovery-action resolution, source-task wake behavior, execution-review preservation, claimant refusal, runtime-service refusal, and coverage for the non-transactional git ordering. - Added board UI controls for the repair action in the recovery card plus a compact failed-run workspace recovery surface that uses the same reconcile handlers. - Hardened Greptile follow-up cases by best-effort restoring the recorded branch after a mid-sequence rescue commit failure and by refusing quarantine restore while attached runtime services are active. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspaces-service.test.ts -t "quarantine_restore"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts -t "workspace dirty quarantine branch repair"` - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts -t "repairs clean unrecorded branch drift|adopts unrecorded forward branch drift"` - `pnpm exec vitest run server/src/__tests__/workspace-runtime-routes-authz.test.ts` - `pnpm --filter @paperclipai/server typecheck` - Earlier PR verification covered the route, service, heartbeat, UI component, and run-page recovery surfaces; Cutter posted public preview screenshots for the repair popover and run-page panel at https://github.com/paperclipai/paperclip/pull/9297#issuecomment-4926934211. - GitHub PR checks are green on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. - Greptile is 5/5 with zero annotations and no unresolved review threads on `dcac76b05f4cf6e1ee16544c2831d83c7857e475`. ## Risks - Moderate risk because the change intentionally runs git commands against local worktrees; the implementation refuses dirty repair when another claimant or active runtime service is detected and records rescue refs for auditability. - Compatibility / release-note callout for self-hosted operators: existing instances that left `enableWorkspaceBranchReconcileForward` unset now get automatic forward branch reconciliation during heartbeat workspace recovery. Operators who want the previous advisory-only behavior can set `experimental.enableWorkspaceBranchReconcileForward` to `false`; dirty quarantine repair can likewise be disabled with `experimental.enableWorkspaceDirtyQuarantineRepair: false`. - If the git rescue succeeds but a later database write fails, the worktree may already be restored while the recovery action remains open; this ordering is documented in code because git side effects cannot participate in the database transaction. - UI risk is limited to the workspace recovery surfaces and covered by component tests plus the existing Cutter visual preview. ## Model Used OpenAI Codex coding agent based on GPT-5, with repository tool use, shell execution, and local test execution. Earlier preserved commits on this branch also show Claude Code / Claude Opus 4.8 assistance in their commit metadata. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] Branch naming exception documented: this PR preserves the existing worktree branch requested for publication while keeping the PR title and body public-facing - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4d898aa7af |
feat(prompt): one-time execution-workspace branch guard in wake prompt (PAP-13326) (#9319)
## Summary
When a task runs in a branch-pinned execution workspace, agents
sometimes switch or rename the workspace branch, which breaks the
worktree contract. This adds a short, one-time prompt hint telling the
agent to stay on the pinned branch.
- **heartbeat.ts**: after the execution workspace is resolved, attach
`executionWorkspace: { branchName }` to the wake payload (only when a
branch pin exists — agent-home runs without a branch are untouched).
- **server-utils.ts (adapter-utils)**: normalize the new payload field
and render one bullet in `renderPaperclipWakePrompt`:
> `- execution workspace branch: you are running in an execution
workspace on branch \`<name>\`. Do not switch, rename, or re-point this
branch; keep all commits on it.`
- The hint renders **only on non-resumed sessions** — resume-delta
prompts skip it, so it appears the first time an issue's session starts,
not on every turn, and it never pollutes the issue thread. One renderer
change covers every adapter (claude, codex, cursor, gemini, grok,
opencode, pi, hermes, acpx engine) with zero per-adapter edits.
## Tests
- `server-utils.test.ts`: branch guard renders on first prompt, absent
on resumed-session prompts, absent when no branch is pinned; payload
round-trips through `stringifyPaperclipWakePayload`.
- `heartbeat-workspace-branch-containment.test.ts`: the finalize-path
adapter mock now asserts the wake payload the adapter receives carries
the branch pin matching `context.paperclipWorkspace.branchName`
(end-to-end heartbeat wiring, embedded postgres). All 6 pass.
- Full `server-utils` (56) and acpx-engine execute (34) suites green;
adapter-utils typechecks clean; no new server tsc errors.
PAP-13326
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
176645187c |
Fix request storm polling and issue-list coalescing (#9190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI keeps issue, agent, activity, and run state fresh through polling across several pages and sidebar surfaces > - When multiple components or browser tabs poll the same company data at the same time, the API can receive bursts of duplicate issue-list requests > - Those duplicate requests increase database and server load without returning meaningfully different data > - This pull request adds server-side compression/coalescing plus client-side visibility-aware and cross-tab shared polling > - The benefit is lower request volume during normal board usage while preserving fresh UI data for active users ## Linked Issues or Issue Description No exact public GitHub issue was found for this request. Problem: - The board can issue redundant polling requests for the same issue-list data from multiple UI surfaces and tabs. - In busy operator sessions, those bursts can trigger request-storm behavior and unnecessary issue-list load. - Expected behavior is to reuse identical in-flight work server-side and reduce hidden-tab or duplicate-tab polling client-side while preserving normal refresh behavior. Related public context found during duplicate search: - #8206 covers a different board UI 404-storm scope. - #5165 covers separate issue-list behavior around page-size truncation. ## What Changed - Added API compression middleware foundation and a company/created-at index for heartbeat run access. - Added server-side issue-list request storm detection and identical in-flight request coalescing. - Added UI fetch metadata, visibility-aware polling, and request deduplication for issue/activity/client calls. - Added cross-tab shared polling primitives and wired them into the sidebar, inbox, dashboard, issue, project, routine, and agent surfaces. - Resolved the latest `master` migration collision by keeping upstream `0140_built_in_managed_resources.sql` and renumbering this branch's heartbeat-run index migration to `0141_heartbeat_runs_company_created_at_index.sql`; the SQL uses `CREATE INDEX IF NOT EXISTS` for idempotency. - Stabilized server heartbeat cleanup tests exposed by the PR check matrix. - Fixed the Greptile compression follow-up by weakening strong ETags on encoded JSON responses and bypassing compression for streamed/download responses. ## Verification - `pnpm exec vitest run ui/src/components/IssuesList.test.tsx ui/src/pages/Inbox.test.tsx` — passed after resolving the latest `master` conflict in `IssuesList.tsx` and updating the 200-result cap expectations. - `pnpm check:token-gates` — passed after the UI conflict resolution. - `jq -e '.entries | length as $n | (map(.idx) | unique | length == $n) and (map(.tag) | unique | length == $n)' packages/db/src/migrations/meta/_journal.json` — passed after renumbering the migration to `0141`. - `pnpm exec vitest run server/src/__tests__/api-compression.test.ts` - passed after the compression follow-up. - `pnpm exec vitest run server/src/__tests__/issue-list-assignee-filter-routes.test.ts` - passed after the compression follow-up. - `pnpm --filter @paperclipai/server typecheck` - passed after the compression follow-up. - Greptile Review for head `8dbddac41ec273fda404100b4981ddb912fad57b` - passed after the latest conflict/migration fix; all Greptile review threads are resolved. - `pnpm exec vitest run server/src/__tests__/api-compression.test.ts server/src/__tests__/issue-list-assignee-filter-routes.test.ts ui/src/api/client.test.ts ui/src/api/issues.test.ts ui/src/lib/polling.test.ts ui/src/lib/cross-tab-poll.test.ts ui/src/pages/Inbox.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/lib/issueDetailCache.test.ts` — passed. - `pnpm exec vitest run ui/src/api/client.test.ts` — passed. - `pnpm exec vitest run ui/src/pages/Inbox.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/lib/issueDetailCache.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-worktree-suppression.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/low-trust-red-team-routes.test.ts` — passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-finalize-branch.test.ts` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - GitHub PR checks for head `8dbddac41ec273fda404100b4981ddb912fad57b`: all GitHub Actions/status checks passed; Greptile, Superagent, Socket, Snyk, build, typecheck/release registry, general tests, serialized server suites, e2e, canary dry run, policy, and commitperclip review are green; Storybook visual regression and security-review were skipped/neutral by policy. - Confirmed this branch does not include `pnpm-lock.yaml` or `.github/workflows` changes. ## Risks - Medium risk: issue-list coalescing changes request timing and cache semantics for a hot API path. - Medium risk: cross-tab polling uses browser coordination primitives, so older or unusual browser environments need fallback behavior to stay correct. - Low migration risk: the new index migration is ordered after current `master` and uses `CREATE INDEX IF NOT EXISTS`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent, GPT-5-based model, tool-enabled with shell/git execution. Exact hosted deployment identifier and context-window size were not surfaced in the agent runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4856558fd9 |
Fix skills routes to skip user secret resolution
Skip user_secret_ref bindings when resolving adapter config for agent skill listing and sync paths, while keeping normal runtime resolution strict. Add route and service regression tests for required user-secret refs in adapter config. |
||
|
|
8b6a06ee25 |
[codex] Add built-in agents and Reflection Coach bundle (#9206)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators need first-party agent capabilities for repeatable company work, not just manually created one-off agents. > - Built-in agents need to behave like normal company-scoped agents while preserving approval gates, permissions, budgets, and audit trails. > - Reflection and coaching work also needs bundled instructions, skill content, and a routine so the feature can be installed and reset predictably. > - The API, database, UI, portability, and tests all need to agree on the built-in lifecycle from not provisioned through setup, approval, ready, paused, and reset. > - This pull request adds built-in agent provisioning and the Reflection Coach bundle end-to-end. > - The benefit is a safer first-party path for Paperclip-managed agents without bypassing the same governance model used for operator-created agents. ## Linked Issues or Issue Description No public GitHub issue was found for this exact built-in agent and Reflection Coach bundle work. Problem/motivation: - Paperclip did not have a first-party built-in agent lifecycle for product-owned agents. - Bundled agent resources such as default instructions, skills, and routines needed managed ownership and reset semantics. - Approval-gated companies needed built-in setup to preserve requested adapter, budget, manager, and permission state through board approval. - The board UI needed clear built-in badges, setup affordances, readiness state, and bundle status without exposing secrets. Proposed solution: - Add a company-scoped built-in agent registry, provisioning/reset/reconcile/status APIs, and Reflection Coach bundled resources. - Track bundled managed resources in the database with idempotent migration behavior. - Reuse existing agent approval, authorization, budget, and activity-log paths instead of creating a bypass. - Add UI setup, badges, gates, bundle panels, and route coverage for built-in agents. Duplicate search: - Searched GitHub PRs for `built-in agents Reflection Coach repo:paperclipai/paperclip`; only this PR was returned. - Searched GitHub issues for the same query; no public issues were returned. ## What Changed - Added built-in agent definitions, lifecycle state derivation, provisioning, reset, reconcile, status, and routine-control routes. - Added the `built_in_managed_resources` migration and schema exports for bundled instructions, skill, and routine ownership. - Added the Reflection Coach built-in bundle with default instructions, skill catalog content, routine template, default permissions, and managed-resource drift handling. - Added approval-aware provisioning behavior that preserves requested adapter config, budgets, manager assignment, and built-in permissions through hire approval. - Added authorization and mutation gates for built-in agent and skill changes, including consented Reflection Coach change paths. - Added UI surfaces for built-in agent setup, roster/detail badges, readiness gates, bundle status, routine controls, and route filtering. - Added company import/export and validator coverage for built-in managed resources and low-trust/red-team presets. - Addressed Greptile follow-ups for pending approval reconciliation, consent-gate error propagation, config-read authorization fallback, approval-path manager preservation, and non-model adapter provisioning. ## Verification Local verification: - `git diff --check public/master..HEAD` passed. - `pnpm check:token-gates` passed with all gates clean. - `pnpm exec vitest run ui/src/components/ConfigureBuiltInAgentModal.test.tsx` passed: 1 file, 4 tests. - `pnpm exec vitest run ui/src/components/EntityRow.test.tsx ui/src/pages/Agents.test.tsx ui/src/components/BuiltInAgentGate.test.tsx ui/src/components/ConfigureBuiltInAgentModal.test.tsx ui/src/components/BuiltInBundlePanel.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx ui/src/pages/Routines.test.tsx` passed: 7 files, 64 tests. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/built-in-agents.test.ts src/__tests__/authorization-service.test.ts src/__tests__/company-skills-routes.test.ts` passed: 3 files, 91 tests. - `pnpm --filter @paperclipai/db check:migrations` passed. - `pnpm -r typecheck` passed after the rebase; `pnpm --filter ui typecheck` passed after the final UI review fix. Remote verification on latest head `1c61f693a4ec881d739022b0e75a8ca8bf8c2cd8`: - Merge state: `CLEAN`. - Greptile: `5/5`, zero unresolved Greptile threads. - PR check rollup: all checks successful, neutral, or skipped as expected. - Passing gates include Build, Typecheck + Release Registry, all server shards, all workspace shards, all serialized server suites, e2e, Canary Dry Run, policy, review, verify, Socket, Superagent, and Snyk. ## Risks - This adds a new managed-resource table and migration; the migration uses idempotent create/add/index guards and passed migration safety checks. - Built-in agent provisioning touches approval and authorization paths; tests cover pending approval preservation, stale retry rejection, consent gates, and config-read fallback behavior. - Reflection Coach creates managed instructions, skill, and routine resources; drift/reset behavior is covered by service tests and redacted API responses. - Non-model adapter setup now provisions a `needs_setup` built-in row before command/endpoint fields are complete; this matches the server lifecycle and is covered by the setup modal regression test. ## Model Used OpenAI Codex coding agent based on GPT-5. Exact hosted model ID, context-window size, and reasoning-mode labels are not exposed in this runtime; tool use, shell execution, GitHub CLI/API access, and local code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b13eb5b2b5 |
Skill Studio: three-pane skill IDE with sandboxed test runs (#9241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Skills Manager gives operators a reusable skill layer, but iteration still required manual edits, ad hoc prompts, and indirect run inspection. > - Skill authors need a focused workflow for editing skill files, saving representative test inputs, and running those inputs through an agent without exposing harness tasks as normal company work. > - The backend therefore needs durable test inputs, reusable run templates, hidden harness issues, scoped run execution, retention metadata, and read-containment rules around hidden work. > - The frontend needs a three-pane Studio that keeps skill files, saved inputs/templates, and run output/history visible together while preserving the existing design system and token rules. > - This pull request ships that Skill Studio surface end to end: database migrations, shared contracts, server APIs/services, hidden harness execution behavior, UI routes/components, and focused tests. > - The benefit is faster and safer skill iteration, with inspectable outputs and fewer ways for internal harness work to leak into normal task lists, costs, or adjacent read APIs. ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Feature request summary: - Problem: Skill authors need to edit and test company skills in one place instead of switching between the skill detail page, task creation, run output, and manual prompt history. - Proposed solution: Add a Skill Studio workbench with saved inputs, reusable templates, hidden sandboxed test runs, live run status, output inspection, run history, rerun/delete controls, and frontmatter-aware editing. - Expected users: Paperclip operators and agent-company maintainers who create, fork, import, and tune skills. - Related public PRs: Supersedes #9205, which was replaced so the public PR branch name follows contributor policy. - Duplicate search: searched public GitHub issues and PRs for "Skill Studio"; no other active public issue or PR directly covers this feature. ## What Changed - Added database migrations for Skill Studio test inputs, test runs, test run retention, and reusable run templates. - Added shared Skill Studio types, validators, route helpers, frontmatter utilities, and status handling. - Added server services and routes for saved inputs, test runs, templates, reruns, terminal-run deletion, hidden harness issue execution, and run-detail hydration. - Strengthened hidden-issue read containment across issue-adjacent routes and cost rollups used by skill test harness work. - Added the Skill Studio UI with skill file editing, frontmatter editing, saved inputs, templates, run creation/cancel/rerun/delete flows, output rendering, history, route support, and responsive pane behavior. - Added focused backend, shared, and UI tests for the new APIs, routing logic, editor/run behavior, hidden-issue containment, and migration safety. - Rebased onto current `master`, removed the generated lockfile diff from the PR, and verified no workflow files are changed. ## Verification - [x] `pnpm --filter @paperclipai/db check:migrations` - [x] `pnpm check:token-gates` - [x] `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/company-skill-test-runs-service.test.ts ui/src/lib/skill-studio.test.ts ui/src/pages/SkillStudio.test.tsx` — 5 files, 132 tests passed - [x] Greptile review on the latest PR head - [x] GitHub PR checks on the latest PR head ## Risks - Medium risk because this is a broad feature touching database schema, server orchestration, issue visibility, and a large UI surface. - Hidden harness issue containment is security-sensitive; this PR includes regression coverage for adjacent read paths and cost rollups. - The new migrations are additive and use idempotent guards where applicable, but deployed databases that previously tested draft migration numbers should still be checked carefully. - The UI depends on a new resizable panels package in `ui/package.json`; the lockfile is intentionally left to repository automation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with shell, git, and GitHub CLI tool use. Earlier feature commits include assistance from other Paperclip coding agents; this PR preparation, rebase, cleanup commit, and PR body were completed by OpenAI Codex in a Paperclip worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1c75a46c10 |
feat(ui): add waiting-on-live-work blocked notice (#9298)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators rely on the issue detail thread to understand whether a task is blocked, live, or waiting on another task. > - A blocked issue can have a healthy blocker chain where downstream work is actively running and the parent will resume automatically. > - Showing that case with the same amber blocked notice as a stalled or attention-needed blocker makes the state look more severe than it is. > - The UI already receives blocker-attention state, blocker summaries, and company live-run ids, so this can be clarified without a new API shape. > - This pull request adds a blue "Waiting on live work" notice for covered blocker chains while preserving the existing amber notice for the other blocked states. > - The benefit is that operators can distinguish healthy queued work from blocked work that needs intervention. ## Linked Issues or Issue Description Refs #3820 Refs #8271 Related PR: #3877 Supersedes #9295 ## What Changed - Added a blue `IssueBlockedNotice` variant when `blockerAttention.state` is `covered` and the blocker chain has live work. - Rendered blocker-chain progress as done, running, and queued steps, including a "Now running" row for live terminal blockers. - Preserved the existing amber blocked notice for stalled, attention-needed, ordinary blocked, and successful-run handoff states. - Plumbed the existing `liveIssueIds`, `blockedBy`, and `blockerAttention` data from issue detail into the chat-thread blocked notice. - Added regression coverage around the covered live-work state, the no-confirmed-live fallback, numeric step ordering, and amber fallback states. - Hardened a low-trust server route test cleanup helper so CI deletes heartbeat run events before deleting heartbeat runs. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run ui/src/components/IssueBlockedNotice.test.tsx` - GitHub PR workflow is green on head `52ab6d9076ce233c183bf7133fa666e8597b6765`. - Greptile check is green on head `52ab6d9076ce233c183bf7133fa666e8597b6765` with zero unresolved review threads. ## Risks Low runtime risk: the product change is frontend-only and uses data already returned to the issue detail page. The main risk is visual regression in the blocked notice; the change keeps non-covered states on the existing amber path and adds focused regression coverage. The server-side change is test-only cleanup for an existing CI shard failure. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex using GPT-5, tool-use enabled in a repository workspace. The runtime did not expose a more specific model build id or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7cf0d3ebb0 |
Require health readiness for Paperclip dev services (#9269)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents often run against managed workspace runtime services, including reusable Paperclip dev servers > - A running process and an open root URL are not enough to prove the Paperclip API is actually ready > - If the API health endpoint is still failing, agents can reuse a service that looks alive but cannot safely serve the board or API clients > - This pull request makes Paperclip dev runtime readiness probe the resolved `/api/health` endpoint > - The benefit is that runtime service reuse waits for the same health signal operators and agents depend on ## Linked Issues or Issue Description No matching public GitHub issue was found. Public duplicate search found no open PR for "workspace runtime health readiness". Bug report: ### What happened? A managed Paperclip dev runtime service could satisfy HTTP readiness at the exposed base URL even when the Paperclip health endpoint was returning an unhealthy status. ### Expected behavior Paperclip dev runtime services should not be considered ready until their health endpoint succeeds. ### Steps to reproduce 1. Start a workspace runtime service named `paperclip-dev` whose base URL responds successfully. 2. Make that same service return HTTP 503 from `/api/health`. 3. Ask Paperclip to ensure the runtime service for a run. 4. Observe that the service can be reused even though the API health endpoint is not ready. ### Paperclip version or commit Current `origin/master` before this PR. ### Deployment mode Local workspace runtime service management. ## What Changed - Resolve Paperclip dev runtime readiness checks to the service health URL before polling. - Surface readiness errors with the actual health URL that failed. - Add a regression test that fails when `/api/health` returns HTTP 503 even if the service process is running. ## Verification - `pnpm exec vitest run server/src/__tests__/workspace-runtime.test.ts` — 81 tests passed. - `git diff --check origin/master...HEAD` — passed. ## Risks - Low to medium risk. This tightens readiness for Paperclip dev runtime services, so a service that previously looked ready while unhealthy will now fail fast instead of being reused. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, tool-enabled shell workflow. Exact hosted model variant and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
719da5f9b5 |
Harden environment deletion and expose delete blast radius (#9250)
## Thinking Path > - Paperclip manages AI agents that each have an associated execution environment (local, Kubernetes, etc.) > - Instance administrators can create and delete environments; currently the DELETE endpoint has no protection against deleting managed or in-use environments > - Deleting the managed local environment or the instance-default environment would break all agents using those environments with no path to recovery > - The endpoint also suffered a TOCTOU race: a check-then-delete pattern allowed the managed-local or default guard to pass if the environment's role changed between the read and the delete > - This pull request adds a blast-radius read endpoint so admins can preview impact, hard-blocks the dangerous deletes atomically, cleans up all dependent references after a valid delete, and fixes a concurrent creation race in ensureLocalEnvironment ## Linked Issues or Issue Description Fixes #9251 ## What Changed - **New endpoint** `GET /api/environments/:id/delete-blast-radius` (instance-admin gated): returns reference counts (agent defaults, workspace selections, issue selections, project selections, secret bindings, active leases, active setup sessions) and blocking reasons — no config, env-var values, or secret data returned. - **Atomic delete guard** `environmentService.removeIfDeletable(id)`: performs the DELETE with an inline `WHERE driver != 'local' AND NOT EXISTS (instanceSettings where defaultEnvironmentId = id)` predicate, eliminating the TOCTOU race between the app-level check and the DB write. - **Route hardening**: `DELETE /environments/:id` now calls `getDeleteBlastRadius` first (app-level check + logging), then calls `removeIfDeletable` (atomic guard). If the atomic guard returns null the route fetches a fresh blast-radius snapshot and rejects with a 409 Conflict carrying `deleteBlockedReasons`. - **Reference cleanup on valid delete**: after a successful delete, the route clears environment selections on all company execution workspaces, issues, and projects; syncs env-var secret bindings to `{}` (removing bindings for the deleted environment); syncs config secret refs to `[]` for the environment target; and removes the SSH private-key secret if one was stored. - **Race fix in `ensureLocalEnvironment`**: the insert-or-nothing path now catches a `environments_name_idx` unique-constraint violation and falls through to the existing SELECT, treating the name conflict as idempotent. - **Shared types**: `EnvironmentDeleteBlastRadius` and `EnvironmentDeleteBlockedReason` exported from `@paperclipai/shared`. - **OpenAPI**: registers the new blast-radius endpoint; updates the delete-environment response schema to document 403/404/409. - **Tests**: 56 existing environment-route and service tests continue to pass; new service-level regression tests assert the atomic guard rejects `local`-driver environments and instance-default environments and succeeds for deletable ones. ## Verification ``` corepack pnpm exec vitest run \ server/src/__tests__/environment-routes.test.ts \ server/src/__tests__/environment-service.test.ts # 56 tests, all passing corepack pnpm --filter @paperclipai/shared typecheck node scripts/ensure-plugin-build-deps.mjs cd server && ../node_modules/.bin/tsc --noEmit ``` ## Risks - **Blast-radius endpoint auth**: guarded by `assertCanAccessInstanceEnvironments`, the same gate as the existing environment-list and delete routes. Non-admin callers receive 401/403 before any data is returned. - **Atomic guard may reject a delete that the app-level check passed**: this is intentional — it means the environment became protected between the read and the write. The caller receives a fresh blast-radius snapshot explaining why. - **Secret cleanup ordering**: cleanup runs after the atomic DELETE succeeds, in parallel across companies. If cleanup partially fails the environment row is already gone; partial-cleanup state is recoverable by re-running the sync operations. Risk: low — these are idempotent upsert/sync operations. - **ensureLocalEnvironment race fix**: swapping a unique-constraint error for an idempotent SELECT adds one extra query on the conflict path. This path is rare (only fires during concurrent boot) and is significantly safer than the previous behavior. - **No migration**: all changes are application-level; no schema changes required. ## Model Used - Provider: Anthropic - Model: claude-sonnet-4-6 (Claude Sonnet 4.6) - Context window: 200k tokens - Mode: agentic tool use via Paperclip agent system (Claude Code) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold.kim@paperclip.ing> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4a7a732476 |
feat(skills): add company skill fork prechecks (#9235)
Adds company skill fork precheck metadata, fork result/reassignment contracts, selected-agent reassignment during fork creation, and targeted server/shared test coverage. |
||
|
|
f616b6746c |
fix(server): clean heartbeat run scratch directories (#9234)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
38cca22b09 |
fix(server): avoid accepted-plan workspace branch freeze before child realization (#9233)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents tackle complex tasks via *plan* flows: a planner decomposes
work into child issues, which are accepted by the board and then
executed
> - When a plan is accepted, `createChild` in `issues.ts` inserts child
issues pre-bound to the parent's already-realized execution workspace —
carrying over its concrete branch ref
> - If the repository's base ref advances between plan acceptance and a
child's first heartbeat, the child inherits a stale branch that no
longer matches the current base
> - At first heartbeat the workspace validator detects the mismatch and
freezes the child ("branch freeze"), blocking it from starting any work
> - The real fix is to strip the concrete workspace binding when
creating accepted-plan children: they should receive only the unresolved
*intent* (mode, baseRef, branchTemplate) and realize a fresh workspace
from the current base on their own first heartbeat
> - This PR implements that strip, adds a regression test that proves a
post-base-advance child realizes cleanly, and also fixes
`parseIssueExecutionWorkspaceSettings` so `environmentId` is not
silently dropped on update round-trips (a latent bug that was masking
the original fix)
## Linked Issues or Issue Description
No public GitHub issue exists for this bug. Bug description follows the
bug-report template:
**What happened:** Accepted-plan decomposition pre-binds child issues to
the parent's realized execution workspace branch (`executionWorkspaceId`
+ `executionWorkspaceBranch`). When `origin/master` advances between
plan acceptance and the child's first heartbeat, the workspace branch
interlock fires and the child is permanently frozen before it can start.
**Expected behavior:** Accepted-plan children should receive only
unresolved workspace intent (mode, git strategy fields) and realize a
fresh isolated worktree from the current base on first heartbeat. A
base-ref advance between acceptance and first-run should be transparent.
**Steps to reproduce:**
1. Accept a plan that decomposes into one or more child issues
(isolated_workspace + git_worktree mode).
2. Allow `origin/master` to advance (new merge).
3. Observe the first child heartbeat: workspace validation fails with a
branch-freeze error.
**Paperclip version:** current `master` (pre-fix).
**Deployment mode:** any (affects all modes that use isolated workspace
+ git worktree strategy).
Supersedes #9227 (earlier attempt, now closed — the fix was incomplete
because `environmentId` was silently dropped during
`parseIssueExecutionWorkspaceSettings` update round-trips, causing the
child workspace to lose its environment binding; this PR includes that
fix).
## What Changed
- **`server/src/issues.ts` — `createChild` / accepted-plan decomposition
path:** strip resolved workspace fields (`executionWorkspaceId`,
concrete branch) when creating accepted-plan children; preserve only
unresolved intent fields (`mode`, `baseRef`, `branchTemplate`,
`environmentId`, runtime/provisioning settings).
- **`server/src/execution-workspace-policy.ts` —
`parseIssueExecutionWorkspaceSettings`:** preserve `environmentId`
through update round-trips (was silently dropped, causing environment to
detach on any workspace settings update).
-
**`server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts`:**
new regression test — accepted-plan child created after `origin/master`
moves realizes a fresh isolated worktree from the moved base and passes
workspace execution.
- **`server/src/__tests__/issues-service.test.ts`:** extended
workspace-linkage and `createChild` tests covering the accepted-plan
strip and the unchanged direct-child path.
## Verification
```sh
pnpm exec vitest run server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts
pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t "workspace linkage|accepted plan decomposition|createChild applies"
pnpm --filter @paperclipai/server typecheck
```
All three pass on this branch.
## Risks
**Low.** The change is scoped to the accepted-plan `createChild` code
path. The direct child / follow-up issue creation path (normal non-plan
decomposition) is unchanged and covered by existing tests. The
`parseIssueExecutionWorkspaceSettings` fix is additive — it now
preserves a field that was previously silently dropped, so no consumer
loses data.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`), Anthropic, 200K context window,
extended tool use + code generation.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (supersedes #9227)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
88bf71b84e |
Reject projectless isolated git worktree tasks (#9231)
## Thinking Path
> - Paperclip is an open-source platform for orchestrating AI agents;
agents run inside execution workspaces that range from a shared
container to full git worktrees cloned from a project repository.
> - Isolated git-worktree workspaces require a project to determine
which repository to clone — without a project the worktree base path
cannot be computed.
> - A task pinned to `isolated_workspace` + `git_worktree` with no
project was previously accepted at creation time but failed late at
dispatch with the opaque `workspace_validation_failed` /
`git_worktree_base_agent_home` error — only after the heartbeat
attempted to provision the workspace.
> - Fail-closed validation should happen in two places: (1) explicit
create/update pins that contradict the requirement are rejected at the
HTTP layer with a structured 422; (2) rows that reach the heartbeat
dispatcher with this invalid combination (e.g. through inheritance or a
retroactively-removed project) are blocked before any heartbeat run or
adapter spawn.
> - This PR adds the shared detection policy, the create/update guard in
the issues service, and the heartbeat pre-dispatch guard, together with
focused unit tests for all three layers.
> - The benefit is deterministic early failure with a clear remediation
message instead of a late, cryptic runtime error.
## Linked Issues or Issue Description
No upstream public GitHub issue — describing the problem inline
(bug-report format).
**What happened?**
Creating an issue with `executionWorkspaceSettings: { mode:
"isolated_workspace", type: "git_worktree" }` and no `projectId` was
accepted without error. The issue then became blocked at dispatch time
with the opaque message `git_worktree_base_agent_home` /
`workspace_validation_failed` — surfaced only after the heartbeat
attempted to provision the workspace.
**Expected behavior**
The platform should reject the invalid combination at create/update time
with a structured 422 that includes a clear remediation message, before
any heartbeat resource is consumed.
**Steps to reproduce**
1. Call `POST /api/issues` (or `PATCH /api/issues/:id`) with
`executionWorkspaceSettings: { mode: "isolated_workspace", type:
"git_worktree" }` and omit `projectId` (or set it to `null`).
2. Observe: request succeeds (200/201).
3. Assign the issue to an agent and watch it enter `blocked` with a
cryptic `workspace_validation_failed` error at dispatch.
**Related prior fix** — Refs #4844 (`fix(validator): reject static cwd
combined with git_worktree strategy`) — same validation area, different
dimension (static cwd vs. missing project).
**Paperclip version / commit**
Latest `master` (pre-this-PR).
**Deployment mode**
Standard (app-global server).
## What Changed
- **`execution-workspace-policy.ts`** — new shared
`detectWorkspaceWorktreeRequiresProject` function returning a stable
`workspace_worktree_requires_project` policy violation when an isolated
git-worktree task has no project, project workspace, or reusable
execution workspace; exports canonical remediation text used by both the
HTTP guard and the heartbeat guard.
- **`issues.ts`** — create and update paths check the new policy before
persisting; explicit pins to `isolated_workspace` / `operator_branch` +
`git_worktree` with no project are rejected with a 422 including the
policy code and remediation text.
- **`heartbeat.ts`** — pre-dispatch preflight checks the same policy for
rows that reach the heartbeat with the invalid combination (e.g. through
inheritance); such rows are marked `blocked` with a skipped wakeup
request, durable issue comment, and activity log before any heartbeat
run or adapter spawn.
- **`execution-workspace-policy.test.ts`** — focused policy-layer unit
tests for detection logic and remediation text.
- **`issues-service.test.ts`** — create/update 422 guard tests for the
new policy.
- **`heartbeat-workspace-branch-containment.test.ts`** — pre-dispatch
blocking test for inherited/ambiguous invalid rows; also fixes a cleanup
race in the existing test suite.
## Verification
```sh
pnpm exec vitest run \
server/src/__tests__/execution-workspace-policy.test.ts \
server/src/__tests__/issues-service.test.ts \
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts
pnpm --filter @paperclipai/server typecheck
# Targeted regression
pnpm exec vitest run \
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts \
-t "blocks projectless isolated git-worktree issues before dispatch"
```
All three test files and typecheck passed locally before this PR was
opened.
## Risks
**Low risk.** The policy detection function is pure with no side
effects. The create/update guard only triggers on explicit
`isolated_workspace` or `operator_branch` + `git_worktree` pins combined
with a missing project — it does not fire on inherited settings (handled
by the heartbeat preflight), so there is no false-positive rejection
risk for valid tasks. The heartbeat guard fires before any resource is
provisioned; the only behavioral change for already-invalid rows is that
they receive a clear `blocked` status and durable comment instead of a
late cryptic error.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k context window, tool use
enabled (agentic coding). Used to implement all server-side changes and
tests in this PR.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
11177c2665 |
Fix resolved blocker wake reconciliation (#9229)
## Thinking Path > - Paperclip is an open source platform for managing AI agent companies — agents pick up issues, do work, and release checkouts in a heartbeat loop > - Issue dependency resolution is handled by `issue_blockers_resolved` wakes: when a blocking issue reaches `done`, dependent blocked issues should be woken so they can resume > - The workspace-finalize path maintained its own special-case loop to retry missed dependency wakes at run completion, duplicating logic that already exists in the shared level-triggered reconciliation backstop > - Additionally, the resolved-blocker reconciliation sweep was gated behind the broader liveness auto-recovery escalation setting — so instances that disabled auto-escalation creation also lost the baseline reconciliation sweep > - This PR routes the workspace-finalize retry through the shared backstop and decouples the reconciliation sweep from the escalation creation gate > - The benefit is simpler code (one authoritative path instead of two), and correct behaviour on instances where escalation creation is disabled ## Linked Issues or Issue Description No existing GitHub issue. Describing inline: **Bug (blocker wake reconciliation):** `issue_blockers_resolved` wakes can be missed when: 1. A `workspace_finalize` heartbeat retries dependency resolution using its own inline loop instead of the shared level-triggered backstop, making them diverge over time. 2. The `blockedByIssueIds` dependency is set (or the issue is moved to `blocked`) *after* the blocker already reached `done` — the PATCH-time wake fires on a non-blocked issue and the dependency reconciliation never catches up. 3. An assignee briefly becomes `null` between the blocker reaching `done` and the reconciliation sweep running — the sweep skips the issue and never retries. Related PRs: - Refs #8009 — adds dedup for `issue_blockers_resolved` re-fires (complementary; prevents over-firing; this PR ensures under-firing is caught) - Refs #6522 — `auto-unblock dependents with no assignee` (related no-assignee edge case) ## What Changed - **`server/src/services/heartbeat.ts`** — remove the inline dependency-wake retry loop from the `workspace_finalize` path; delegate to the shared `reconcileResolvedDependencyWakeups` helper instead - **`server/src/services/recovery/service.ts`** — split the resolved-blocker wake reconciliation sweep out from under the `liveness_escalation_auto_recovery` feature flag; the sweep runs unconditionally while escalation *creation* remains behind the flag - **`server/src/__tests__/issue-dependency-wakeups-routes.test.ts`** — regression: `blockedBy` set after blocker already done still triggers a reconciliation wake - **`server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts`** — regression: assignee-null churn before reconciliation; stale skipped dependency wakes do not suppress a fresh reconciliation wake ## Verification ``` pnpm exec vitest run \ server/src/__tests__/issue-dependency-wakeups-routes.test.ts \ server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts ``` 23 tests, all passing locally. `pnpm --filter @paperclipai/server typecheck` clean. ## Risks Low risk. Bounded blast radius: - The reconciliation sweep only considers non-hidden `blocked` issues with an agent assignee - Uses keyset pagination with an existing 500-candidate cap - Reuses dependency readiness/finalize gating logic unchanged - Skips issues with existing active/queued runs and pending interactions - Deduplicates against live/completed `issue_blockers_resolved` wakes (skipped/cancelled wakes intentionally do not suppress a fresh reconciliation) - Observability: healed reconciliations emit `issue.blockers_resolved_wake_emitted` activity and log the healed issue ids and source ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) with tool use and extended context. Paperclip AI agent runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
562567fcd6 |
[codex] Improve work timeline activity story (#9222)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The work timeline helps operators understand when agent and user activity actually happened across a project. > - The timeline view needs clearer interaction context so activity is easier to inspect and reason about. > - The existing story coverage did not fully exercise the denser activity states needed to review this UI safely. > - This pull request expands the work timeline data shape, service behavior, UI rendering, tests, and Storybook story so the activity timeline is easier to verify. > - The benefit is a more inspectable timeline for project activity, backed by targeted server and UI coverage. ## Linked Issues or Issue Description No public GitHub issue found, so this PR describes the feature inline following the feature request template. **Subsystem affected** Cross-cutting: `server/`, `packages/shared`, and `ui/`. **Problem or motivation** Project operators need a clearer timeline view that shows when work activity happened, how much agent time is represented inside the selected window, and enough realistic activity states for safe visual review. Sparse mock data and unbounded summary calculations make it harder to trust the timeline when inspecting historical or capped windows. **Proposed solution** Enrich the work timeline activity data returned by the service, render clearer top-level timeline summary stats, clamp duration calculations to the returned window, prorate token totals for partially visible spans, and add Storybook/test coverage with realistic timeline activity data. **Alternatives considered** Keeping the existing sparse timeline story was considered, but it would leave dense activity layouts and selected-window summary behavior under-reviewed. Counting full span usage for partially visible spans was also considered, but it makes historical windows report activity outside the displayed range. **Roadmap alignment** Searched `ROADMAP.md` for timeline/activity references and found no conflicting planned core work. **Additional context** This PR does not include migrations and does not commit generated design screenshots or images. ## What Changed - Extended shared work timeline activity types and server timeline service behavior. - Updated the timeline page and work timeline chart for richer activity rendering. - Clamped timeline runtime summary calculations to the returned window and prorated summary token usage for clipped spans. - Added and updated targeted server/UI tests for timeline activity behavior. - Added Storybook timeline mock coverage and Storybook preview setup needed by the story. ## Verification - `git rebase origin/master` completed cleanly after fetching `paperclipai/paperclip:master`. - `git diff --check origin/master...HEAD` - `pnpm exec vitest run server/src/__tests__/work-timeline-service.test.ts ui/src/components/timeline/WorkTimelineChart.test.tsx ui/src/pages/Timeline.test.tsx` — latest run: 3 files passed, 28 tests passed. - Greptile review completed at 5/5 with no unresolved Greptile threads after fixes. - GitHub checks completed green on the latest head SHA; Storybook visual regression was skipped by the workflow. - `pnpm check:token-gates` currently fails locally on existing `origin/master` violations in `ui/src/components/ActivityCharts.tsx` and `ui/src/components/IssueRecoveryActionCard.tsx`; this PR does not modify those files. ## Risks Low to moderate risk. The change affects the work timeline service response shape and timeline UI rendering, so regressions would likely show up as missing/incorrect timeline activity display. Targeted service and UI tests cover the changed behavior. No migrations are included. ## Model Used OpenAI Codex running GPT-5 as a tool-enabled coding agent with local shell and GitHub CLI access. Exact runtime model ID/context-window size was not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5d0de3499d |
[codex] Fix collapsed starred project indentation (#9215)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board sidebar is a high-frequency navigation surface for companies, projects, and work queues. > - Starred projects render as child rows under Projects in the expanded sidebar. > - The collapsed rail should align every nav icon in the same rail column. > - The starred-project child indent was still applied in the collapsed rail, which pushed the project glyph out of alignment. > - This pull request keeps the expanded hierarchy indent while removing it only for the collapsed rail. > - The benefit is a cleaner collapsed sidebar without changing expanded sidebar hierarchy. ## Linked Issues or Issue Description No public GitHub issue exists. ## What happened? In the collapsed sidebar rail, starred project rows kept the expanded child indentation. That pushed the project glyph out of alignment with the rest of the collapsed sidebar icons. ## Expected behavior Collapsed starred project icons should align with the other sidebar rail icons while the expanded sidebar should keep the child-row indentation under Projects. ## Steps to reproduce 1. Open Paperclip with at least one starred project. 2. Collapse the sidebar into rail mode. 3. Compare the starred project glyph position with the other collapsed sidebar glyphs. ## Paperclip version or commit Reproduced on the PR base before this branch; fixed on commit |
||
|
|
555391fed7 |
fix: run restart recovery, workspace self-heal, quota-aware retries, failed-run metrics (#9183)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run in heartbeat runs orchestrated by the server; run
lifecycle, retry scheduling, and the dashboard's run-activity metrics
are the subsystems involved
> - A spike in "failed" tasks traced to three causes: server restarts
killing in-flight runs and mislabeling them as failures, deterministic
workspace-validation loops when a worktree's branch diverged, and
provider quota/usage-limit errors being classified as generic transient
failures (putting agents into error state and polluting metrics)
> - Killed-then-recovered runs and quota waits are not product failures,
so both the runtime behavior and the reporting needed to distinguish
them
> - This pull request drains runs gracefully on shutdown with idempotent
restart retries, self-heals workspace branch mismatches, adds a
quota-aware failure class with reset-time retry, separates recovered
restart kills from true failures on the dashboard, and documents restart
hygiene for operators
> - The benefit is fewer spurious failures, automatic recovery instead
of manual repair, and dashboard metrics that reflect real failure rates
## Linked Issues or Issue Description
No public GitHub issue exists; describing the bug inline per the
bug-report template:
**What happened?**
In-flight heartbeat runs are marked `failed` when the server restarts,
even though a retry later succeeds. Worktrees whose checked-out branch
diverges from the issue branch fail workspace validation on every
subsequent run with no recovery path. Provider quota/usage-limit
responses are treated as generic transient upstream errors, putting
agents into an error state and retrying before the quota window resets.
The dashboard counts all of these as true failures, inflating failure
metrics.
**Expected behavior**
Graceful shutdown should interrupt (not fail) running runs and chain
exactly one recovery retry. Workspace validation should repair
recoverable branch mismatches automatically. Quota errors should get
their own error class with the retry scheduled at the provider reset
time and the agent left idle. The dashboard should report recovered
restart kills separately from true failures.
**Steps to reproduce**
1. Start a heartbeat run, then restart the server (SIGTERM) while it is
in flight — the run lands as `failed` with a process-loss error code
even when its retry succeeds
2. Check out an issue whose worktree branch has diverged (e.g. after a
force-moved branch) — every subsequent run fails
`workspace_validation_failed` deterministically
3. Drive an agent into a provider usage-limit window — the run fails as
a generic transient upstream error and the agent enters an error state
instead of idling until the reset time
**Paperclip version or commit**
master (base
|
||
|
|
3b16ac3804 |
fix(server): use run.id for activity_log in heartbeat invoke/resume (#3424)
## Summary - The heartbeat invoke and resume endpoints log activity with `actor.runId` (the caller's auto-generated run ID from JWT), which hasn't been registered in `heartbeat_runs` yet - This causes a FK constraint violation: `activity_log.run_id → heartbeat_runs.id` - Fix: use `run.id` (the newly created heartbeat_run) instead, which is guaranteed to exist in the table ## Test plan - [ ] Trigger a heartbeat invoke via the API — verify no FK constraint error in logs - [ ] Trigger a heartbeat resume — verify activity_log row is created successfully - [ ] Verify existing activity_log queries still return correct results 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
83f5f59842 |
[codex] Hide goals sidebar link behind experiment (#9189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board sidebar is the primary navigation surface for operators scanning companies, projects, tasks, agents, and related control-plane tools. > - Goals still has a route and product surface, but keeping the top-level sidebar link always visible makes it part of the default navigation whether or not that surface is ready for every operator. > - Instance experimental settings already provide a controlled place to expose optional UI surfaces while they are being evaluated. > - This pull request adds a dedicated experimental setting for restoring the Goals sidebar link. > - The benefit is a quieter default sidebar with an explicit escape hatch for operators who still need the Goals entry point. ## Linked Issues or Issue Description No public GitHub issue exists for this internal task, so the feature request is described inline. **Subsystem affected** Cross-cutting: `ui/`, `server/`, and `packages/shared`. **Problem or motivation** The Goals route remains available, but the top-level Goals sidebar entry makes that surface part of the default operator navigation. While the goals surface is still being evaluated, operators need a quieter default sidebar without losing an escape hatch for teams that still rely on the link. **Proposed solution** Add a boolean instance experimental setting, `enableGoalsSidebarLink`, default it to `false`, and render the Goals sidebar link only when the setting is enabled. Expose the toggle in Instance Experimental Settings so operators can restore the link without changing routes or rebuilding the app. **Alternatives considered** - Remove the Goals route entirely: rejected because this task only asks to hide the sidebar entry point and preserve access for teams evaluating goals. - Keep the sidebar link always visible: rejected because it does not provide the requested quieter default navigation. - Hard-code a local UI flag: rejected because instance experimental settings already provide the expected operator-controlled pattern. **Roadmap alignment** Checked `ROADMAP.md`; no overlapping goals/sidebar/experimental roadmap entry was found. **Additional context** The `/goals` route is preserved. This PR only gates the sidebar navigation item. ## What Changed - Added `enableGoalsSidebarLink` to the shared instance experimental settings type and validator, defaulting to `false`. - Normalized the new setting in the server instance settings service. - Hid the Goals sidebar nav item unless the new setting is enabled. - Added a Goals Sidebar Link toggle to the Instance Experimental Settings page. - Updated shared, server, sidebar, and settings page tests for the new setting. ## Verification - `pnpm exec vitest run packages/shared/src/validators/instance.test.ts server/src/__tests__/instance-settings-service.test.ts server/src/__tests__/instance-settings-routes.test.ts ui/src/components/Sidebar.test.tsx ui/src/pages/InstanceExperimentalSettings.test.tsx` - `git diff --check origin/master...HEAD` - `git merge-tree --write-tree HEAD origin/master` - Searched for duplicate/related PRs by title and `enableGoalsSidebarLink`; none found. - Checked `ROADMAP.md` for overlapping goals/sidebar/experimental entries; none found. ## Risks Low risk. The main behavior shift is that operators who depended on the sidebar Goals link need to enable the new experimental toggle. The `/goals` route itself is not removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based Paperclip CodexCoder session with repository tool access and command execution. Exact API model identifier and context window were not exposed by the Paperclip harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bb87047fe6 |
Add auto-forward execution workspace branch reconciliation (#9172)
## Thinking Path > - Paperclip manages execution workspaces for AI agent runs, each associated with a git branch so the agent always works in a known code state. > - Workspace runtime reconciles an agent's checkout branch against the workspace's recorded branch when a heartbeat resumes; without forward-ancestry detection, any divergence fails closed and blocks the run. > - When the workspace branch has moved forward (e.g. after a feature merge), the recorded branch is an ancestor of the current HEAD — a safe, forward-only case that the previous implementation refused even though it carries no safety risk. > - The gap means legitimate forward-advancing deployments require manual operator intervention to unblock agents every time, creating operational friction and interrupting automated workflows. > - This pull request adds a flag-gated reconcile-forward path that detects when the current branch is a strict forward descendant of the workspace branch and auto-reconciles, while preserving fail-closed behavior for all non-forward or flag-off cases. > - It also threads the active execution workspace id through restore and finalize call sites so the reconcile verdict can be persisted durably across heartbeats. > - The benefit is that agents resume automatically from forward-advancing workspace branches without operator intervention, while adversarial and backward branch changes continue to fail closed. ## Linked Issues or Issue Description This PR adds an auto-forward reconcile path for execution workspace branch tracking. When a workspace's recorded branch is a strict ancestor of the current HEAD (a forward-only advancement), the runtime now auto-reconciles rather than hard-blocking. The feature is gated behind an explicit runtime flag, defaults to off, and falls back to fail-closed behavior for all non-forward or flag-off cases. The prior implementation treated all branch divergences identically: any mismatch between the recorded workspace branch and the current HEAD failed closed. This prevented agents from resuming after routine forward deployments (e.g. after a feature branch merges into the workspace branch), requiring manual operator action to unblock every affected run. ## What Changed - Added `reconcileForward` flag-gated path in workspace runtime reconciliation logic that allows auto-reconciliation when the workspace branch is a strict ancestor of the current HEAD. - Threaded active execution workspace id through `restore` and `finalize` call sites so reconcile verdicts are persisted durably. - Added `plainLanguageReason` and `ancestryVerdict` evidence fields to the reconciliation result structure for operator visibility. - Stabilized a branch containment test that exposed a late run-linked activity FK cleanup race during the focused Vitest rerun. - All new paths remain fail-closed when the flag is off or when the branch relationship is not strictly forward. ## Verification ```bash pnpm --filter @paperclipai/shared typecheck pnpm --filter @paperclipai/server typecheck pnpm exec vitest run \ server/src/__tests__/workspace-runtime.test.ts \ server/src/__tests__/heartbeat-workspace-branch-containment.test.ts \ server/src/__tests__/execution-workspaces-service.test.ts git diff --check origin/master..HEAD ``` All 3 test files / 98 tests pass. Typecheck passes for both shared and server packages. > **Note:** This is a stacked PR on top of PR #9170 (Add execution workspace branch reconciliation route). The diff shown targets that branch; the combined change builds on the reconciliation route infrastructure it provides. ## Risks - **Flag-off default:** The reconcile-forward path is off by default. No behavior change for existing workspaces unless the flag is explicitly enabled by an operator. - **Ancestry check correctness:** The forward-only guard uses git ancestry verification; a branch that is not a strict ancestor of HEAD remains fail-closed. Adversarial or concurrent branch resets are not auto-reconciled. - **FK cleanup race (stabilized):** A late run-linked activity FK cleanup race in the containment test was exposed during the Vitest rerun. The stabilization commit addresses the non-deterministic ordering without changing production behavior. - **Stacking dependency:** This PR must not be merged before PR #9170 merges, as it is built on top of the reconciliation route infrastructure. ## Model Used - Provider: Anthropic - Model: Claude Sonnet 4.6 (`claude-sonnet-4-6`) - Context: 200k token context window - Mode: Agentic tool use with code execution and git operations ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f17202b571 |
Add execution workspace branch reconciliation route (#9170)
## Thinking Path > - Paperclip is an open-source app that lets teams run AI agents for work tasks; each agent session uses an execution workspace — a git checkout — to track the agent's active code state. > - Every execution workspace has an expected target branch (`PAPERCLIP_WORKSPACE_BRANCH`). The workspace git HEAD should always point to that branch so agents commit in the right place. > - When workspace git HEAD diverges from the expected branch — for example after a harness branch-name fix or an accidental `checkout -b` during a CI-retrigger — the discrepancy must be corrected before agents can continue safely. > - Operators (board users) need a controlled, audited path to reconcile a workspace's live branch back to the expected target, with an override escape-hatch for cases where the normal forward path is blocked. > - This pull request adds a board-only `POST /api/execution-workspaces/:id/reconcile-branch` service operation and route that validates safety preconditions, resolves matching recovery-action fingerprints, posts source-issue audit comments, and records the reconciliation outcome. > - The benefit is that operators can correct branch divergence through the API with a full audit trail, instead of via raw database edits. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Context below follows the feature-request template format. **Subsystem affected** server/ — REST API & orchestration services; packages/shared — request validation. **Problem or motivation** Execution workspaces have an expected branch record that must match the checked-out worktree branch. When the live git branch and stored branch record drift apart, operators currently lack a first-class, audited API to reconcile the record. The fallback is manual database repair or workspace replacement, both of which are risky and hard to audit. **Proposed solution** Add a board-only execution workspace branch reconciliation operation. `forward` mode re-inspects the server-side git state and only updates the branch record when the stored branch is an ancestor of the checked-out branch. `override` mode is a break-glass path that requires board access and an operator reason. Both modes require a clean, idle workspace, write audit details, post a source-issue audit comment, and resolve the matching workspace-validation recovery action. **Alternatives considered** Manual database edit (no durable audit trail and easy to mistype), recreating the workspace (heavier operational disruption), or trusting client-supplied ancestry evidence (unsafe because the server must verify the git state itself). **Roadmap alignment** This is incremental hardening for execution-workspace recovery and operator controls. It does not duplicate a public roadmap item. **Additional context** The endpoint is intended for operator recovery, not normal agent control flow, so the generated OpenAPI metadata and runtime route both classify it as board-only. ## What Changed - Added `reconcileExecutionWorkspaceBranchSchema` discriminated-union validator (`forward` with optional reason, `override` requiring a non-empty reason string) to `packages/shared/src/validators/execution-workspace.ts` - Exported `ReconcileExecutionWorkspaceBranch` type and the new schema from the shared package index - Added board-only reconcile-branch service operation in the execution workspaces service: safety checks, recovery-action fingerprint resolution, source-issue audit comment, and outcome recording - Added clean-worktree and stopped-runtime-service preconditions before branch-record mutation. - Marked the reconcile route as board-only in OpenAPI generated auth metadata. - Added `POST /api/execution-workspaces/:id/reconcile-branch` route wired to the new service operation with board-permission gate - Extended `execution-workspaces-routes.test.ts` and `execution-workspaces-service.test.ts` to cover: safety-check rejection, override-reason validation, audit-comment posting, and recovery-action fingerprint resolution (2 files / 19 tests) ## Verification ```sh pnpm --filter @paperclipai/shared typecheck pnpm --filter @paperclipai/server typecheck pnpm exec vitest run server/src/__tests__/execution-workspaces-routes.test.ts server/src/__tests__/execution-workspaces-service.test.ts pnpm exec vitest run server/src/__tests__/execution-workspaces-service.test.ts server/src/__tests__/openapi-routes.test.ts ``` ## Risks - **Board-only gate:** the operation is gated behind the board permission; no agent can trigger it without operator authorization. - **Override requires reason:** the `override` mode requires a non-empty reason string so every bypass is audited. - **Idempotent recovery-action resolution:** re-running with the same fingerprint is safe; duplicate resolution is a no-op. - **No execution-state mutation:** the route records a reconciliation intent and updates the branch record; it does not restart the workspace or modify running agent state. - Overall risk: **low**. ## Model Used - Provider: Anthropic - Model ID: `claude-sonnet-4-6` (Claude Sonnet 4.6) - Context window: 200 K tokens - Capabilities: tool use, code execution, multi-turn context Follow-up safety commit: - Provider: OpenAI - Model ID: `codex` / GPT-5 with tool use and code execution ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`feat/execution-workspace-branch-reconciliation-route`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bdd3aa2110 |
build(deps): bump sharp from 0.35.2 to 0.35.3 (#9061)
Bumps [sharp](https://github.com/lovell/sharp) from 0.35.2 to 0.35.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/lovell/sharp/releases">sharp's releases</a>.</em></p> <blockquote> <h2>v0.35.3</h2> <ul> <li> <p>Tighten verification of <code>text</code> dimensions, TIFF tile dimensions and <code>extend</code> values.</p> </li> <li> <p>Improve code bundler support by resolving path to libvips binary.</p> </li> <li> <p>Increase default concurrency when use of <code>MALLOC_ARENA_MAX</code> is detected.</p> </li> <li> <p>Emit warning about binaries provided by Electron for use on Linux.</p> </li> <li> <p>Add <code>hasAlpha</code> property to output <code>info</code>. <a href="https://redirect.github.com/lovell/sharp/issues/4500">#4500</a></p> </li> <li> <p>TypeScript: Return more precise <code>Buffer<ArrayBuffer></code> from <code>toBuffer</code>. <a href="https://redirect.github.com/lovell/sharp/pull/4520">#4520</a> <a href="https://github.com/Andarist"><code>@Andarist</code></a></p> </li> <li> <p>Bound <code>clahe</code> width and height to avoid signed overflow. <a href="https://redirect.github.com/lovell/sharp/pull/4551">#4551</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Bound <code>trim</code> margin to avoid signed overflow. <a href="https://redirect.github.com/lovell/sharp/pull/4552">#4552</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Reject infinite values when validating numbers. <a href="https://redirect.github.com/lovell/sharp/pull/4553">#4553</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Bound extract region to libvips coordinate limit. <a href="https://redirect.github.com/lovell/sharp/pull/4555">#4555</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Verify background colour values are numbers. <a href="https://redirect.github.com/lovell/sharp/pull/4556">#4556</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Bound create and raw input dimensions to coordinate limit. <a href="https://redirect.github.com/lovell/sharp/pull/4558">#4558</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> <li> <p>Tighten recomb and affine matrix verification. <a href="https://redirect.github.com/lovell/sharp/pull/4560">#4560</a> <a href="https://github.com/chatman-media"><code>@chatman-media</code></a></p> </li> <li> <p>Verify cache memory limit to avoid overflow. <a href="https://redirect.github.com/lovell/sharp/pull/4561">#4561</a> <a href="https://github.com/metsw24-max"><code>@metsw24-max</code></a></p> </li> </ul> <h2>v0.35.3-rc.2</h2> <ul> <li>Tighten verification of <code>text</code> dimensions, TIFF tile dimensions and <code>extend</code> values.</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/lovell/sharp/commit/1018449164723ba0203c1beffaba0e21f7829c18"><code>1018449</code></a> Release v0.35.3</li> <li><a href="https://github.com/lovell/sharp/commit/ba303a799de6b0639a108e82465870ce0722ee4a"><code>ba303a7</code></a> Prerelease v0.35.3-rc.2</li> <li><a href="https://github.com/lovell/sharp/commit/4f94fc5162a488187be17872ba513d7cadfe22d6"><code>4f94fc5</code></a> Upgrade to sharp-libvips v1.3.2</li> <li><a href="https://github.com/lovell/sharp/commit/c5e7a3ff2043922b110dc9c00700c6f17d478c33"><code>c5e7a3f</code></a> Bump devDeps, fix Deno/Windows smoke tests</li> <li><a href="https://github.com/lovell/sharp/commit/9a8d00268893cf163eea638fcc05aaed65fe01ea"><code>9a8d002</code></a> Docs: Add changelog entry and note about transferable <a href="https://redirect.github.com/lovell/sharp/issues/4520">#4520</a></li> <li><a href="https://github.com/lovell/sharp/commit/8694db0bac0d227f183d05585ebac7d17048d123"><code>8694db0</code></a> TypeScript: Return more precise <code>Buffer\<ArrayBuffer></code> from <code>toBuffer</code> (<a href="https://redirect.github.com/lovell/sharp/issues/4520">#4520</a>)</li> <li><a href="https://github.com/lovell/sharp/commit/e000d0b5e128e05bb6c499600f37e2a6a4b74314"><code>e000d0b</code></a> Prerelease v0.35.3-rc.1</li> <li><a href="https://github.com/lovell/sharp/commit/9554ca95535e8385ec36815d7c793ce96739bf1c"><code>9554ca9</code></a> Prerelease v0.35.3-rc.0</li> <li><a href="https://github.com/lovell/sharp/commit/6a29fd55db0c569276f143a7b480c62573a7aa16"><code>6a29fd5</code></a> Emit warning about native binaries on Linux Electron</li> <li><a href="https://github.com/lovell/sharp/commit/540d2eada4613f954aa541ffac8c12485375967e"><code>540d2ea</code></a> Increase default concurrency when use of MALLOC_ARENA_MAX detected</li> <li>Additional commits viewable in <a href="https://github.com/lovell/sharp/compare/v0.35.2...v0.35.3">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
077ba611fa |
build(deps-dev): bump @types/multer from 2.1.0 to 2.2.0 (#9065)
Bumps [@types/multer](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/multer) from 2.1.0 to 2.2.0. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/multer">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
390627b46e |
[codex] Suppress worktree heartbeat scheduling (#9163)
Suppress heartbeat scheduling in worktree and restore runtimes while keeping routine ticks and setup cleanup active. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
57a7da81ee |
[codex] Isolate run JWTs by control-plane instance (#9162)
Bind local agent run JWT signing and validation to the issuing Paperclip instance while preserving rollout compatibility for legacy tokens. Co-Authored-By: Paperclip <noreply@paperclip.ing> |