mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 01:24:44 +02:00
efcce9cc8e17e2b30f6f2edf30e00a77b0bf19bb
344
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
efcce9cc8e |
fix(adapters): record unpriced CLI usage (#9505)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Budgets and spend telemetry are control-plane safety features, not just reporting > - Local Codex and Claude adapters can execute through either ACP or their native CLI engines > - The ACP lane records usage and reported cost, but CLI JSON output often reports tokens without a price > - The CLI lane was either losing per-run usage semantics or coercing missing cost to zero, making real usage indistinguishable from a genuinely free run > - This pull request preserves CLI usage as per-run totals and records token-bearing runs without a reported price as explicitly unpriced ledger events > - The benefit is accurate usage accounting and a visible pricing gap instead of silently misleading zero-cost telemetry ## Linked Issues or Issue Description Refs #9471 Refs #9230 **Bug description** A `codex_local` run using the CLI engine can emit a final `turn.completed` event with millions of input tokens and tens of thousands of output tokens while the agent's spend ledger remains indistinguishable from a true zero-usage, zero-cost run. Claude CLI output has the same missing-price edge case. **Expected behavior** Token-bearing CLI runs should persist their usage. If the adapter reports a price, the ledger should record it as reported; if the CLI reports usage but no price, the ledger should explicitly mark the event as unpriced rather than silently treating missing price data as a reported `$0` cost. **Reproduction shape** 1. Configure `codex_local` with `engine: cli`. 2. Run a task that produces a `turn.completed` usage payload. 3. Observe token usage in the run stream. 4. Before this change, missing price data is represented as ordinary zero-cost spend and the CLI usage basis is not consistently propagated. ## What Changed - Mark Codex and Claude native CLI usage totals as `per_run` and propagate that basis through success and failure results. - Stop coercing missing Claude CLI cost to `0`. - Add `cost_status` to cost events with `reported` and `unpriced` values, including an idempotent migration and shared validation/types. - Persist token-bearing runs without a reported price as `unpriced` ledger events while retaining zero cents until an authoritative price exists. - Add parser, execute-path, heartbeat-accounting, and cost-service regression coverage for both local CLI adapters. - Document the cost-status invariant and CLI accounting behavior. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/parse.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests passed. - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` — includes migration numbering and safety checks. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Existing cost rows default to `reported`, preserving current interpretation; only new token-bearing events with absent cost are marked `unpriced`. - This change does not invent model pricing. Budget hard stops still cannot charge an unknown amount, but operators and evals can now distinguish missing pricing from a genuinely reported zero cost. - Consumers that enumerate cost-event fields should tolerate the additive `costStatus` field. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use and code execution; default reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e21a27301 |
fix: forward onSpawn to hermes and process adapters for PID persistence (#8722)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter layer (hermes-local, process adapters) delegates agent execution to child processes via `runChildProcess()` > - `runChildProcess()` accepts an `onSpawn` callback to report child PID and process group info, but the hermes and process adapters were not forwarding `ctx.onSpawn` to this call > - Without PID persistence, the orphan reaper cannot distinguish live runs from abandoned processes, causing false-positive reaps and 5-minute timeout errors for active runs > - This pull request adds `onSpawn: ctx.onSpawn` to both adapter call sites and declares the option in the `runChildProcess` wrapper type > - The benefit is that the orphan reaper can now correctly track live child processes, eliminating false-positive reaps ## Linked Issues or Issue Description Fixes #8723 Fixes false-positive orphan reaps in hermes-local and process adapters by forwarding the `onSpawn` callback to `runChildProcess()`. All other adapters (claude-local, codex-local, cursor-local, gemini-local, grok-local, opencode-local, pi-local) already forward `ctx.onSpawn` — these two were the only ones missing it. ## What Changed - `server/src/adapters/utils.ts`: Added `onSpawn?` to the `runChildProcess()` options type so callers can forward the callback - `server/src/adapters/process/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` - `packages/adapters/hermes/src/server/execute.ts`: Forward `ctx.onSpawn` to `runChildProcess()` ## Verification - `pnpm -r typecheck` passes across all packages - Confirmed all other adapters already forward `ctx.onSpawn` (12 grep matches across 9 adapter files) - The 3-line diff is additive only — no existing behavior is changed, only a previously-ignored callback is now forwarded ## Risks Low risk. This is a 3-line additive change. The `onSpawn` parameter is optional (`?`) so existing callers are unaffected. The callback is already well-established across all other adapters. ## Model Used Hermes Agent (by Nous Research) — xiaomi/mimo-v2.5-pro via OpenRouter, with tool use (file editing, git, GitHub API). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass (typecheck passes) - [x] I have added or updated tests where applicable (N/A — type-level fix only, no behavioral change) - [x] I have updated relevant documentation to reflect my changes (N/A — internal fix) - [x] I have considered and documented any risks above --------- Co-authored-by: Zephyr <zephyr@motoyuki.dev> |
||
|
|
c8253e3641 |
fix(adapters): inject execution contract once per fresh heartbeat (#9469)
## Thinking Path > - Paperclip coordinates AI-agent work through repeated heartbeat runs. > - Adapter prompts combine a default heartbeat template with scoped wake context. > - Fresh heartbeats received the same execution contract from both layers, wasting prompt tokens and obscuring which layer owns the contract. > - Resume deltas and template-less adapters do not share that composition path, so removing the wake-payload copy unconditionally would drop required guidance. > - Empty comment batches also emitted instructions and metadata that only matter when comments exist. > - This pull request makes execution-contract inclusion explicit by prompt path, preserves OpenClaw gateway behavior, and suppresses no-op comment boilerplate. > - The benefit is one contract per heartbeat path and roughly 300 fewer prompt tokens on a fresh zero-comment wake. ## Linked Issues or Issue Description - Fixes #9221 - Refs #9200 - Refs #7634 ## What Changed - Stop emitting the execution-contract paragraph from fresh scoped wake payloads because the default heartbeat template already contains the full contract. - Keep the contract in resume deltas, and add `includeExecutionContract` for adapters that do not render the default heartbeat template. - Opt `openclaw-gateway` into wake-payload contract rendering so template-less gateway runs retain the guidance. - Omit comment-batch acknowledgement/fetch guidance and empty `pending comments` / `latest comment id` metadata when a fresh wake has no pending comments. - Add regression and acceptance coverage proving composed fresh prompts contain `Execution contract` exactly once while resume and template-less paths retain it. Measured effect: the fresh zero-comment wake block drops from 1,840 to 855 characters (about 300 tokens saved per fresh heartbeat; about 220 on comment wakes), and the composed fresh prompt contains `Execution contract` once instead of twice. ## Verification - `npx vitest run packages/adapter-utils/src/server-utils.test.ts` — 63 passed - `npx vitest run server/src/__tests__/codex-local-execute.test.ts` — 13 passed - `npx vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/openclaw-gateway-adapter.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts` — 27 passed - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` — passed ## Risks - Low risk: prompt text and adapter composition only; no database or API migration. - The main compatibility risk is a template-less adapter losing the contract. The explicit option and OpenClaw gateway regression coverage protect the known template-less path. - External adapters that call `renderPaperclipWakePrompt` directly can opt into `includeExecutionContract: true` when they do not render the default template. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, reasoning mode with tool use and code execution; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9e7e84e3fe |
fix(adapters): propagate ACP-lane usage and cost into spend telemetry (#9471)
## Thinking Path > - Paperclip is the open source control plane for running and governing AI-agent companies. > - Adapter executions feed token usage, billing identity, and run cost into the control plane's spend telemetry. > - The default ACP execution lane for local Claude and Codex adapters did not propagate per-turn usage or cost, so paid runs could be recorded with zero spend and no tokens. > - Claude CLI result events could also undercount output tokens by reading only the main-loop usage block instead of the complete per-model ledger. > - The shared executor needs to distinguish per-run usage from session-cumulative usage so the server does not apply the wrong delta heuristic. > - This pull request captures ACP usage and cumulative-cost deltas, resolves adapter billing identity, uses Claude's complete model-usage ledger, and preserves per-run usage in server normalization. > - The benefit is accurate token and cost accounting across the default paid Claude and Codex execution paths. ## Linked Issues or Issue Description ### What happened? Paid `claude_local` and `codex_local` runs using the default ACP engine can complete successfully while the control plane records zero or null cost and missing token usage. Claude CLI result parsing can additionally undercount output tokens when subagent or sidechain usage is present. ### Steps to reproduce 1. Run a paid Claude or Codex local adapter through the ACP engine. 2. Complete a turn that reports usage and cumulative cost through ACP status/events. 3. Inspect the execution result and normalized run telemetry. ### Expected behavior The execution result contains per-turn token usage, a per-run USD cost delta, and the correct billing identity. Server normalization records those per-run values without applying a session-cumulative delta a second time. ### Actual behavior before this change ACP execution results returned no usage and `costUsd: null` with unknown billing. The server therefore recorded zero spend and no tokens for paid runs. Claude CLI parsing could use an incomplete usage block. ## What Changed - Capture ACP usage from runtime status and `usage_update` events, reporting it as `usageBasis: per_run`. - Convert agent-reported cumulative ACP cost into a per-turn delta, including counter-reset and no-report safeguards. - Add a shared billing-identity resolver and map Claude and Codex authentication/provider modes to control-plane billing types. - Prefer Claude result-event `modelUsage` totals so subagent and sidechain tokens are included. - Skip the server's session-cumulative usage delta when an adapter explicitly reports per-run usage. - Add regression coverage for usage capture, event fallback, cost resets, stale reports, billing identities, model-usage totals, and server spend normalization. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/acp.test.ts server/src/__tests__/costs-service.test.ts server/src/__tests__/monthly-spend-service.test.ts` — 6 files, 126 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `pnpm --filter @paperclipai/adapter-claude-local typecheck` — passed. - `pnpm --filter @paperclipai/adapter-codex-local typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - A broader Claude-local suite has a pre-existing rate-limit classification failure in `test.probe.test.ts`; it also fails on clean `master` and is unrelated to this change. ## Risks - Cost reporting depends on the agent's cumulative counter semantics; reset handling falls back to the post-turn amount and is covered by regression tests. - Incorrect billing-mode inference could misclassify spend; provider/auth mappings mirror each adapter's existing CLI behavior and have focused tests. - The new `usageBasis` contract changes server normalization only when adapters explicitly opt into `per_run`; existing adapters retain prior behavior. - No database migration, workflow, lockfile, or UI changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation commit: Anthropic Claude Fable 5, tool-enabled coding workflow (exact context window and runtime configuration were not recorded in the commit metadata). - PR preparation and verification: OpenAI Codex, tool-enabled coding agent (runtime model ID and context window are not exposed to this session). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local> |
||
|
|
1fe89eb8f8 |
Enforce durable external-wait liveness (#9373)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat/recovery subsystem decides whether an agent run has a durable continuation path after the process stops. > - External waits need stricter semantics than local background watchers: a killed local process is not durable, while a first-class blocker/monitor/scheduled wake is. > - Without that distinction, recovery can repeatedly treat adapter-failed continuations as live work and obscure the real reason a task stopped. > - This pull request adds explicit durable external-wait liveness handling and documents the expected execution semantics. > - It also improves operator-visible recovery evidence so invalid external-wait paths explain why they were rejected. > - The benefit is clearer recovery behavior, fewer duplicate continuation recoveries, and a safer contract for monitor-backed external waits. ## Linked Issues or Issue Description - Refs #5978 - Related PRs: #4988, #7495, #8502 ## What Changed - Added durable external-wait liveness classification so local/background watchers are not accepted as durable live paths after the owning process exits. - Preserved first-class blocker/monitor/scheduled wake paths as valid external-wait continuations. - Added backend regression coverage for killed watcher failure, monitor-backed durable wait resumption, normal completion, blocker behavior, and no duplicate recovery. - Added adapter utility coverage for terminal cleanup behavior used by local process adapters. - Surfaced invalid external-wait recovery evidence in the recovery action card and run ledger. - Updated execution semantics documentation and the V1 implementation contract. ## Verification - `pnpm check:token-gates` passed. - `pnpm -r typecheck` passed. - `node scripts/run-vitest-stable.mjs --mode general --group general-server` equivalent lane passed in CI-clean env: 238 files, 2164 tests passed, 1 skipped. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-a` passed in fully Paperclip-env-clean env: UI 305 files / 2430 tests; CLI 43 files / 230 tests. - `node scripts/run-vitest-stable.mjs --mode general --group general-workspaces-b` passed in fully Paperclip-env-clean env: shared/db/adapters/plugin packages all green. - `node scripts/run-vitest-stable.mjs --mode serialized` passed in fully Paperclip-env-clean env: 107 serialized server suites green, including 84/84 heartbeat-process-recovery tests. - `pnpm build` passed in fully Paperclip-env-clean env. Notes: running `pnpm test:run` directly inside the Paperclip heartbeat environment exposed local harness env contamination in existing tests (`PAPERCLIP_CONFIG`, `PAPERCLIP_DB_BACKUP_DIR`, and `PAPERCLIP_WORKTREE_START_POINT`). Re-running the same lanes with inherited `PAPERCLIP_*` and port env removed produced the CI-equivalent green results above. ## Risks - Medium behavioral risk: this changes recovery classification for stopped local external-wait processes, so adapters relying on unmanaged background watchers must use blockers, monitors, scheduled wakes, or explicit durable handoff instead. - Low UI risk: recovery-card copy changes are covered by component tests and Storybook screenshot QA. - No database migration is included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5-based coding agent, tool-enabled terminal/code execution. Exact context-window metadata was not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a02fe8d575 |
Update Codex adapter GPT-5.6 defaults (#9352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Codex local is the adapter subsystem that exposes OpenAI Codex CLI model choices to agents and issue overrides. > - OpenAI has GPT-5.6 Codex-capable models that should appear in Paperclip's built-in Codex model list and refresh behavior. > - Paperclip's server model listing falls back to the adapter metadata and merges OpenAI refresh results with known Codex defaults. > - This pull request updates the Codex default model metadata to include GPT-5.6 options and adds regression coverage for fallback and refresh paths. > - The benefit is that operators can select the new Codex models without relying on manual model IDs, and refresh behavior keeps known GPT-5.6 options visible. ## Linked Issues or Issue Description Refs #9322. Refs #9342. Refs #9346. ### Agent or provider Codex CLI (OpenAI). ### Why this adapter is useful OpenAI's GPT-5.6 Codex-capable models should be available in Paperclip's Codex adapter defaults and model refresh path. ### How the agent is invoked `codex` ## What Changed - Changed the `codex_local` default model metadata from `gpt-5.5` to `gpt-5.6`. - Added `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` to the built-in Codex adapter model list. - Updated adapter and server model-listing tests to cover GPT-5.6 fallback and refresh behavior. - Aligned Codex Fast mode support and helper text with the new `gpt-5.6` default, while preserving GPT-5.5, GPT-5.4, and manual model ID support. ## Verification - `git diff --check origin/master...HEAD` - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts server/src/__tests__/adapter-models.test.ts server/src/__tests__/adapter-model-refresh-routes.test.ts` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` ## Risks Medium risk because changing `DEFAULT_CODEX_LOCAL_MODEL` from `gpt-5.5` to `gpt-5.6` changes the adapter's default model selection for new blank configurations. The model-list additions are otherwise low risk and covered by adapter/server metadata tests. This PR intentionally overlaps related PRs #9342 and #9346, so reviewers may prefer to close or fold it into one of those branches. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, GitHub CLI, and repository editing tool use. Exact served model ID and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d3e26a8d02 |
docs: point Paperclip docs links at docs.paperclip.ing (#9300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The public README files are owned discovery surfaces for users who arrive from GitHub or npm. > - Some documentation links still used the old `https://paperclip.ing/docs` redirect path. > - Redirect hops are worse for users and for SEO because crawlers and readers do not land on the canonical docs host immediately. > - Package metadata should still point at the GitHub repository, because npm package homepages are expected to identify the source/project page. > - This pull request updates only explicit documentation links to the canonical docs subdomain. > - The benefit is a smaller, clearer link sweep with no package homepage metadata change. ## Linked Issues or Issue Description - No public GitHub issue exists for this small documentation maintenance change. - Problem: public README documentation links used a redirecting docs URL instead of the canonical docs host. - Expected behavior: README documentation links should point directly at `https://docs.paperclip.ing`. - Scope: root README, CLI README, and the Hermes adapter README docs reference. - Related search results reviewed: #793, #592, #675, and this PR. No open duplicate PR was found for this README-only canonical docs URL sweep. ## What Changed - Updated the root README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the CLI README Docs navigation link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Updated the Hermes adapter README Paperclip Docs link from `https://paperclip.ing/docs` to `https://docs.paperclip.ing`. - Kept all `package.json` homepage fields pointing at the Paperclip GitHub repository or package-specific GitHub README pages. ## Verification - `git diff --check origin/master...HEAD` - `rg 'https://paperclip\.ing/docs|https://docs\.paperclip\.ing' README.md cli/README.md packages/adapters/hermes/README.md` - `rg '"homepage": "https://docs\.paperclip\.ing"' -g 'package.json'` returned no matches. - Reviewed the final diff against `origin/master`; only the three README files changed. ## Risks - Low risk: this is a documentation-only URL update. - The main review risk is scope creep into npm metadata; that was explicitly avoided by keeping package homepages on GitHub. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5-based Codex coding agent with local shell, git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3571b6c38b |
feat(models): add gpt-5.4-mini to Codex and OpenCode selection (and openai/gpt-5.5 to OpenCode) (#4357)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtimes are selected through adapters, and each adapter exposes the model IDs an operator can pick in the UI > - The `codex_local` and `opencode_local` adapters hardcode those lists, so a newly released model stays unreachable until it is added > - `gpt-5.4-mini` is available via both the Codex CLI and the OpenCode CLI, but neither adapter lists it; `openai/gpt-5.5` is likewise missing from `opencode_local` > - This pull request adds those entries, and deliberately keeps `gpt-5.4-mini` out of the Codex Fast mode allowlist because the model does not support Fast mode > - The benefit is that operators can select these models from the UI instead of falling back to a manual model ID, and Fast mode fails closed rather than sending unsupported overrides to the CLI ## Linked Issues or Issue Description No existing issue covers this. Describing it inline, following the feature request template: ### Problem or motivation `gpt-5.4-mini` is absent from the `models` list of both the `codex_local` and `opencode_local` adapters, and `openai/gpt-5.5` is absent from `opencode_local`. (`codex_local` already ships `gpt-5.5` — it is the default model on master.) Operators who want these models must type a manual model ID. For `codex_local` that has a real side effect. `isCodexLocalFastModeSupported` treats any *unknown* model as Fast-mode-capable and passes `service_tier="fast"` and `features.fast_mode=true` through to the CLI. Because `gpt-5.4-mini` does not support Codex Fast mode, an agent configured with a manual `gpt-5.4-mini` model ID and `fastMode` enabled silently sends overrides the CLI cannot honor. ### Proposed solution Add the three missing entries to the two `models` lists, and leave `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` untouched. Listing `gpt-5.4-mini` in `models` is precisely what makes it a *known* model, so `isCodexLocalFastModeSupported` returns `false`, `buildCodexExecArgs` omits the Fast mode overrides, and `fastModeIgnoredReason` is surfaced to the operator. ### Alternatives considered Adding `gpt-5.4-mini` to `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` as well — rejected, because the model does not support Fast mode and the overrides would be rejected at run time. Leaving the models unlisted so operators keep using manual IDs — rejected, because that is the path that silently enables Fast mode for a model that cannot use it. ### Roadmap alignment Not core roadmap work. `ROADMAP.md` does not plan adapter model-list maintenance; this is routine upkeep as upstream CLIs ship new models. ## What Changed - Add `gpt-5.4-mini` to the `codex_local` adapter's `models` list, positioned after `gpt-5.4` (newest-first ordering). - Add `openai/gpt-5.5` and `openai/gpt-5.4-mini` to the `opencode_local` adapter's `models` list. - Add a `buildCodexExecArgs` test asserting Fast mode is ignored for `gpt-5.4-mini`. `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS` is intentionally unchanged. No behavior changes to existing models or adapter logic. ## Verification ``` pnpm --filter @paperclipai/adapter-codex-local --filter @paperclipai/adapter-opencode-local typecheck npx vitest run packages/adapters/codex-local packages/adapters/opencode-local ``` Both pass: typecheck clean on both packages, and 22 test files / 133 tests green, including the new `ignores fast mode for gpt-5.4-mini` case. ## Risks Low risk. The change is additive: three entries appended to two model-selection lists, plus one test. No default model changes, no adapter logic changes, no migrations. One behavioral shift is intended. An operator who had `gpt-5.4-mini` configured as a *manual* model ID with `fastMode` enabled was getting Fast mode overrides passed through to the Codex CLI. After this change `gpt-5.4-mini` is a known model, so those overrides are dropped and `fastModeIgnoredReason` explains why. ## Model Used - OpenAI Codex CLI with GPT-5 / GPT-5.5-assisted code editing (the original commits on this branch). - Anthropic Claude Opus 4.8 (`claude-opus-4-8`, 1M context, extended thinking, tool use) for the master merge, conflict resolution, and the scope reduction in the latest commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched the GitHub PR list for similar or duplicate PRs and confirmed this one is not a duplicate - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (N/A; the adapter docs describe Fast mode support, which is unchanged) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc17f29e7e |
fix(timeouts): raise sandbox wall-clock backstop to 4h and make acpx_local timeouts self-describing (#9232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs execute through adapters (e.g. `acpx_local`), which can run locally, over SSH, or inside sandbox execution targets, each with a wall-clock execution timeout > - Sandbox-backed runs defaulted to a 30-minute wall-clock backstop (`DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC = 1800`), which kills healthy long agent runs that are still making progress — long before the recovery watchdog's 4h critical threshold would even consider them stuck > - On top of that, `acpx_local` resolved its timeout directly from `adapterConfig.timeoutSec` instead of the shared execution-target resolver, and its timeout failures surfaced as a bare `Timed out after Ns` — giving operators no clue which timer fired or which knob raises it > - This pull request raises the sandbox backstop to 4h (aligned with the recovery watchdog), routes `acpx_local` through the shared timeout resolver, logs the effective timeout and its source at run start, and makes every timeout error message self-describing > - The benefit is that long-running sandbox agent runs no longer die at 30 minutes, and when a wall-clock timeout does fire, the run log states exactly which timer fired and how to configure it ## Linked Issues or Issue Description Refs #4535 (related: wall-clock execution timeouts killing agent runs that are still making progress — that issue covers a different hardcoded 600s timer, but the operator pain is the same). No exact public issue exists for this one, so describing it in-PR: **Bug:** A long sandbox-backed `acpx_local` agent run was killed with a bare `Timed out after 1800s` even though the agent was actively working. - **What happened:** The run hit the 30-minute `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` backstop. `acpx_local` never consulted the shared execution-target timeout resolution (it read `adapterConfig.timeoutSec` directly, default 0), so on sandbox targets the sandbox-provider default applied with no adapter-level say. The resulting error named neither the timer that fired nor the knob that controls it. - **Expected:** Healthy long runs should not be killed by a 30-minute wall-clock backstop when the recovery watchdog only treats runs as critically stuck after 4h of output silence; and any timeout error should say which timeout fired and how to raise it. - **Impact:** Long, legitimate agent runs in sandboxes fail mid-work; operators waste time reverse-engineering which of several timers produced "Timed out after Ns". ## What Changed - `packages/adapter-utils/src/execution-target.ts` - `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` raised from `1_800` to `14_400` (4h), with a comment explaining it intentionally matches the recovery watchdog's `ACTIVE_RUN_OUTPUT_CRITICAL_THRESHOLD_MS` (4h) so the adapter backstop never fires before the watchdog path. Output-inactivity monitors remain the primary hang detectors. - New `resolveAdapterExecutionTargetTimeout(target, configuredTimeoutSec)` returns `{ timeoutSec, source }` where `source` is `configured` / `sandbox_default` / `unlimited`. The existing `resolveAdapterExecutionTargetTimeoutSec` is preserved as a thin wrapper, so current callers are unaffected. - New `formatAdapterExecutionTimeoutErrorMessage(resolution)` and `formatAdapterExecutionTimeoutStartLogLine(resolution)` produce self-describing messages that name the timer that fired and the `adapterConfig.timeoutSec` knob that controls it. - `packages/adapters/acpx-local/src/server/execute.ts` - `buildRuntime` now resolves the wall-clock timeout through the shared resolver: sandbox targets default to the 4h backstop, local/SSH keep the historical "0 = no adapter timeout", and a configured `adapterConfig.timeoutSec` always wins. - The executor logs the effective timeout and its source at run start (`[paperclip] Adapter execution timeout: …`), so a later timeout is diagnosable from the run log alone. - All three bare timeout messages (timer cancel reason, turn result `errorMessage`, catch-path `messageOverride`) now use the self-describing format. - `packages/adapters/acpx-local/src/index.ts` — the adapter configuration doc for `timeoutSec` states the sandbox default and that the output-inactivity monitor remains the primary hang detector. - Tests: `packages/adapter-utils/src/execution-target-sandbox.test.ts` and `packages/adapters/acpx-local/src/server/execute.test.ts` (see Verification). ## Verification - `pnpm --filter @paperclipai/adapter-utils typecheck` — passes - `pnpm --filter @paperclipai/adapter-acpx-local typecheck` — passes - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapters/acpx-local/src/server/execute.test.ts` — 2 files, 40 tests, all pass - New/updated test coverage: - sandbox default resolves to 4h (and the constant is asserted to be `4 * 60 * 60`) - `resolveAdapterExecutionTargetTimeout` reports `configured` / `sandbox_default` / `unlimited` sources with the correct precedence (configured > sandbox default; local/SSH stay unlimited) - exact wording of the self-describing error message and the start-of-run log line - `acpx_local` runtime picks up the sandbox default into `timeoutMs`, keeps the unlimited local default, honors configured-over-default precedence, emits the start-of-run log line, and surfaces the self-describing `errorMessage`/cancel reason when the wall-clock timer kills a turn ## Risks - **Behavioral shift:** sandbox-backed adapter runs that previously hit the 30-minute backstop now run up to 4h before the adapter kills them. Genuinely hung runs are still caught much earlier by the adapters' output-inactivity monitors and by the recovery watchdog; the wall-clock timer is a last-resort kill switch. Operators who relied on the 30-minute default can restore it explicitly via `adapterConfig.timeoutSec`. - **Error-message consumers:** any tooling that pattern-matched the exact `Timed out after Ns` string from `acpx_local` will see the new self-describing message instead. - No API or schema changes; `resolveAdapterExecutionTargetTimeoutSec` keeps its exact signature and behavior (modulo the raised sandbox default). > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) via Claude Code CLI — model ID `claude-fable-5`, extended thinking enabled, agentic tool use (file edits, shell, test execution) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
555391fed7 |
fix: run restart recovery, workspace self-heal, quota-aware retries, failed-run metrics (#9183)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run in heartbeat runs orchestrated by the server; run
lifecycle, retry scheduling, and the dashboard's run-activity metrics
are the subsystems involved
> - A spike in "failed" tasks traced to three causes: server restarts
killing in-flight runs and mislabeling them as failures, deterministic
workspace-validation loops when a worktree's branch diverged, and
provider quota/usage-limit errors being classified as generic transient
failures (putting agents into error state and polluting metrics)
> - Killed-then-recovered runs and quota waits are not product failures,
so both the runtime behavior and the reporting needed to distinguish
them
> - This pull request drains runs gracefully on shutdown with idempotent
restart retries, self-heals workspace branch mismatches, adds a
quota-aware failure class with reset-time retry, separates recovered
restart kills from true failures on the dashboard, and documents restart
hygiene for operators
> - The benefit is fewer spurious failures, automatic recovery instead
of manual repair, and dashboard metrics that reflect real failure rates
## Linked Issues or Issue Description
No public GitHub issue exists; describing the bug inline per the
bug-report template:
**What happened?**
In-flight heartbeat runs are marked `failed` when the server restarts,
even though a retry later succeeds. Worktrees whose checked-out branch
diverges from the issue branch fail workspace validation on every
subsequent run with no recovery path. Provider quota/usage-limit
responses are treated as generic transient upstream errors, putting
agents into an error state and retrying before the quota window resets.
The dashboard counts all of these as true failures, inflating failure
metrics.
**Expected behavior**
Graceful shutdown should interrupt (not fail) running runs and chain
exactly one recovery retry. Workspace validation should repair
recoverable branch mismatches automatically. Quota errors should get
their own error class with the retry scheduled at the provider reset
time and the agent left idle. The dashboard should report recovered
restart kills separately from true failures.
**Steps to reproduce**
1. Start a heartbeat run, then restart the server (SIGTERM) while it is
in flight — the run lands as `failed` with a process-loss error code
even when its retry succeeds
2. Check out an issue whose worktree branch has diverged (e.g. after a
force-moved branch) — every subsequent run fails
`workspace_validation_failed` deterministically
3. Drive an agent into a provider usage-limit window — the run fails as
a generic transient upstream error and the agent enters an error state
instead of idling until the reset time
**Paperclip version or commit**
master (base
|
||
|
|
aac7e2ed15 |
build(deps): bump acpx from 0.11.2 to 0.12.0 (#9062)
Bumps [acpx](https://github.com/openclaw/acpx) from 0.11.2 to 0.12.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/openclaw/acpx/releases">acpx's releases</a>.</em></p> <blockquote> <h2>acpx 0.12.0</h2> <h2>v0.12.0</h2> <h3>Changes</h3> <ul> <li>Agents/built-ins: add Grok Build via <code>grok agent stdio</code>, including cached-login and <code>XAI_API_KEY</code> authentication selection. Thanks <a href="https://github.com/TheAngryPit"><code>@TheAngryPit</code></a>. (<a href="https://redirect.github.com/openclaw/acpx/pull/426">#426</a>)</li> </ul> <h3>Fixes</h3> <ul> <li>CLI/queue: drain active turns before releasing queue-owner leases and preserve typed retryable shutdown responses while terminating agent bridges. Thanks <a href="https://github.com/superWorldSavior"><code>@superWorldSavior</code></a>. (<a href="https://redirect.github.com/openclaw/acpx/pull/430">#430</a>)</li> <li>CLI/quiet output: emit exactly one structured stderr diagnostic for direct and queued prompt failures without adding diagnostics to stdout. Thanks <a href="https://github.com/superWorldSavior"><code>@superWorldSavior</code></a>. (<a href="https://redirect.github.com/openclaw/acpx/pull/431">#431</a>)</li> </ul> <h3>Verification</h3> <ul> <li>npm: <a href="https://www.npmjs.com/package/acpx/v/0.12.0">https://www.npmjs.com/package/acpx/v/0.12.0</a></li> <li>Registry tarball: <a href="https://registry.npmjs.org/acpx/-/acpx-0.12.0.tgz">https://registry.npmjs.org/acpx/-/acpx-0.12.0.tgz</a></li> <li>Integrity: <code>sha512-APYpN04XFWrCGuSBvM4HTKWWFH8uSIuzc+qI7aCGeVdP9o4euZeBosFEkmNUHvBOop0XBemg6d8RsNvzXN3Mgw==</code></li> <li>Candidate CI: <a href="https://github.com/openclaw/acpx/actions/runs/28718352566">https://github.com/openclaw/acpx/actions/runs/28718352566</a></li> <li>Trusted publish: <a href="https://github.com/openclaw/acpx/actions/runs/28718550655">https://github.com/openclaw/acpx/actions/runs/28718550655</a></li> <li>Full local tests, coverage, docs, conformance, mutation, security audits, packed-install CLI smoke, runtime export smoke, lifecycle proof, and fresh autoreview passed before tagging.</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/openclaw/acpx/blob/main/CHANGELOG.md">acpx's changelog</a>.</em></p> <blockquote> <h2>2026.7.4 (v0.12.0)</h2> <h3>Changes</h3> <ul> <li>Agents/built-ins: add Grok Build via <code>grok agent stdio</code>, including cached-login and <code>XAI_API_KEY</code> authentication selection. Thanks <a href="https://github.com/TheAngryPit"><code>@TheAngryPit</code></a>.</li> </ul> <h3>Breaking</h3> <h3>Fixes</h3> <ul> <li> <p>CLI/queue: drain active turns before releasing queue-owner leases and preserve typed retryable shutdown responses while terminating agent bridges. Thanks <a href="https://github.com/superWorldSavior"><code>@superWorldSavior</code></a>.</p> </li> <li> <p>CLI/quiet output: emit exactly one structured stderr diagnostic for direct and queued prompt failures without adding diagnostics to stdout. Thanks <a href="https://github.com/superWorldSavior"><code>@superWorldSavior</code></a>.</p> </li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/openclaw/acpx/commit/6a24a546d2349cbe71ed032d52d07cab611e320c"><code>6a24a54</code></a> chore(release): prepare acpx 0.12.0</li> <li><a href="https://github.com/openclaw/acpx/commit/ffd8549355adbb28bfc36ffd1087643d2df29fdb"><code>ffd8549</code></a> fix: surface quiet-mode failures on stderr (<a href="https://redirect.github.com/openclaw/acpx/issues/431">#431</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/87327c6e4ac28243b94945afb922c6cbb0dd0f98"><code>87327c6</code></a> fix: make queue owner shutdown lossless (<a href="https://redirect.github.com/openclaw/acpx/issues/430">#430</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/7d05c37b470cc357f0e5d6cefac7c74fdc434829"><code>7d05c37</code></a> feat: add Grok Build ACP agent (<a href="https://redirect.github.com/openclaw/acpx/issues/426">#426</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/95d75cd000a2f9f99b6e14e1d55e34ff4bf3cced"><code>95d75cd</code></a> chore(deps): bump tsx to 4.23.0 (<a href="https://redirect.github.com/openclaw/acpx/issues/429">#429</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/bce1c96d9d1a760cfdc233e19b040492c64fbeae"><code>bce1c96</code></a> chore(deps-dev): update development toolchain (<a href="https://redirect.github.com/openclaw/acpx/issues/428">#428</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/8a643e725ab31f9e2ec7f800c17ec8b845d61a14"><code>8a643e7</code></a> chore(deps-dev): refresh TypeScript native preview (<a href="https://redirect.github.com/openclaw/acpx/issues/427">#427</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/a18bf74c4b3a774357f993a83a0cfaf40465b044"><code>a18bf74</code></a> chore(deps): bump <code>@agentclientprotocol/sdk</code> from 0.28.1 to 1.1.0 (<a href="https://redirect.github.com/openclaw/acpx/issues/421">#421</a>)</li> <li><a href="https://github.com/openclaw/acpx/commit/1d882575e34e18621e59229f0e711723cef223ae"><code>1d88257</code></a> chore(deps-dev): bump the development group with 7 updates</li> <li><a href="https://github.com/openclaw/acpx/commit/c2689c025f36a8ebb76590e85ab44d499fac68ac"><code>c2689c0</code></a> chore(deps-dev): bump <code>@types/node</code> from 25.9.3 to 26.0.1</li> <li>Additional commits viewable in <a href="https://github.com/openclaw/acpx/compare/v0.11.2...v0.12.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
773bf45720 |
build(deps): bump @zed-industries/codex-acp from 0.12.0 to 0.16.0 (#9068)
Bumps [@zed-industries/codex-acp](https://github.com/zed-industries/codex-acp) from 0.12.0 to 0.16.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/zed-industries/codex-acp/releases">@zed-industries/codex-acp's releases</a>.</em></p> <blockquote> <h2>Release 0.16.0</h2> <h2>What's Changed</h2> <ul> <li>Update Codex dependencies to rust-v0.137.0 by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/320">zed-industries/codex-acp#320</a></li> <li>Use thread store for session listing by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/321">zed-industries/codex-acp#321</a></li> <li>chore(acp): Update to ACP 0.14.0 by <a href="https://github.com/mrjones2014"><code>@mrjones2014</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/317">zed-industries/codex-acp#317</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/mrjones2014"><code>@mrjones2014</code></a> made their first contribution in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/317">zed-industries/codex-acp#317</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/zed-industries/codex-acp/compare/v0.15.0...v0.16.0">https://github.com/zed-industries/codex-acp/compare/v0.15.0...v0.16.0</a></p> <h2>Release 0.15.0</h2> <h2>What's Changed</h2> <ul> <li>Update Codex to 0.133.0 by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/303">zed-industries/codex-acp#303</a></li> <li>Stop output buffer resend in terminal_interaction stdin path by <a href="https://github.com/frozename"><code>@frozename</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/270">zed-industries/codex-acp#270</a></li> <li>fix: Detach pending permission request tasks on cancel by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/304">zed-industries/codex-acp#304</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/zed-industries/codex-acp/compare/v0.14.0...v0.15.0">https://github.com/zed-industries/codex-acp/compare/v0.14.0...v0.15.0</a></p> <h2>Release 0.14.0</h2> <h2>What's Changed</h2> <ul> <li>Bump openssl from 0.10.78 to 0.10.79 in the cargo group across 1 directory by <a href="https://github.com/dependabot"><code>@dependabot</code></a>[bot] in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/265">zed-industries/codex-acp#265</a></li> <li>Update to codex 0.129 by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/268">zed-industries/codex-acp#268</a></li> <li>Fix O(N²) memory growth in exec_command_output_delta fallback by <a href="https://github.com/frozename"><code>@frozename</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/269">zed-industries/codex-acp#269</a></li> <li>Emit image generation tool calls by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/271">zed-industries/codex-acp#271</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/frozename"><code>@frozename</code></a> made their first contribution in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/269">zed-industries/codex-acp#269</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/zed-industries/codex-acp/compare/v0.13.0...v0.14.0">https://github.com/zed-industries/codex-acp/compare/v0.13.0...v0.14.0</a></p> <h2>Release 0.13.0</h2> <h2>What's Changed</h2> <ul> <li>Upgrade to codex 0.128.0 by <a href="https://github.com/benbrandt"><code>@benbrandt</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/261">zed-industries/codex-acp#261</a></li> <li>Reload auth file before failing check_auth() by <a href="https://github.com/anvilpete"><code>@anvilpete</code></a> in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/259">zed-industries/codex-acp#259</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/anvilpete"><code>@anvilpete</code></a> made their first contribution in <a href="https://redirect.github.com/zed-industries/codex-acp/pull/259">zed-industries/codex-acp#259</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/zed-industries/codex-acp/compare/v0.12.0...v0.13.0">https://github.com/zed-industries/codex-acp/compare/v0.12.0...v0.13.0</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/zed-industries/codex-acp/commit/bb590500e8646f6daf879b8b3c6a659fbd29017d"><code>bb59050</code></a> Bump version to 0.16.0</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/c81fa46d897275ed014065a3947596bf0fd1975b"><code>c81fa46</code></a> chore(acp): Update to ACP 0.14.0 (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/317">#317</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/4a853a1cf65309751a88d425c029ee81d3a33c48"><code>4a853a1</code></a> Use thread store for session listing (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/321">#321</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/9841e8b5a82f950f4ce30563884d18d93457e6dc"><code>9841e8b</code></a> Update Codex dependencies to rust-v0.137.0 (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/320">#320</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/863d433fc91855d0b5427372bf635c894bf68cb6"><code>863d433</code></a> v0.15.0</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/f67ca5f35feb82232ff736ba326c2b07dbf8cdd4"><code>f67ca5f</code></a> fix: Detach pending permission request tasks on cancel (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/304">#304</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/8aef91bc08de288531fc694248b5a370d5a3ade5"><code>8aef91b</code></a> Stop output buffer resend in terminal_interaction stdin path (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/270">#270</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/0c2d8280f26cc9583e44ca31d0a11f3f3f38d0b5"><code>0c2d828</code></a> Update README.md</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/d9bf1c157994feb9ebdc6117c8de0bc73dfe696f"><code>d9bf1c1</code></a> Update Codex to 0.133.0 (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/303">#303</a>)</li> <li><a href="https://github.com/zed-industries/codex-acp/commit/156cb0da12f6c7b1c697f90b5f22d5e14be31165"><code>156cb0d</code></a> Emit image generation tool calls (<a href="https://redirect.github.com/zed-industries/codex-acp/issues/271">#271</a>)</li> <li>Additional commits viewable in <a href="https://github.com/zed-industries/codex-acp/compare/v0.12.0...v0.16.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
ec2d87d353 |
fix(openclaw-gateway): bump adapter PROTOCOL_VERSION to 4 to match gateway (#5984)
## Summary Local OpenClaw gateways since v2026.5.x require `MIN_CLIENT_PROTOCOL_VERSION = 4` (see [openclaw/src/gateway/protocol/version.ts](https://github.com/NousResearch/openclaw/blob/main/src/gateway/protocol/version.ts)). The Paperclip openclaw_gateway adapter at v0.3.1 still sends `PROTOCOL_VERSION = 3`, which produces a WebSocket close (code 1002 \`protocol mismatch\`) before any auth challenge is issued. ## Symptom Every \`openclaw_gateway\` agent run fails with: - \`errorCode: openclaw_gateway_request_failed\` - \`error: protocol mismatch\` OpenClaw gateway journal: \`\`\` [ws] protocol mismatch conn=... remote=127.0.0.1 client=gateway-client backend vpaperclip [ws] closed before connect conn=... peer=...->127.0.0.1:18789 code=1002 reason=protocol mismatch \`\`\` ## Fix One-line bump in [\`packages/adapters/openclaw-gateway/src/server/execute.ts:89\`](packages/adapters/openclaw-gateway/src/server/execute.ts#L89): \`PROTOCOL_VERSION = 3\` → \`PROTOCOL_VERSION = 4\`. No protocol semantics changed — the adapter's existing frames are compatible with v4. ## Test plan - [x] \`pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck\` clean - [x] \`pnpm --filter @paperclipai/adapter-openclaw-gateway build\` clean - [x] Verified locally: openclaw_gateway agent connects + receives challenge + completes auth handshake after the bump ## Related - Same root cause affects every openclaw_gateway-backed agent in the field. Sparkeros companies SparkEros, Inc. and SparkEros AOS Inc. both hit it. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
70c86d2c73 |
fix(hermes): strip ANSI escape codes from terminal output in UI parsers (#8731)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Hermes adapter produces terminal output with ANSI color codes on
stdout
> - These escape sequences flow through the UI parsers untouched and
render as raw garbage text
> - This PR adds ANSI stripping at the entry point of all four Hermes
parse-stdout entry points
> - The same regex is already proven in claude-local adapter
> - The benefit is clean, readable terminal output for Hermes agents
## Linked Issues or Issue Description
No existing issue. This is a bug report:
**What happened**
Hermes terminal output displayed ANSI color codes as raw text in the
Paperclip UI, making agent output unreadable.
**Expected behavior**
Terminal output in run transcripts should be clean text without
invisible control characters.
**Steps to reproduce**
1. Connect a Hermes agent to Paperclip
2. Create and assign a task to the agent
3. View the run transcript — ANSI escape codes appear as raw garbage
**Paperclip version or commit**
|
||
|
|
ad961227f5 |
feat(secrets): add user-specific runtime secrets (#8825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs often need provider credentials, API tokens, and other environment-bound secrets. > - Company-level secrets work for shared credentials, but they do not model values that should differ by human operator. > - Without a user-scoped model, a run can dispatch without knowing whether the responsible human has supplied the needed value. > - Paperclip also needs run attribution to make those user-scoped runtime checks deterministic and auditable. > - This pull request adds user-specific secret definitions, per-user values, environment bindings, responsible-user attribution, and runtime resolution gates. > - The benefit is that teams can define the secret once, let each user provide their own value, and block runs before dispatch when required user secrets or active definitions are unavailable. ## Linked Issues or Issue Description Refs #224 Refs #6057 This PR implements user-specific secret support as a core secret-management capability rather than a one-off adapter setting. It is related to existing public work on company secrets UI and runtime secret refs, but is distinct because the value is owned by the responsible user and resolved at run dispatch time. Related PR search before opening found existing secrets work such as #1550, #8256, #8614, #8634, and #8647; none of those add the full user-secret definition/value/runtime gate covered here. ## What Changed - Added user-secret definitions and per-user "My secrets" values, keeping stored values out of access metadata. - Added `user_secret_ref` environment bindings and UI affordances to pick them alongside existing secret refs. - Added responsible-user runtime resolution so user-secret refs resolve against the human responsible for the run. - Added pre-dispatch missing-secret gates so runs fail before adapter dispatch when required user values are absent or definitions are inactive. - Added low-trust allowlist hardening for user-secret runtime access. - Added issue, routine, run, and agent API key responsible-user attribution and fail-closed dispatch behavior when attribution cannot be resolved. - Added denial-copy mapping so responsible-user authorization failures surface as actionable run outcomes instead of opaque setup failures. - Added OpenAPI documentation for the user-secret routes. - Rebases cleanly on current `master`; migrations were renumbered incrementally as `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant` after upstream `0126`/`0127` migrations. - Removed previously committed local design screenshots so the PR contains code/docs/tests only. ## Verification - PASS: PR head `2527febd106bcf3ca264ca0da7fca491084192d6` is based on `paperclipai/paperclip:master`. - PASS: `git diff --check` - PASS: `git diff --name-only public/master...HEAD | rg '^(pnpm-lock\\.yaml|\\.github/workflows/|screenshots/)' || true` produced no files. - PASS: migration journal audit confirmed unique indexes through `130` with tail entries `0126_issue_comment_derived_attribution`, `0127_environment_custom_images_instance_scoped`, `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant`. - PASS: `pnpm --filter @paperclipai/ui typecheck` - PASS: `pnpm --filter @paperclipai/server typecheck` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-active-run-output-watchdog.test.ts src/__tests__/heartbeat-stale-queue-invalidation.test.ts src/__tests__/heartbeat-workspace-finalize-branch.test.ts src/__tests__/issue-monitor-scheduler.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-comment-wake-batching.test.ts src/__tests__/heartbeat-retry-scheduling.test.ts src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts src/__tests__/heartbeat-plugin-environment.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/low-trust-red-team-routes.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` (55 tests) - PASS: `pnpm vitest run server/src/__tests__/secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` (89 tests after final Greptile cleanup fixes) - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-issue-liveness-escalation.test.ts` (17 tests after the final rebase CI fix) - PASS: focused server Vitest batches covering heartbeat recovery, project env, plugin env, routines, low-trust, pipelines, monitors, watchdog, and stale queue paths. - PASS: GitHub checks are green on `2527febd106bcf3ca264ca0da7fca491084192d6`, including Typecheck + Release Registry, Build, General tests, serialized server suites, e2e, Canary Dry Run, verify, security checks, and Greptile Review. - PASS: Greptile Review completed successfully on `2527febd106bcf3ca264ca0da7fca491084192d6` with Confidence Score 5/5, and GraphQL review-thread audit returned zero unresolved non-outdated threads. ## Risks - Runtime behavior now depends on a run having a correct responsible user; missing or incorrect responsibility assignment can block runs before adapter dispatch. - `user_secret_ref` bindings intentionally expose metadata without values, but UI/API callers may need to handle the new binding kind explicitly. - External secret providers and IAM policies are not automatically provisioned by this PR; operators still need to configure provider-side access for non-local vaults. - The PR is broad across db/shared/server/UI/runtime paths, so release validation should include both API and UI secret workflows before merge. - The migration renumbering is intentionally incremental after upstream migrations; the branch migrations use guarded column/table/index/constraint creation so users who tested the older draft numbering should not hit duplicate DDL for the existing objects. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), Codex local adapter with shell/tool use and code execution. Context window and internal reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c48feee190 |
Improve live agent feedback during sandboxed runs (#8915)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A core part of that experience is watching active agent runs without dropping into raw logs first > - Local and sandbox-backed adapters already record useful run output, progress, and tool activity > - But active issue threads could sit visually stale while the agent was syncing workspaces, tailing sandbox output, or emitting incremental tool-call updates > - Operators need timely, human-readable progress while preserving the raw transcript underneath > - This pull request streams sandbox run-log progress into runtime status, keeps visible issue threads refreshed, and folds repeated ACPX tool updates into stable transcript cards > - The benefit is that long-running agent work becomes easier to supervise without changing the task/comment control-plane model ## Linked Issues or Issue Description No public GitHub issue exists for this exact change. Problem/motivation: - During long-running sandboxed agent work, the issue UI can appear idle even though the agent is actively syncing, running tools, or producing incremental output. - Operators need realtime feedback at the issue-thread layer, not only after opening raw logs or waiting for the final heartbeat result. - Related public context: #1808 previously added live-run status dots to Projects; #4362 touches heartbeat wakeup behavior but is not a duplicate of this runtime/UI feedback change. ## What Changed - Added sandbox run-log streaming support and defaulted sandbox-capable local adapters into the richer live-feedback path. - Surfaced environment/sandbox sync progress through heartbeat runtime status with bounded, redacted snippets. - Added live issue-thread cache patching so visible active runs update as progress events arrive. - Folded repeated ACPX `tool_call` updates into one transcript card instead of stacking duplicate cards. - Updated adapter docs and added focused regression coverage for sandbox log streaming, runtime status, ACPX parsing, live updates, transcript rendering, and issue chat messages. ## Verification - `pnpm install --frozen-lockfile` - `pnpm exec vitest run ui/src/context/LiveUpdatesProvider.test.ts` - `pnpm exec vitest run server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts ui/src/context/LiveUpdatesProvider.test.ts` - `pnpm exec vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/agent-live-run-routes.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts packages/adapters/acpx-local/src/ui/parse-stdout.test.ts ui/src/context/LiveUpdatesProvider.test.ts ui/src/components/transcript/RunTranscriptView.test.tsx ui/src/lib/issue-chat-messages.test.ts ui/src/components/IssueChatThread.test.tsx` - GitHub PR workflow on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`: `verify`, build, typecheck/release-registry, e2e, general shards, serialized server shards, and canary dry run passed. - Greptile Review on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`: Confidence Score 5/5, no unresolved review threads. ## Risks - Live issue-thread cache patching could miss an edge case for a route shape not covered by tests. - Surfacing active-run snippets needs continued care around redaction; this PR keeps snippets bounded and adds redaction-focused coverage. - More frequent active-run UI refreshes could expose performance issues on very large issue threads, though updates are scoped to visible run/query caches. ## Model Used OpenAI GPT-5 via Codex, operating as a tool-enabled coding agent with shell, git, and repository-editing capabilities. Context window size is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bf982c8c83 |
Normalize adapter display labels (#8913)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter names are part of the board-facing agent setup and management experience. > - The product now treats adapters as harnesses, while execution environments are modeled separately. > - Several built-in adapter labels still carried legacy local wording from the older harness-by-environment model. > - That wording makes the UI noisier and implies a distinction users no longer need to reason about. > - This pull request normalizes adapter display labels while keeping persisted adapter type identifiers unchanged. > - The benefit is clearer adapter selection and management copy without a database migration. ## Linked Issues or Issue Description No public GitHub issue was found for this exact cleanup. Related public PRs: - Supersedes #8910, an earlier branch for the same cleanup that did not include the later docs/gateway/Cursor alignment. - Refs #8819, which is related display-registry work for external multi-segment adapter labels, but not a duplicate of this built-in label cleanup. Feature request details: - Subsystem affected: Cross-cutting (`ui/`, `packages/adapters`, and docs). - Problem or motivation: user-facing adapter names include legacy local qualifiers even though adapters map to harnesses and environments are first-class elsewhere. - Proposed solution: remove the legacy local wording from built-in display labels, keep machine-readable adapter type ids unchanged, and keep gateway disambiguation where it is useful. - Alternatives considered: changing persisted adapter type ids was ruled out because it would create migration and compatibility risk; one-off UI replacements were ruled out because the display registry is already the correct central label boundary. - Roadmap alignment: this is small adapter UX polish, not a new roadmap-level core feature. ## What Changed - Updated the adapter display registry so known adapter labels are final and no built-in local adapter renders a legacy local suffix. - Preserved clean derived labels for unknown plugin local types while keeping gateway disambiguation for unknown gateway types. - Updated `AdapterManager` to prefer registry labels when the server reports raw adapter type ids for built-ins. - Removed legacy local wording from built-in adapter metadata labels in UI and adapter packages. - Aligned Cursor adapter metadata with the central display registry label. - Updated adapter docs and Storybook fixtures to match the new display names. - Added focused registry coverage for built-in labels and unknown plugin suffix behavior. ## Verification - `pnpm check:tokens` - `git diff --check origin/master...fix/adapter-display-labels` - Patch-addition scan for added secrets, private paths, and internal links: no matches. - GitHub duplicate search for open adapter-label/local-suffix issues and PRs; #8910 was identified as the older superseded public PR. - `pnpm exec vitest run ui/src/adapters/adapter-display-registry.test.ts` - `pnpm --filter @paperclipai/ui typecheck` - Stale-label scan found no remaining user-facing display-label suffixes; remaining local wording is operational/test terminology such as adapter ids, docs about running locally, and test descriptions. ## Risks Low risk. The change is display-label and documentation focused, and adapter type ids remain unchanged. The main risk is ambiguous gateway naming, mitigated by keeping explicit gateway labels where variants need disambiguation. ## Model Used OpenAI GPT-5 via Codex, tool-enabled coding agent in a local repository workspace. Context window size is not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
41059841f1 |
feat(gemini-local): add Gemini 3.1 Pro models to gemini-local adapter (#1602)
**Thinking path** The gemini_local adapter's model dropdown (packages/adapters/gemini-local/src/index.ts) still lists only Gemini 2.x. Google has since shipped the Gemini 3.1 Pro family, so users can't pick the current flagship from the dashboard and instead hit ModelNotFoundError when they type an ID by hand (#1506). The adapter passes the selected string straight to `gemini --model`, so the fix is to surface the valid 3.1 Pro IDs in the list. **What I did** Added two entries to the `models` array, above the existing 2.x entries: - `gemini-3.1-pro-preview` — Gemini 3.1 Pro (Preview) - `gemini-3.1-pro-preview-customtools` — custom-tools variant, tuned for agentic/tool use **Why it matters** Users can select the current flagship 3.1 Pro (and its custom-tools endpoint) directly, instead of guessing IDs and hitting ModelNotFoundError. **How to verify** Open the gemini_local model dropdown in the dashboard; both entries appear above the 2.5 entries and run against `gemini --model <id>` without error. **Risks** Minimal — additive, single-file change to a static list; nothing removed. Both IDs are confirmed-valid Google API identifiers. Fixes #1506. |
||
|
|
5bd6c6ec3c |
build(deps): bump acpx from 0.6.1 to 0.11.2 (#8741)
Bumps [acpx](https://github.com/openclaw/acpx) from 0.6.1 to 0.11.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/openclaw/acpx/releases">acpx's releases</a>.</em></p> <blockquote> <h2>acpx 0.11.0</h2> <h2>v0.11.0</h2> <h3>Changes</h3> <ul> <li>Agents/built-ins: bump the default Claude ACP adapter range to <code>@agentclientprotocol/claude-agent-acp@^0.37.0</code>. Thanks <a href="https://github.com/trumpyla"><code>@trumpyla</code></a>.</li> <li>Runtime/embedding: surface cost, token usage breakdowns, and advertised command metadata on runtime status/events. Thanks <a href="https://github.com/DaniAkash"><code>@DaniAkash</code></a>.</li> <li>Agents/built-ins: add <code>fast-agent</code> as a built-in fast-agent ACP adapter via <code>uvx fast-agent-mcp acp</code>.</li> <li>Agents/built-ins: add <code>mux</code> as a built-in coder/mux ACP adapter via <code>npx -y mux@^0.27.0 acp</code>. Thanks <a href="https://github.com/ThomasK33"><code>@ThomasK33</code></a>.</li> <li>CLI: add <code>acpx compare</code> to run one prompt across multiple agents and summarize timing, token usage, stop reason, permissions, and final output side by side. Thanks <a href="https://github.com/mvanhorn"><code>@mvanhorn</code></a>.</li> </ul> <h3>Fixes</h3> <ul> <li>CLI/Claude: isolate built-in Claude ACP sessions from user settings by default so globally enabled channel and daemon plugins cannot interfere with a spawned session. Set <code>ACPX_CLAUDE_INCLUDE_USER_SETTINGS=1</code> to restore user settings deliberately. Fixes <a href="https://redirect.github.com/openclaw/acpx/issues/361">#361</a>.</li> <li>ACP/models: support SDK 0.25 model config options while preserving <code>session/set_model</code> compatibility for adapters that explicitly advertise legacy model metadata.</li> <li>CLI/Claude: let Claude Code adjudicate model selectors missing from a stale advertised model list on later persistent turns, and preserve the adapter-reported current model after model switches. Thanks <a href="https://github.com/oakif"><code>@oakif</code></a>.</li> <li>Client/ACP: advertise scoped Devin/Windsurf-compatible client metadata and handle Devin extension requests/notifications without noisy method-not-found logs. Thanks <a href="https://github.com/LivioGama"><code>@LivioGama</code></a>.</li> <li>Runtime/sessions: treat corrupt public file-session records as missing while preserving genuine filesystem errors. Thanks <a href="https://github.com/KrasimirKralev"><code>@KrasimirKralev</code></a>.</li> </ul> <h3>Verification</h3> <ul> <li>npm: <a href="https://www.npmjs.com/package/acpx/v/0.11.0">https://www.npmjs.com/package/acpx/v/0.11.0</a></li> <li>Registry tarball: <a href="https://registry.npmjs.org/acpx/-/acpx-0.11.0.tgz">https://registry.npmjs.org/acpx/-/acpx-0.11.0.tgz</a></li> <li>Integrity: <code>sha512-l42LJFmd6kvbr1UytvwWmr5Mdy/v9l3FM6Necs01PWbjUIkxjCdxg97duqoRfRqxtDAfnNPb1IlgIf2ZgMZQqA==</code></li> <li>Candidate CI: <a href="https://github.com/openclaw/acpx/actions/runs/27676359224">https://github.com/openclaw/acpx/actions/runs/27676359224</a></li> <li>Trusted publish: <a href="https://github.com/openclaw/acpx/actions/runs/27676823793">https://github.com/openclaw/acpx/actions/runs/27676823793</a></li> <li>Packed CLI and real Codex ACP adapter E2E passed before tagging.</li> </ul> <h2>2026.5.23 (v0.10.0)</h2> <h3>Changes</h3> <ul> <li>CLI/sessions: add <code>sessions export</code> and <code>sessions import</code> for moving portable session archives between machines. Thanks <a href="https://github.com/mvanhorn"><code>@mvanhorn</code></a>.</li> </ul> <h3>Release Proof</h3> <ul> <li>npm: <a href="https://www.npmjs.com/package/acpx/v/0.10.0">https://www.npmjs.com/package/acpx/v/0.10.0</a></li> <li>registry tarball: <a href="https://registry.npmjs.org/acpx/-/acpx-0.10.0.tgz">https://registry.npmjs.org/acpx/-/acpx-0.10.0.tgz</a></li> <li>integrity: <code>sha512-hd48XV03gG3sd409T1lDrOKJTTz1ap4g0wrndXjxQ590tN85pBYlvfNLyerybvGRrtUGsZjNdt99r1jpIt6ukA==</code></li> <li>release workflow: <a href="https://github.com/openclaw/acpx/actions/runs/26323055145">https://github.com/openclaw/acpx/actions/runs/26323055145</a></li> <li>CI: <a href="https://github.com/openclaw/acpx/actions/runs/26323053524">https://github.com/openclaw/acpx/actions/runs/26323053524</a></li> </ul> <h2>2026.5.22 (v0.9.0)</h2> <h3>Changes</h3> <ul> <li>Tooling: add Slophammer TypeScript quality gates for coverage, complexity, unsafe types, mutation testing, DRY checks, and dependency boundaries.</li> <li>Agents/built-ins: switch the default Codex adapter to <code>@agentclientprotocol/codex-acp</code>, with Codex model selection handled through advertised ACP model ids, and bump the default Claude ACP adapter range.</li> <li>Tooling: add a repo-local autoreview skill and helper for Codex-first</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/openclaw/acpx/blob/main/CHANGELOG.md">acpx's changelog</a>.</em></p> <blockquote> <h2>2026.6.23 (v0.11.2)</h2> <h3>Changes</h3> <h3>Breaking</h3> <h3>Fixes</h3> <ul> <li>Runtime/status: persist token usage reported on successful prompt responses, including adapters that only provide a sparse <code>usage_update</code>.</li> </ul> <h2>Unreleased</h2> <h3>Changes</h3> <h3>Breaking</h3> <h3>Fixes</h3> <h2>2026.6.23 (v0.11.1)</h2> <h3>Changes</h3> <ul> <li>Runtime/embedding: preserve per-agent environment variables across ACP session creation, queue handoff, persistence, and reconnects. Thanks <a href="https://github.com/zhangguiping-xydt"><code>@zhangguiping-xydt</code></a>.</li> </ul> <h3>Breaking</h3> <h3>Fixes</h3> <ul> <li>CLI/queue: harden command parsing, queue-owner startup, stale process cleanup, and release/CI checks found by <code>clawpatch</code>.</li> <li>Windows/Claude: only export a native <code>.exe</code> as <code>CLAUDE_CODE_EXECUTABLE</code>; unresolved <code>.cmd</code>, <code>.bat</code>, and <code>.ps1</code> shims now fall back to the Claude ACP adapter's bundled native binary. Fixes <a href="https://redirect.github.com/openclaw/openclaw/issues/93465">openclaw/openclaw#93465</a>.</li> <li>Client/ACP: ignore non-object JSON lines from adapter stdout before ACP dispatch, preventing primitive frames from crashing the SDK message path.</li> <li>ACP/models: call the current SDK <code>session/set_model</code> method for legacy model metadata instead of the generic extension fallback.</li> <li>CLI/config: add <code>--mcp-config</code> for session-scoped MCP servers without writing a project config file. Live persistent sessions reject MCP config changes until closed. Fixes <a href="https://redirect.github.com/openclaw/acpx/issues/387">#387</a>.</li> </ul> <h2>2026.6.17 (v0.11.0)</h2> <h3>Changes</h3> <ul> <li>Agents/built-ins: bump the default Claude ACP adapter range to <code>@agentclientprotocol/claude-agent-acp@^0.37.0</code>. Thanks <a href="https://github.com/trumpyla"><code>@trumpyla</code></a>.</li> <li>Runtime/embedding: surface cost, token usage breakdowns, and advertised command metadata on runtime status/events. Thanks <a href="https://github.com/DaniAkash"><code>@DaniAkash</code></a>.</li> <li>Agents/built-ins: add <code>fast-agent</code> as a built-in fast-agent ACP adapter via <code>uvx fast-agent-mcp acp</code>.</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/openclaw/acpx/commits/v0.11.2">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1b180fae1f |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.48.0 to 0.52.0 (#8745)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.48.0 to 0.52.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.52.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.51.0...v0.52.0">0.52.0</a> (2026-06-25)</h2> <h3>Features</h3> <ul> <li>Add version flag handling (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/813">#813</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9616bdac47505e4a14c36d667fcffc9ae97e1f2a">9616bda</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/809">#809</a></li> <li><strong>deps-dev:</strong> bump expect-type from 1.3.0 to 1.4.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/814">#814</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61272acb30dcafaa2455d334b11ce2ad97339707">61272ac</a>)</li> <li><strong>deps:</strong> Update <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.191 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/810">#810</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/228f02ecfb23be16e59e121c6b42c0f2b2f40a4e">228f02e</a>)</li> <li>Push session title updates at turn end (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/812">#812</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1fe7ec09a3a7bcb7501231dae4b0ffe6ef9b70a4">1fe7ec0</a>)</li> </ul> <h2>v0.51.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.50.0...v0.51.0">0.51.0</a> (2026-06-24)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/807">#807</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8f6ebd1d9198edf723f4c8c1aa2b49b906c46646">8f6ebd1</a>)</li> </ul> <h2>v0.50.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.49.0...v0.50.0">0.50.0</a> (2026-06-23)</h2> <h3>Features</h3> <ul> <li><strong>acp:</strong> Handle ACP request cancellation signals (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/801">#801</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9013d1d46883a7f3774a63a9dca16c4a0f634a97">9013d1d</a>)</li> <li><strong>deps:</strong> bump actions/checkout from 6.0.3 to 7.0.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/803">#803</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/044c43e0c894082b9e747c01e9dbde6e21036823">044c43e</a>)</li> <li><strong>deps:</strong> upgrade to <code>@anthropic-ai/claude-agent-sdk</code><a href="https://github.com/0"><code>@0</code></a>.3.186 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/806">#806</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a7e6137f6877b72b8daa39e74971c3559db8d28f">a7e6137</a>)</li> </ul> <h2>v0.49.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.48.0...v0.49.0">0.49.0</a> (2026-06-22)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.185 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/798">#798</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8dc8c864263fa50d03bec7ebff7aee776604cb3d">8dc8c86</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Deduplicate streamed assistant blocks by content (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/800">#800</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/960f62d76582ae5c9c5575ab66974809049ce1d0">960f62d</a>)</li> <li>Infer 1M context from model descriptions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/799">#799</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/508453c288b4a12701abd507199e7fa0ab172171">508453c</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.51.0...v0.52.0">0.52.0</a> (2026-06-25)</h2> <h3>Features</h3> <ul> <li>Add version flag handling (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/813">#813</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9616bdac47505e4a14c36d667fcffc9ae97e1f2a">9616bda</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/809">#809</a></li> <li><strong>deps-dev:</strong> bump expect-type from 1.3.0 to 1.4.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/814">#814</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61272acb30dcafaa2455d334b11ce2ad97339707">61272ac</a>)</li> <li><strong>deps:</strong> Update <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.191 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/810">#810</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/228f02ecfb23be16e59e121c6b42c0f2b2f40a4e">228f02e</a>)</li> <li>Push session title updates at turn end (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/812">#812</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1fe7ec09a3a7bcb7501231dae4b0ffe6ef9b70a4">1fe7ec0</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.50.0...v0.51.0">0.51.0</a> (2026-06-24)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/807">#807</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8f6ebd1d9198edf723f4c8c1aa2b49b906c46646">8f6ebd1</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.49.0...v0.50.0">0.50.0</a> (2026-06-23)</h2> <h3>Features</h3> <ul> <li><strong>acp:</strong> Handle ACP request cancellation signals (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/801">#801</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9013d1d46883a7f3774a63a9dca16c4a0f634a97">9013d1d</a>)</li> <li><strong>deps:</strong> bump actions/checkout from 6.0.3 to 7.0.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/803">#803</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/044c43e0c894082b9e747c01e9dbde6e21036823">044c43e</a>)</li> <li><strong>deps:</strong> upgrade to <code>@anthropic-ai/claude-agent-sdk</code><a href="https://github.com/0"><code>@0</code></a>.3.186 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/806">#806</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a7e6137f6877b72b8daa39e74971c3559db8d28f">a7e6137</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.48.0...v0.49.0">0.49.0</a> (2026-06-22)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.185 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/798">#798</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8dc8c864263fa50d03bec7ebff7aee776604cb3d">8dc8c86</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Deduplicate streamed assistant blocks by content (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/800">#800</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/960f62d76582ae5c9c5575ab66974809049ce1d0">960f62d</a>)</li> <li>Infer 1M context from model descriptions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/799">#799</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/508453c288b4a12701abd507199e7fa0ab172171">508453c</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e9163f855398511c6eec0116216d8c06c52ec658"><code>e9163f8</code></a> chore(main): release 0.52.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/811">#811</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/61272acb30dcafaa2455d334b11ce2ad97339707"><code>61272ac</code></a> feat(deps-dev): bump expect-type from 1.3.0 to 1.4.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/814">#814</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9616bdac47505e4a14c36d667fcffc9ae97e1f2a"><code>9616bda</code></a> feat: Add version flag handling (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/813">#813</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1fe7ec09a3a7bcb7501231dae4b0ffe6ef9b70a4"><code>1fe7ec0</code></a> feat: Push session title updates at turn end (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/812">#812</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/228f02ecfb23be16e59e121c6b42c0f2b2f40a4e"><code>228f02e</code></a> feat(deps): Update <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.191 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/810">#810</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/23626c9a43b4fa2b4e98cf1abb25c55985711075"><code>23626c9</code></a> chore(main): release 0.51.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/808">#808</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8f6ebd1d9198edf723f4c8c1aa2b49b906c46646"><code>8f6ebd1</code></a> feat(deps): bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/807">#807</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/07911601ccd2b8d8628aed545f7715c6aa0fa429"><code>0791160</code></a> chore(main): release 0.50.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/802">#802</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a7e6137f6877b72b8daa39e74971c3559db8d28f"><code>a7e6137</code></a> feat(deps): upgrade to <code>@anthropic-ai/claude-agent-sdk</code><a href="https://github.com/0"><code>@0</code></a>.3.186 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/806">#806</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/044c43e0c894082b9e747c01e9dbde6e21036823"><code>044c43e</code></a> feat(deps): bump actions/checkout from 6.0.3 to 7.0.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/803">#803</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.48.0...v0.52.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
8d9f9fd240 |
Add reusable sandbox custom images (#8794)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A growing part of that work runs in sandboxed environments rather than on the operator's local machine. > - Today sandbox providers can start fresh workspaces and run probes, but they do not have a shared contract for capturing and reusing prepared sandbox state. > - Operators need a way to set up tools, credentials, and project dependencies once, then reuse that prepared image for later agent runs. > - This pull request adds reusable sandbox custom images across the provider contract, server runtime, and board UI. > - It also keeps probes and sandbox copy flows aligned with pre-authenticated/custom-image environments. > - The benefit is faster, more reliable sandbox runs without repeatedly rebuilding the same environment setup. ## Linked Issues or Issue Description No public GitHub issue was found for this change. Inline feature request follows. ### Problem or motivation Sandboxed agents need reusable prepared runtime state so repeated runs do not require manual setup every time. Operators often need system packages, CLIs, SDKs, dependency caches, credentials, and project tooling available before an agent can work productively. ### Proposed solution Add a provider-level custom-image capability, server-side setup/capture lifecycle, Daytona/fake provider support, and board UI controls for creating, testing, selecting, and deleting custom images. ### Alternatives considered Leaving this as provider-specific setup outside Paperclip would keep the control plane blind to image state and would not give agents consistent environment metadata. Re-running setup commands for every lease is simpler, but slower and less reliable for interactive or credentialed setup. ### Roadmap alignment Checked `ROADMAP.md`; this aligns with the Cloud / Sandbox agents roadmap area and does not duplicate any related public issue or PR found by search. Additional context: - Subsystem affected: cross-cutting (`packages/db`, `packages/shared`, `packages/plugins`, `server`, `ui`). - Duplicate search: searched GitHub for `sandbox custom image` and `sandbox template environment`; no related public issues or PRs were found. ## What Changed - Added custom-image shared types, validators, constants, API paths, and database schema/migration. - Added server services/routes for custom-image templates and setup sessions, including runtime cleanup and provider metadata handling. - Extended plugin/sandbox provider capabilities for interactive setup, template capture, and template deletion. - Implemented custom-image support in the fake sandbox provider and Daytona provider. - Updated environment runtime/config handling so active custom images flow into leases, probes, and agent execution. - Added board UI controls and API client support for custom-image setup, capture, selection, status, and error states. - Hardened sandbox copy/probe behavior for insecure clipboard contexts and pre-authenticated sandbox images. - Added targeted coverage across shared validators, DB schema, server routes/services, provider plugins, adapter probes, and UI flows. ## Verification - `pnpm install --frozen-lockfile --ignore-scripts` - `pnpm vitest run packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/test.ts packages/adapters/codex-local/src/server/test.remote.test.ts packages/adapters/codex-local/src/server/test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm vitest run packages/db/src/environment-custom-images-schema.test.ts packages/shared/src/environment-custom-images.test.ts packages/shared/src/validators/plugin.test.ts server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/workspace-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` - `pnpm -r typecheck` - `pnpm build` - `rm -rf packages/db/dist && pnpm test:run` - Public-safety scan of the final diff found no internal Paperclip issue links, private instance URLs, or real secret patterns. ## Risks - Adds a database migration and new environment runtime tables, so migration ordering and rollback need care. - Provider implementations may differ in how reliably they can capture/delete images; unsupported providers surface capability-gated UI states. - Custom-image state can contain operator-prepared tooling and credentials inside the provider image, so providers must enforce their own access controls and cleanup semantics. - Broad surface area across shared contracts, server runtime, plugins, adapters, and UI means CI and Greptile review should be watched closely. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (`gpt-5`) via Codex CLI with tool use and code execution. Assisted with branch cleanup, conflict resolution, local verification, and PR preparation. Earlier branch implementation work was assisted by Paperclip-managed Claude/Codex agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0c2ec7deb4 |
Fix sandboxed Claude and Codex probe behavior (#8775)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters are the bridge between Paperclip's control plane and provider CLIs such as Claude Code and Codex. > - Those adapters can run either on the host machine or inside a remote/sandbox execution target. > - Sandbox probes need to validate the same auth/config path that real sandbox execution will use. > - The previous probe paths could surface misleading Claude errors, rely on host-only Codex state, or upload far more Codex home state than the probe needed. > - This pull request fixes the Claude and Codex sandbox probe/runtime behavior together while keeping provider-specific sandbox image work out of scope. > - The benefit is faster, clearer adapter health checks that better match real sandbox execution. ## Linked Issues or Issue Description No public GitHub issue was found for this exact bug during duplicate search. Bug report: **What happened?** Sandboxed Claude/Codex adapter tests could diverge from real runtime auth/config behavior. Claude sandbox probes could show the leading stream init line instead of the real final error, and Codex sandbox probes could upload full managed home state or mask a sandbox-local login with an empty uploaded `CODEX_HOME`. **Expected behavior** Sandbox probes should exercise the remote runtime contract, preserve useful sandbox credentials, avoid relying on unrelated host state, and report actionable probe failures. **Steps to reproduce** 1. Configure a remote/sandbox execution target for `claude_local` or `codex_local`. 2. Run the environment Test/probe path where host credentials differ from the sandbox's runtime credentials or the managed Codex home contains session history. 3. Observe that probe behavior can differ from the actual sandbox runtime path or surface an unhelpful Claude stream initialization line. **Paperclip version or commit** Current `master` before this PR, based on `4a2447da3`. **Deployment mode** Local development/control-plane deployment with remote sandbox execution targets. Related search performed: - Public issues: `Claude sandbox probe`, `Codex CODEX_HOME sandbox` returned no matches. - Public PRs: `Claude Codex sandbox probe`, `codex home sandbox`, `claude auth sandbox` returned no matches. ## What Changed - Made Claude sandbox Test probes materialize the same Paperclip-managed Claude config seed path used by sandbox execution. - Preserved sandbox-local Claude credentials when materializing remote Claude config and expanded auth-required detection for `/login` API-key failures. - Improved Claude hello-probe diagnostics so the final result/error is surfaced instead of the unhelpful stream init event, with transient upstream failures downgraded to warnings. - Changed Codex probe behavior to upload only minimal auth/config files instead of the full managed `CODEX_HOME`. - Let Codex sandbox probes leave `CODEX_HOME` unset when the host has no credentials, so pre-authenticated sandbox images can be tested directly. - Excluded bulky host-local Codex session/shell state from sandbox runtime home uploads. - Switched the Codex local default model away from the ChatGPT-unsupported `gpt-5.3-codex` option. - Added regression coverage for Claude parsing/probe paths, Codex adapter metadata/argument/probe behavior, and server-level Claude sandbox environment behavior. ## Verification Passed locally: - `pnpm install --frozen-lockfile` - `pnpm vitest run packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/test.probe.test.ts server/src/__tests__/claude-local-adapter-environment.test.ts` - `pnpm vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/server/test.remote.test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks - Adapter configuration behavior is sensitive to local vs sandboxed execution mode, so review should focus on environment detection, argument construction, and any state written during probe/test runs. - The Codex default-model change may affect newly created agents that rely on the adapter default instead of an explicit model. - Excluding Codex session/shell state from sandbox uploads should be safe for fresh sandbox runs, but reviewers should confirm no runtime resume path depends on that host-local state. - Provider-specific setup/capture behavior is intentionally left to separate work. ## Model Used OpenAI GPT-5 Codex via Paperclip `codex_local`; tool-enabled local coding session with terminal access. Context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d77fab6aae |
fix(adapters/hermes-gateway): improve onboarding configuration (#8678)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs agents through adapters, including built-in local and gateway-style adapters. > - Hermes gateway users need to connect Paperclip to an already-running Hermes API server. > - The gateway setup flow was missing clear non-local adapter configuration fields and accepted fewer URL shapes than operators naturally paste from Hermes. > - It also surfaced sparse diagnostics when the gateway was unreachable or when Paperclip generated onboarding prompts for gateway agents. > - This pull request tightens Hermes gateway configuration, URL normalization, diagnostics, and onboarding defaults. > - The benefit is that Hermes gateway setup is easier to complete and easier to debug without affecting unrelated adapters. ## Linked Issues or Issue Description No public issue found for this exact follow-up. Related prior/in-flight Hermes work: - Refs #2363 - Refs #4359 - Refs #6473 Problem statement: - **Type:** Adapter follow-up / setup reliability - **Adapter:** `hermes_gateway` - **Motivation:** Operators configure `hermes_gateway` against a running Hermes API server, but the UI and onboarding flow did not expose enough gateway-specific configuration or diagnostics. - **Expected behavior:** Paperclip should render the gateway fields, normalize common Hermes dashboard/API URL inputs, preserve sensible gateway onboarding defaults, and report reachability failures with actionable detail. - **Deployment mode:** Built-in adapter package in the Paperclip monorepo. ## What Changed - Added UI config fields for non-local Hermes gateway settings, including tests for rendering and field behavior. - Accepted Hermes dashboard URLs by normalizing them to gateway API URLs for execution. - Improved gateway reachability and run URL diagnostics. - Updated Hermes gateway onboarding text/default behavior so join prompts preserve gateway configuration. - Added focused server, UI, and adapter tests for the gateway configuration and onboarding paths. ## Verification - `pnpm install --frozen-lockfile --prefer-offline` - `pnpm --filter @paperclipai/hermes-paperclip-adapter exec vitest run src/gateway/server/execute.test.ts` — 20 passed - `pnpm exec vitest run server/src/__tests__/invite-accept-gateway-defaults.test.ts server/src/__tests__/invite-onboarding-text.test.ts ui/src/adapters/hermes-gateway/config-fields.test.tsx ui/src/components/AgentConfigForm.render.test.tsx ui/src/lib/agent-onboarding-prompt.test.ts` — 23 passed across 5 files - Confirmed the PR diff excludes `pnpm-lock.yaml` and `.github/workflows`. ## Risks Low to moderate risk. The changes are scoped to Hermes gateway configuration/onboarding and generic non-local adapter field rendering. The main risk is rejecting an unusual Hermes URL shape that should be accepted; the normalization tests cover dashboard and API URL variants added here. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent in a tool-enabled local CLI environment, with shell/GitHub/Paperclip API access. Exact runtime model identifier was not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fd2f82ac5b |
[codex] Add built-in Hermes adapters (#8543)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters are the boundary between the control plane and the runtimes that actually do work. > - Hermes support needs to be available as first-class local and gateway adapters while still preserving the adapter-manager override path for external packages. > - The adapter work touches runtime execution, UI adapter metadata, onboarding prompts, scoped credentials, release packaging, and smoke coverage, so the handoff needs concrete verification rather than only unit tests. > - This pull request adds built-in Hermes local and Hermes gateway support, keeps external adapter overrides compatible, and documents/tests the gateway flow end to end. > - The benefit is that operators can hire Hermes-backed agents without a manual plugin install, while self-hosted installs can still override/shadow the built-ins through Adapter manager packages. ## Linked Issues or Issue Description No public GitHub issue exists for this exact Hermes built-in adapter, gateway onboarding, and release-source work. Problem description: - Hermes local and gateway adapters need a public, reviewable source path in the monorepo so package artifacts and built-in adapter behavior match the application source. - Operators need built-in `hermes_local` and `hermes_gateway` adapter choices without losing the ability to install external Hermes packages as overrides. - Gateway onboarding needs secure defaults for API server URLs, API keys, and generated agent setup text. - Hermes-originated task bridge credentials need narrower API-key scope configuration. - Related public PRs found during duplicate search include #3027, #2363, #7544, #7950, #8095, and #8543. ## What Changed - Added the unified Hermes adapter package with local and gateway server/UI/CLI exports, config schemas, transcript parsing, model detection, and package metadata. - Registered `hermes_local` and `hermes_gateway` as built-in adapters across shared constants, server registries, CLI packaging, and UI adapter registries. - Kept the external adapter override path compatible so installed Hermes packages can shadow built-ins and restore the built-in parser when disabled. - Added Hermes gateway onboarding docs, board-operator docs, Docker smoke assets, and shell smoke harnesses for join/e2e validation. - Added scoped task-bridge API-key support, authorization checks, issue-origin handling, and tests for Hermes-created Paperclip tasks. - Hardened gateway transport and redaction behavior for API keys, headers, session data, and smoke diagnostics. - Updated release packaging/bootstrap checks for the Hermes packages while leaving `pnpm-lock.yaml` out of the PR per repository policy. ## Verification Targeted local verification recorded before PR handoff: - `pnpm --filter @paperclipai/hermes-paperclip-adapter exec vitest run src/gateway/server/execute.test.ts` — 14/14 passed. - `pnpm test:hermes-gateway-smoke` — 6/6 passed. - Hermes package typecheck/build checks passed. - Focused server/UI adapter tests passed — 31/31. - Release helper Node tests passed — 18/18. - `git diff --check origin/master..HEAD` passed. Fresh Docker E2E smoke evidence: - Ran `pnpm smoke:hermes-gateway-e2e` on 2026-06-26 with a fresh state directory and fresh Docker container against a live Paperclip dev server. - Hermes direct execution reached `completed`. - Hermes stop/cancel path reached `cancelled`. - Hermes gateway created a Paperclip task, Paperclip ran the Hermes agent, and the task reached `done` with the expected marker response. - Temporary board auth keys, token files, smoke state, and Docker containers were cleaned up after the run. PR checks on head `b5eae40ce`: - GitHub Actions passed: `policy`, `review`, `Typecheck + Release Registry`, all general test shards, all serialized server shards, `Build`, `Canary Dry Run`, `e2e`, and aggregate `verify`. - External checks passed: Snyk and Socket Project Report. - External Socket Pull Request Alerts remained pending after the first-party CI matrix completed. ## Risks - Medium risk: this spans adapter registration, package publishing, gateway execution, onboarding docs, API-key scoping, and UI adapter metadata. - Migration risk is low: the scope-config migration adds a nullable column and does not rewrite existing keys. - Gateway execution depends on operator-provided Hermes API configuration; the smoke covers the Docker gateway path but real deployments may differ by network/auth setup. - Direct Greptile review on the latest expanded diff is file-count limited, although the commitperclip review gate passed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool use enabled in a local repository workspace. Context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Commitperclip review gate is green; direct Greptile review is file-count limited on the latest expanded diff - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef1422c23e |
fix(cursor): coalesce streamed assistant text into prose blocks (#8544)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runs stream their transcripts through per-adapter stdout
parsers into the chat/run transcript UI
(`ui/src/adapters/transcript.ts`)
> - The Cursor CLI (local) streams assistant text as many small `text`
events (often a token or word each), and the parser emitted one
assistant entry per event and trimmed each
> - As a result the chat rendered one bubble per token ("every line a
new token") and dropped inter-token whitespace, making Cursor runs hard
to read
> - The render layer already coalesces consecutive `delta` entries
(`appendTranscriptEntry`), but the Cursor parser never tagged streamed
text as a delta
> - This pull request tags streamed `text` as a delta (without trimming)
so the existing render-time coalescer merges them into one assistant
block, while a `tool_call`/`tool_result` between deltas still breaks the
run
> - The benefit is readable Cursor transcripts with correct spacing and
preserved tool boundaries, with no change to the canonical event stream
(raw view unaffected)
## Linked Issues or Issue Description
No existing public issue — describing the bug inline (per
`.github/ISSUE_TEMPLATE/bug_report.yml`):
**What happened**
In the chat/run transcript, Cursor (local) assistant messages render as
one bubble per token/word, and inter-token spaces are dropped, making
the transcript unreadable. Root cause:
`packages/adapters/cursor-local/src/ui/parse-stdout.ts` (`type: "text"`
branch) emitted `{ kind: "assistant" }` per streamed `text` event
without `delta: true` and trimmed each, so the render-time coalescer
(`ui/src/adapters/transcript.ts`) never merged them and whitespace was
lost.
**Expected behavior**
Streamed assistant text should render as a single contiguous prose
block, with tool calls preserved as boundaries between blocks.
**Steps to reproduce**
1. Run a Cursor (local) agent that streams a multi-word assistant
message.
2. Open the run transcript in the chat UI.
3. Observe each streamed token/word rendered as its own bubble, with
inter-token spaces missing.
**Paperclip version**
Reproduced on current `master` (cutover base `e68188c43`).
**Deployment mode**
Self-hosted, `cursor_local` adapter.
## What Changed
- `packages/adapters/cursor-local/src/ui/parse-stdout.ts`: tag streamed
`text` events as `{ kind: "assistant", delta: true }` and stop trimming,
so the existing `appendTranscriptEntry` coalescer merges consecutive
deltas into one block.
- `ui/src/adapters/cursor-coalescing.test.ts` (new): dual-shape golden
fixtures (Cursor local + cloud) exercising the full render-time
projection via `buildTranscript`.
## Verification
- `pnpm --filter @paperclipai/ui exec vitest run
src/adapters/cursor-coalescing.test.ts src/adapters/transcript.test.ts`
→ **10/10 pass**.
- `pnpm --filter @paperclipai/ui --filter
@paperclipai/adapter-cursor-local typecheck` → **green**.
- The golden fixtures assert the run `text → tool_call → tool_result →
text → consolidated final` renders as exactly **two prose blocks with
the tool between them**, **no duplication** of the consolidated final,
and **inter-token whitespace preserved** across coalesced deltas.
## Risks
- **Low risk.** Pure classification at parse time; the canonical event
stream and the raw view are unchanged — only the "nice" render-time
projection changes. The coalescing logic (`appendTranscriptEntry`) is
pre-existing and already covered by tests. No schema, migration, or
behavioral change outside transcript rendering.
## Model Used
- **Claude Opus 4.8** (Anthropic), extended/high reasoning mode, driven
via the Cursor agent with tool use + code execution. Diagnosis and
fixtures grounded in the repo's actual parser/render code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change
(`fix/cursor-transcript-coalescing`) and contains no internal Paperclip
ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (N/A —
no documented behavior changes)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (pending CI run)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending review)
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Sebastian Heyneman <sebastian@joinnova.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
||
|
|
3d7a4f12f9 |
fix(claude-local): use direct 4.5 model aliases (#3800)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The `claude_local` adapter is the control-plane surface that turns Paperclip agent settings into concrete Claude CLI invocations. > - Issue #3777 reports that selecting Claude Haiku 4.5 from the Paperclip UI causes Claude Code to request `claude-haiku-4-5-20251001`, which some enterprise-scoped keys are not allowed to access. > - In the current adapter model list, the direct Anthropic 4.5 entries are version-pinned IDs, unlike the 4.6 entries which already use the stable short aliases. > - This pull request switches the direct 4.5 `claude_local` model options to the short aliases Claude Code account access expects, while leaving Bedrock-native IDs unchanged in the Bedrock-specific model list. > - The benefit is that selecting Claude Sonnet 4.5 or Claude Haiku 4.5 from Paperclip no longer forces the CLI onto a more restrictive versioned direct model name. ## What Changed - Updated the direct `claude_local` adapter model list to use `claude-sonnet-4-5` instead of `claude-sonnet-4-5-20250929`. - Updated the direct `claude_local` adapter model list to use `claude-haiku-4-5` instead of `claude-haiku-4-5-20251001`. - Left the Bedrock-specific model IDs in `src/server/models.ts` unchanged so AWS Bedrock routing still uses region-qualified native IDs. ## Verification - `pnpm install --frozen-lockfile` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - Confirmed the only source change is `packages/adapters/claude-local/src/index.ts`. ## Risks - Low risk. This only changes the direct `claude_local` model IDs advertised by Paperclip for the two 4.5 options. - Bedrock behavior is unchanged because the Bedrock-native identifiers are defined separately and were not modified in this PR. - Existing 4.6 model options are untouched. ## Model Used - OpenAI Codex on a GPT-5-class coding model with terminal tool use and local code execution in this workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge Fixes #3777 |
||
|
|
a27e5ad002 |
Add ephemeral sandbox runtime status plumbing (#8593)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandboxed agent runs can spend meaningful time preparing a remote workspace before the agent transcript shows useful output. > - Operators need short, current progress text for those setup phases, but that text should not become durable run history. > - The existing live-run websocket path already carries run updates to the UI, so the backend can reuse that channel instead of adding polling. > - This pull request adds an ephemeral runtime-progress contract, a process-local status store, and heartbeat integration for sandbox-managed runs. > - The benefit is a clearer active-run experience without database migrations or persistent progress rows. ## Linked Issues or Issue Description Refs #248 No exact public GitHub issue was found for this status-message plumbing. The underlying problem is that active sandboxed runs currently have setup phases, such as workspace sync and restore, where the operator cannot see concise current progress through the live run state. This PR addresses that gap for the backend/runtime layer while keeping progress messages ephemeral. GitHub search performed for related or duplicate work: `sandbox runtime status`, `sandbox restore index`, and `runtime progress`. No direct duplicate PR was found. ## What Changed - Added shared runtime-progress types and the `heartbeat.run.progress` live event type. - Added a process-local heartbeat run runtime-status store with TTL, bounded/redacted messages, and terminal cleanup. - Threaded runtime progress callbacks through heartbeat execution and active/live run serialization. - Emitted sandbox-managed runtime phase updates for sync, adapter startup, restore/export, and finalization paths. - Added backend and adapter-utils tests for ephemeral status behavior, terminal cleanup, live serialization, and sandbox progress callbacks. ## Verification - `pnpm install --frozen-lockfile` - Local PII scan before push: high-confidence secret patterns, internal issue links, local user paths, and private URL patterns checked across all three split diffs; no real secrets or internal links found. The only secret-like text is an intentional fake test fixture (`sk-test-secret`). - `git diff --check origin/master..feat/sandbox-runtime-status` - `pnpm exec vitest run server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/agent-live-run-routes.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 4 files, 23 tests passed. - `pnpm run typecheck` passed on both top stacks that include this branch: `feat/sandbox-status-ui` and `fix/sandbox-restore-index-sync`. - `pnpm run build` passed on both top stacks that include this branch; Vite reported existing CSS `::highlight` and chunk-size warnings. - `pnpm run test:run` was attempted on `fix/sandbox-restore-index-sync`; it failed in two unrelated broad-suite tests. One depends on this host's Git default branch behavior, and one depends on local Claude model-discovery environment. The changed focused suites above pass. ## Risks - Runtime progress is process-local by design, so status disappears after TTL, terminal cleanup, or server restart. - Clients that do not consume `heartbeat.run.progress` simply keep existing behavior. - Message redaction is intentionally generic; overly specific phase details should stay out of runtime-progress payloads. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent, with shell/tool execution in a local worktree. Exact context-window metadata is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip CTO <cto@paperclip.local> Co-authored-by: Paperclip CTO <noreply@paperclip.ing> |
||
|
|
51ffbb380f |
Exclude transient Codex home dirs from sandbox sync (#8581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters can run agents against sandboxed execution targets by syncing the workspace and selected runtime assets into the sandbox. > - The Codex local adapter includes a managed Codex home asset so sandboxed Codex runs can use the expected auth, config, skills, and session state. > - That asset follows symlinks, which is useful for real Codex home content but unsafe for transient launcher directories. > - Transient `tmp` and `.tmp` directories can contain symlinks to large host binaries, so the sandbox archive can inline large executable targets instead of just the small home directory content. > - This pull request excludes transient Codex home directories from the sandbox home asset while preserving the required Codex home files. > - The benefit is a much smaller and more predictable sandbox setup upload without changing the runtime files Codex actually needs. ## Linked Issues or Issue Description No public issue was found for this exact sandbox archive-size bug. Bug description: - What happened: sandboxed `codex_local` runs sync the managed Codex home as a `home` asset with `followSymlinks` enabled. If transient Codex home dirs such as `tmp` or `.tmp` contain symlinks to a large host binary, the archive can inline that binary and make `Syncing home to sandbox` much larger than the managed home directory itself. - Expected behavior: sandbox setup should include the Codex home files needed for auth, config, skills, and session continuity, but should not archive transient launcher scratch directories. - Reproduction shape: create a managed Codex home with normal auth/config/skills files and a `tmp/arg0` or `.tmp` symlink to a large host executable, then start a sandboxed `codex_local` run. The home asset archive grows by the symlink target size. - Version/commit: observed on local `master` before this change. - Related public context: #5028 covers a different managed Codex home reliability issue around stale auth files; this PR addresses sandbox archive bloat from transient symlink targets. ## What Changed - Excluded `tmp` and `.tmp` from the Codex `home` asset that is uploaded for sandboxed runs. - Added regression coverage proving transient symlinked home dirs are excluded from the tar while required auth/config/skills files remain included. - Kept `followSymlinks` behavior for the rest of the Codex home asset so existing non-transient symlink behavior is preserved. ## Verification - `git diff --check` - Local PII/secret pattern scan over the committed diff - `pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` ## Risks Low risk. The exclusion is limited to transient Codex home scratch directories, and the regression test verifies the files needed in the sandbox are still archived. The main compatibility risk is if a user intentionally placed required persistent Codex state under `tmp` or `.tmp`; those paths are treated as volatile scratch space by this change. ## Model Used OpenAI Codex coding agent based on GPT-5, with shell, git, and GitHub CLI tool use. Exact hosted model build and context-window size were not exposed in the local adapter runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b1332b61c |
fix(openclaw-gateway): drop root paperclip params (#4416)
Fixes #5997, fixes #4081, fixes #4723, fixes #6625, fixes #3923 Refs #6606 — this PR removes the rejected root `paperclip` field, but #6606 also requires the protocol v3→v4 bump, which is out of scope here; referencing rather than closing it. ## Thinking Path > - Paperclip is the control plane that wakes and coordinates agent workers across company-scoped execution flows. > - The `openclaw_gateway` adapter is part of that wake path, so its outbound payload contract has to match the gateway's validated `agent` schema. > - `master` currently reintroduces a previously fixed regression by sending a top-level `paperclip` property in `agentParams` (see #3923, which reverts the original fix in #626). > - The gateway rejects unknown root params, which means OpenClaw wakes fail before the remote agent can start work. > - The actual wake context already rides in the generated `message`, so the extra root property is both redundant and harmful. > - This pull request removes that leaked root property, adds a focused regression test around param construction, and updates affected server expectations/docs to the supported contract. > - The benefit is that OpenClaw Gateway agents wake successfully again without losing inline wake context. ## What Changed - Removed the top-level `paperclip` field from OpenClaw Gateway `agentParams` and extracted `buildAgentParams()` so the contract is easy to test. - Added a package-level regression test that proves `payloadTemplate.paperclip` is stripped while explicit `agentId`/`timeout` behavior stays intact. - Updated server tests that inspect OpenClaw Gateway payloads to assert wake data is delivered in `message` instead of a rejected root field. - Updated the adapter configuration docs to state that wake context is embedded in the generated message text, not sent as a top-level param. ### Rebase onto current `master` (conflict resolution) This branch was opened against an older `master`; re-merging current `master` required: - Resolving conflicts in `execute.ts` — `master` hoisted `configuredAgentId` and moved the agentId/timeout precedence inline; this PR keeps the `buildAgentParams()` extraction that strips the gateway-rejected root `paperclip`. - Updating tests `master` added **after** this branch's base that assert the old root-`paperclip` contract. These suites use the OpenClaw gateway adapter purely as a delivery harness (`adapterType: "openclaw_gateway"` + a mock gateway) and observe wake content via the gateway payload, so dropping the root field requires them to read wake context from `message` instead: - `server/src/__tests__/heartbeat-comment-wake-batching.test.ts` - `server/src/__tests__/low-trust-red-team-routes.test.ts` (redaction guarantees preserved — sanitized body + `expectNoCanary` on the raw canary) - Replaced brittle JSON-substring assertions (flagged by Greptile) with a shared `parseWakePayloadFromMessage()` helper + `toMatchObject`, robust to serialization/key-order changes. The strict contract is confirmed upstream: OpenClaw's `AgentParamsSchema` is `Type.Object(..., { additionalProperties: false })` with no `paperclip` field, so a root `paperclip` is rejected (`invalid agent params: at root: unexpected property 'paperclip'`). ## Verification - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` — clean - `pnpm --filter @paperclipai/server typecheck` — clean - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/openclaw-gateway-adapter.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts` — 26 passed (7 + 11 + 8) - `packages/adapters/openclaw-gateway/src/server/execute.test.ts` — 6 passed (run via a local temp vitest config because the root `vitest.config.ts` does not include this package) ## Risks - Low risk: this narrows the outbound payload to the gateway-supported contract and keeps wake context in the already-supported `message` channel. - Any downstream consumer that incorrectly depended on a top-level `paperclip` field from the gateway mock payloads would need to follow the supported `message` contract instead. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent via API authored the original change; exact underlying model ID and context window were not exposed in that environment. - Rebase/conflict resolution and the test-assertion migration were done with Claude Code (Claude Opus 4.8, 1M context). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked the related issues above - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: serenakeyitan via breeze-runner <serenakeyitan@users.noreply.github.com> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b86d60b802 |
build(deps): bump @cursor/sdk from 1.0.18 to 1.0.19 (#8462)
Bumps [@cursor/sdk](https://github.com/cursor/cursor) from 1.0.18 to 1.0.19. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/cursor/cursor/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
957f30cd72 |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.47.0 to 0.48.0 (#8469)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.47.0 to 0.48.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.48.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.47.0...v0.48.0">0.48.0</a> (2026-06-19)</h2> <h3>Features</h3> <ul> <li>Agent selection dropdown in config options (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/794">#794</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/5729c471f92e1d2a191eb2a6d18c2be47821e9ec">5729c47</a>)</li> <li><strong>deps:</strong> bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/787">#787</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ad3b5fe74527f83964695a1e8056d010b981fc79">ad3b5fe</a>)</li> <li>Update to claude-agent-sdk 0.3.183 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/791">#791</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/744b2d41128f67091d4136283594c0db11ea4db5">744b2d4</a>)</li> <li>Update to new ACP SDK patterns (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/790">#790</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2554c7bf980472760a6e7810b826f030d2c3af25">2554c7b</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>duplicate assistant text when turn activates mid-message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/789">#789</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1c80bf8e56a9279dc799e7bbdcae87241e99c18b">1c80bf8</a>)</li> <li>Skip empty thinking chunks (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/793">#793</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15fdf26fc7d3a6c51e89d3e56fc91dd00cb3d7ae">15fdf26</a>)</li> <li>surface Bash tool image output instead of dropping it (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/617">#617</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a759e64ef6d9b7c5bfe9a6e6b182db835f8ed3b6">a759e64</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.47.0...v0.48.0">0.48.0</a> (2026-06-19)</h2> <h3>Features</h3> <ul> <li>Agent selection dropdown in config options (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/794">#794</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/5729c471f92e1d2a191eb2a6d18c2be47821e9ec">5729c47</a>)</li> <li><strong>deps:</strong> bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/787">#787</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ad3b5fe74527f83964695a1e8056d010b981fc79">ad3b5fe</a>)</li> <li>Update to claude-agent-sdk 0.3.183 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/791">#791</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/744b2d41128f67091d4136283594c0db11ea4db5">744b2d4</a>)</li> <li>Update to new ACP SDK patterns (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/790">#790</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2554c7bf980472760a6e7810b826f030d2c3af25">2554c7b</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>duplicate assistant text when turn activates mid-message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/789">#789</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1c80bf8e56a9279dc799e7bbdcae87241e99c18b">1c80bf8</a>)</li> <li>Skip empty thinking chunks (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/793">#793</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15fdf26fc7d3a6c51e89d3e56fc91dd00cb3d7ae">15fdf26</a>)</li> <li>surface Bash tool image output instead of dropping it (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/617">#617</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a759e64ef6d9b7c5bfe9a6e6b182db835f8ed3b6">a759e64</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b4b7f561e2bc495132e624c9256b6e3cb9d9b477"><code>b4b7f56</code></a> chore(main): release 0.48.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/788">#788</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/41cde99fd7685e1b5b74f589bac9a68f3b79778f"><code>41cde99</code></a> test: isolate CLAUDE_CONFIG_DIR in settings tests to stop user-tier leak (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/769">#769</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/5729c471f92e1d2a191eb2a6d18c2be47821e9ec"><code>5729c47</code></a> feat: Agent selection dropdown in config options (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/794">#794</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a759e64ef6d9b7c5bfe9a6e6b182db835f8ed3b6"><code>a759e64</code></a> fix: surface Bash tool image output instead of dropping it (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/617">#617</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15fdf26fc7d3a6c51e89d3e56fc91dd00cb3d7ae"><code>15fdf26</code></a> fix: Skip empty thinking chunks (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/793">#793</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/9f38cb67dfad31d8a1265c2d2449c16c1b03ff8b"><code>9f38cb6</code></a> test: Improve tmp dirs in tests (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/792">#792</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/744b2d41128f67091d4136283594c0db11ea4db5"><code>744b2d4</code></a> feat: Update to claude-agent-sdk 0.3.183 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/791">#791</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/2554c7bf980472760a6e7810b826f030d2c3af25"><code>2554c7b</code></a> feat: Update to new ACP SDK patterns (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/790">#790</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1c80bf8e56a9279dc799e7bbdcae87241e99c18b"><code>1c80bf8</code></a> fix: duplicate assistant text when turn activates mid-message (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/789">#789</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ad3b5fe74527f83964695a1e8056d010b981fc79"><code>ad3b5fe</code></a> feat(deps): bump the minor group with 11 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/787">#787</a>)</li> <li>See full diff in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.47.0...v0.48.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2e2da3bc2f |
fix: isolate remote Claude config from local credentials (#7676)
Fix remote Claude sandbox config isolation and permissions handling.\n\nPR: https://github.com/paperclipai/paperclip/pull/7676 |
||
|
|
33353ce62b |
feat(skills): remove bundled paperclip-dev skill and retire required skill attribute (#7029)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local adapters (Claude, Codex, Cursor, Gemini, Grok, OpenCode, Pi, ACPX) ship bundled "skills" — opinionated Markdown prompt bundles materialized into the agent's runtime > - One of those bundled skills, `paperclip-dev`, existed to let agents develop Paperclip itself; it has now moved to its own external repo and no longer belongs in the core tree > - The adapter skill model also carried a `required` / `requiredReason` attribute plus a `paperclip_required` `AdapterSkillOrigin` variant, all of which only existed to mark bundled skills as non-optional in the UI and adapter sync logic > - With `paperclip-dev` gone, no bundled skill is "required" anymore, and the type / runtime surface for `required` is dead weight — but it is computed at request time and never persisted, so a clean removal is safe (no compatibility shim needed) > - This pull request deletes `skills/paperclip-dev/` and removes every trace of the `required` / `requiredReason` field and the `paperclip_required` origin across shared types, validators, adapter-utils, all eight local adapters, server routes, the company-skills service, the UI, the storybook fixtures, and the test suite > - The benefit is a smaller, simpler adapter-skill surface: one origin (`company_managed`) for managed bundled skills, `resolvePaperclipDesiredSkillNames` collapses to "just the configured desired set", and the AgentDetail skills tab no longer renders a "Required by Paperclip" section that no longer applies ## Linked Issues or Issue Description <!-- No existing public GitHub issue; describing the underlying work inline (feature_request template fields). --> **Summary** Remove the bundled `paperclip-dev` skill (now maintained in its own external repo) and retire the `required` / `requiredReason` skill attribute and the `paperclip_required` skill origin, which only existed to support it. **Problem or motivation** `paperclip-dev` is the only bundled skill that was ever marked "required". Now that it lives in a separate repository, shipping it inside the core tree is wrong, and the entire `required` surface (a type field, a validator field, a synthesized `paperclip_required` origin, UI "Required by Paperclip" section, and required-skill merging in the desired-skills calculation) becomes dead weight. The `required` value is computed at request time and never persisted, so it can be removed cleanly without a migration or compatibility shim. **Proposed solution** Delete `skills/paperclip-dev/`, drop the `required` / `requiredReason` fields and `paperclip_required` origin everywhere they are produced or consumed, collapse managed-skill origin to a single `company_managed` value, and simplify `resolvePaperclipDesiredSkillNames` to return only the configured desired set. **Alternatives considered** Keeping the `required` attribute as a no-op for forward compatibility — rejected because it is request-time only (nothing persists it), so leaving it in place is pure dead surface area with no callers. **Roadmap alignment** Internal cleanup / dead-code removal that simplifies the adapter-skill surface; it does not introduce or duplicate any planned core feature in ROADMAP.md. ## What Changed - Deleted bundled `skills/paperclip-dev/` (moved to a separate repo). - Dropped `required`, `requiredReason`, and the `paperclip_required` origin from `packages/shared/src/types/adapter-skills.ts`, `packages/shared/src/validators/adapter-skills.ts`, and `packages/adapter-utils/src/types.ts`. - In `packages/adapter-utils/src/server-utils.ts`: removed `readSkillRequired()`; dropped `required`/`requiredReason` from `listPaperclipSkillEntries()`, `normalizeConfiguredPaperclipRuntimeSkills()`, `buildPersistentSkillSnapshot()`, and `PaperclipSkillEntry`; collapsed `buildManagedSkillOrigin()` to always return `company_managed`; simplified `resolvePaperclipDesiredSkillNames()` to return only the configured desired set (signature preserved so adapter call sites are untouched). - Walked all eight local adapters (`acpx-local`, `claude-local`, `codex-local`, `cursor-local`, `gemini-local`, `grok-local`, `opencode-local`, `pi-local`) and removed every remaining `requiredReason` / `paperclip_required` reference. - `server/src/services/company-skills.ts`: dropped the `required = sourceKind === "paperclip_bundled"` synthesis when listing runtime skill entries. - `server/src/routes/agents.ts`: removed required-skill merging from the desired-skills calculation in the persist-config path and the unsupported-snapshot path (keeping the current version-aware `desiredSkillEntries` structure). - `ui/src/pages/AgentDetail.tsx`: dropped required-based filters, the required tooltip, and the entire "Required by Paperclip" section from the agent skills tab; storybook fixtures in `ui/storybook/stories/acpx-local.stories.tsx` cleaned up to match. - Tests: deleted the `required: false` case in `paperclip-skill-utils.test.ts` and the "keeps required bundled skills installed" case in every `*-local-skill-sync.test.ts`; `acpx-local-execute.test.ts`, `cursor-local-execute.test.ts`, `cursor-local-skill-sync.test.ts`, `agent-skills-routes.test.ts`, and `packages/adapter-utils/src/server-utils.test.ts` were updated to drop removed fields and map `origin: "paperclip_required"` → `"company_managed"`. - `server/src/adapters/registry.ts`: two `as unknown as ServerAdapterModule["..."]` casts on `hermesListSkills` / `hermesSyncSkills` (matching the existing `executeHermesLocal` pattern). `hermes-paperclip-adapter@0.2.0` still depends on the published `@paperclipai/adapter-utils` which keeps the retired `paperclip_required` variant; the cast bridges the workspace-vs-published type mismatch at the registry seam and can drop once hermes upgrades. ## Verification Run from the workspace root: ```sh grep -rn "skills/paperclip-dev" . grep -rn "paperclip_required" --include="*.ts" --include="*.tsx" . grep -rn "requiredReason" --include="*.ts" --include="*.tsx" . pnpm -w typecheck pnpm --filter @paperclipai/server exec vitest run paperclip-skill-utils pnpm --filter @paperclipai/server exec vitest run skill-sync ``` The first three greps return only the explanatory comment in `server/src/adapters/registry.ts` (no live `paperclip_required` / `requiredReason` usage) and zero `skills/paperclip-dev` source hits. Locally: - `pnpm -w typecheck` → all packages this PR touches pass (adapter-utils, shared, server, ui, cli, and the cursor/gemini/opencode/pi adapters). - Affected vitest suites pass: `paperclip-skill-utils`, `server-utils`, all eight `*-local-skill-sync`, `agent-skills-routes`, and the `acpx`/`cursor`/`pi` execute suites. ## Risks - Behavioral shift in the agent skills UI: the "Required by Paperclip" section disappears. No bundled skill is required anymore, so this only affects environments that previously surfaced `paperclip-dev` as a forced-on row; those installs will see the skill move into the regular "company-managed" list (and be uninstalled on next sync unless explicitly listed as desired). - Existing agents may still have the string `"paperclip-dev"` in their persisted `desiredSkills`. That entry is inert (no source for it to install from); a one-time DB cleanup is out of scope. Low risk. - Hermes adapter type bridge: two casts in `registry.ts` paper over a type-only divergence between the workspace `@paperclipai/adapter-utils` and the published version still pinned by `hermes-paperclip-adapter@0.2.0`. Runtime behavior is unaffected because the retired `paperclip_required` value is no longer produced by anything in this tree. The casts can be removed once hermes upgrades its dependency. ## Model Used - Provider: Anthropic - Model: Claude Opus 4.7 (`claude-opus-4-7`) - Capability: agent tool use via Paperclip's `claude_local` adapter ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7aa212296e |
codex_local adapter output inactivity monitor (#5017)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - The `codex_local` adapter runs `codex exec` as a child process and streams JSONL events from its stdout > - NEE-79 caught a real-world `codex_local` orphan: codex sat in `read()` on stdin for 1h+, no JSON events emitted past startup, no LLM call in flight. The adapter had no inactivity timer; the only safety net was the platform-level 1h silent-run detector > - This is precisely the failure shape NEE-80 scoped: an adapter-detected fault that should be killed *by the adapter* and surfaced as `failed`, not waited out for an hour > - This pull request adds an output-inactivity watchdog inside the codex-local adapter that resets on every parsed JSONL event from stdout, kills the child via SIGTERM → 5s grace → SIGKILL when it fires, and resolves the run with a structured watchdog failure > - The benefit is that NEE-79-class hangs shorter than 1h stop reaching the platform-level safety net — they fail fast and visibly at the adapter layer, with diagnostic logs that don't require host shell access ## Linked Issues or Issue Description No existing GitHub issue covers this; describing the underlying bug in-PR per the bug report template (`.github/ISSUE_TEMPLATE/bug_report.yml`): - **What happened?** A `codex_local` run sat in `read()` on stdin for over an hour: codex emitted its startup JSONL events, then nothing — no further events, no LLM call in flight, no exit. The adapter kept the child alive indefinitely; the run was only reaped by the platform-level 1h silent-run safety net, an hour after it had effectively died. - **Expected behavior:** The adapter should detect that its child has stopped producing output long before the platform-level safety net, kill it, and surface the run as `failed` with a diagnostic that explains what happened. - **Steps to reproduce:** Run any `codex_local` issue where `codex exec` hangs after startup (e.g. codex blocks reading stdin and never emits another JSONL event). Observe the run stays alive until the 1h platform safety net fires. - **Paperclip version or commit:** reproduced on master prior to this branch. - **Agent adapter(s) involved:** codex_local. Related PRs found while searching for duplicates (none implement an event-aware inactivity watchdog inside `codex_local`): - #4004 — generic idle + wall watchdogs for `runChildProcess` in adapter-utils; complementary, operates below the JSONL parse layer and is not codex-event-aware - #4742 — fail-fast on codex-local *startup* hang; this PR covers the post-startup hang class - #6861 — zombie-run termination for codex-local; reaping after exit, not inactivity detection - #7811 — the analogous output-idle timeout for grok-local ## What Changed - `packages/adapters/codex-local/src/server/watchdog.ts` *(new)* — pure watchdog primitive with injectable timers/clock: `resolveCodexInactivityTimeout`, `createCodexInactivityWatchdog`, `formatWatchdogErrorMessage`. Default `7 * 60_000` ms; honors `null` as the disabled escape hatch - `packages/adapters/codex-local/src/server/execute.ts` — `runAttempt` now wraps `onSpawn` to capture `pid`/`processGroupId`, feeds stdout chunks through `noteStdoutChunk`, and on watchdog fire sends SIGTERM to the process group, schedules SIGKILL after 5s, and returns an `AdapterExecutionResult` with `exitCode: null`, `signal: SIGTERM|SIGKILL`, `errorMessage: "watchdog: no codex output for {N}m {S}s"`, `errorCode: "codex_output_inactivity_watchdog"`. With `errorMessage` set and `timedOut: false`, `heartbeat.ts:5860` maps the run to `outcome === "failed"` (not `cancelled`) - `packages/adapters/codex-local/src/index.ts` — `agentConfigurationDoc` documents `outputInactivityTimeoutMs` (number ms; `null` disables; non-positive falls back to default with a warning log at spawn) - `packages/adapters/codex-local/src/server/watchdog.test.ts` *(new)* — 13 tests covering acceptance criteria 2 and 3 plus supporting cases: fires after silence, no-fire across 12× (threshold − 1s) cycles, multi-event chunks, non-JSON ignoring, single-fire idempotency, formatter shape, full resolution table, `null` → disabled - `packages/adapters/codex-local/src/server/watchdog.integration.test.ts` *(new)* — real Node subprocess that prints one JSONL event then sleeps; `runChildProcess` reaps it within `threshold + 6s`; signal is SIGTERM or SIGKILL; `parsedEventCount === 1` (acceptance criteria 1 and 4) ## Verification ``` $ pnpm --filter @paperclipai/adapter-codex-local typecheck > tsc --noEmit # clean $ pnpm --filter @paperclipai/adapter-codex-local exec vitest run ✓ src/server/quota-spawn-error.test.ts (1 test) ✓ src/server/codex-home.test.ts (3 tests) ✓ src/server/codex-args.test.ts (3 tests) ✓ src/server/watchdog.test.ts (13 tests) ✓ src/server/parse.test.ts (9 tests) ✓ src/server/execute.remote.test.ts (4 tests) ✓ src/ui/parse-stdout.test.ts (3 tests) ✓ src/ui/build-config.test.ts (1 test) ✓ src/server/watchdog.integration.test.ts (1 test) Test Files 9 passed (9) Tests 38 passed (38) ``` Acceptance criteria check: 1. ✅ Simulated child emits one event then sleeps → killed at threshold; result `errorMessage` matches `watchdog: no codex output for {N}m {S}s` (`watchdog.integration.test.ts`) 2. ✅ Child emits events every (threshold − 1s) → not killed (`watchdog.test.ts: "does not fire when events arrive every (threshold - 1s)"`) 3. ✅ `outputInactivityTimeoutMs: null` disables the watchdog (`resolveCodexInactivityTimeout` returns `disabled`; `execute.ts` skips watchdog construction and logs a startup warning) 4. ✅ Real subprocess reaped well within `threshold + 6s` — 250 ms threshold, 290 ms wall clock in CI 5. ✅ Adapter-level fault → `outcome === "failed"` per `heartbeat.ts:5860`. NEE-79-class hangs <1h get caught at the adapter, not the platform-level safety net Post-rebase verification (head `2a8703fab`, rebased onto master `69a368ed5`): ``` $ pnpm --filter @paperclipai/adapter-codex-local typecheck # clean $ pnpm --filter @paperclipai/adapter-codex-local exec vitest run Test Files 11 passed (11) Tests 64 passed (64) ``` ## Risks - Low risk. Behind a default-on watchdog that only fires after 7m of zero parsed JSON events. Operators can disable it with `outputInactivityTimeoutMs: null` for known-slow tasks - The kill path reuses the same `process.kill(-pgid, signal)` pattern that `runChildProcess` already uses for its terminal-result cleanup, so signal semantics match the existing code path - `timedOut: false` is preserved on watchdog fire — the platform-level timeout outcome is unchanged, only the `failed`-vs-success classification flips. No behavioral shift for already-failing runs - Sandbox/SSH execution targets: the watchdog fires and emits the structured log, but the kill is best-effort because remote pids aren't owned by this process. The platform-level 1h safety net still applies. Out of scope for NEE-81 by design ## Model Used - Provider: Anthropic Claude - Model: claude-opus-4-7 - Mode: Claude Code (extended thinking, tool use) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI changes) - [x] I have updated relevant documentation to reflect my changes (`agentConfigurationDoc` for `outputInactivityTimeoutMs`) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (re-running on the rebased head) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (P2 addressed; review threads resolved) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Neeraj Kumar Singh <b.nirajkumarsingh@hotmail.com> Co-authored-by: Devin Foley <devin@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
67ca8287e8 |
fix(codex-local): seed managed auth into isolated homes (#8403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter isolates local Codex runs by assigning managed `CODEX_HOME` state per company and per agent > - PR #8272 tightened that isolation, but it left a gap: once `CODEX_HOME` became explicit, the adapter treated it like a user-managed override and skipped auth seeding > - That meant newly isolated agents could launch with no usable `auth.json`, hit OpenAI unauthenticated, and fail with `401 Missing bearer` > - Users who had already persisted one of those broken managed homes could remain stranded even after config changes unless Paperclip repaired the home itself > - This pull request teaches Paperclip to seed managed homes correctly, backfill already-stranded managed homes on startup, and reject credential-less managed homes before they reach the provider > - The benefit is that affected managed `codex_local` agents recover automatically after upgrade and restart, without manual `CODEX_HOME` surgery ## Linked Issues or Issue Description - Fixes #497 - Refs #5028 - Related PR: #8272 - Related PR: #8399 ## What Changed - Distinguished Paperclip-managed `CODEX_HOME` paths from genuine external overrides and always seeded auth into managed homes, even when `CODEX_HOME` is explicit in config. - Wrote API-key-backed `auth.json` files for managed homes when `OPENAI_API_KEY` is configured, otherwise symlinked the shared Codex auth for subscription/OAuth flows. - Added a startup reconciliation pass that backfills already-isolated managed homes created by the broken release so upgrade plus restart repairs stranded agents automatically. - Preserved previously resolved API-key auth when the stored `OPENAI_API_KEY` binding is secret-backed and startup cannot resolve the secret value directly. - Hardened the managed-home preflight to require a credential-bearing `auth.json`, not just file presence, and documented the recovery behavior. - Added regression tests covering managed-home seeding, fail-fast behavior, and server-side startup reconciliation. ## Verification ```bash pnpm exec vitest run packages/adapters/codex-local/src/server/codex-home.test.ts packages/adapters/codex-local/src/server/execute.auth.test.ts server/src/__tests__/codex-auth-reconciliation.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/server-startup-feedback-export.test.ts pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/server typecheck pnpm check:tokens pnpm exec vitest run server/src/__tests__/server-startup-feedback-export.test.ts ``` GitHub Actions `PR` workflow is green on latest head `e5b1e08d4`, including policy, typecheck, test shards, build, e2e, serialized server suites, and canary dry run. Greptile Review is green on latest head with 0 comments added. No UI changes. ## Risks - Startup reconciliation now mutates persisted managed Codex homes at boot. Risk is low because it only touches Paperclip-managed company/agent home paths and no-ops when a home already has usable auth. - Genuine external `CODEX_HOME` overrides remain intentionally self-managed, so those users still own repair steps inside their custom home. - Hosts with neither shared Codex auth nor an explicit per-agent API key now fail earlier with a clearer adapter error instead of surfacing a downstream `401`, which changes timing but not capability. ## Model Used - OpenAI Codex, GPT-5-based coding agent in a local Codex session; exact served model ID/context window were not exposed to the session. Tool use and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c0743482bc |
fix(claude-local): tolerate sandboxes whose Claude CLI lacks --effort (#8393)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The claude-local adapter launches the Claude CLI inside execution environments (local, SSH, and ephemeral sandboxes such as Daytona) > - Newer adapter code passes `--effort` to the CLI, but the Claude binary baked into some sandbox images is older and rejects it with `error: unknown option '--effort'`, so every run in those environments fails > - This needs addressing because the failure is environment-dependent and silent from the operator's perspective — the run just dies with a CLI usage error > - This pull request probes the in-sandbox CLI for `--effort` support once, caches the result per environment, and strips the flag (with a warning) when the CLI does not support it > - The benefit is that sandboxes with older Claude CLIs keep working instead of failing, with negligible probe overhead because the capability check is cached and reused across ephemeral leases ## Linked Issues or Issue Description No public GitHub issue exists. Describing the bug inline following the bug report template: ### What happened? Runs using the `claude-local` adapter inside certain sandbox/execution environments fail with `error: unknown option '--effort'`. The adapter unconditionally appends `--effort` to the Claude CLI invocation, but the Claude CLI version present in some sandbox base images predates that flag, so the process exits with a usage error and the run dies. ### Expected behavior The adapter should detect that the target environment's Claude CLI does not support `--effort` and degrade gracefully — drop the flag and emit a warning — rather than failing the run. ### Steps to reproduce 1. Configure an execution environment (e.g. a sandbox image) whose bundled Claude CLI is old enough to predate the `--effort` option. 2. Run any claude-local task that resolves to an effort level (so `--effort` is appended). 3. Observe the run fail immediately with `error: unknown option '--effort'`. ### Paperclip version or commit `master` at the time of this PR (branch forked from current `master`). ### Deployment mode Self-hosted / local instance using execution environments (reproducible with ephemeral sandbox providers such as Daytona where `reuseLease: false`). ### Agent adapter(s) involved Claude Code (`claude-local`). ## What Changed - Add a CLI capability probe (`cli-capabilities.ts`) that runs the target Claude binary's `--help` inside the execution environment to detect `--effort` support. - Strip `--effort` from the CLI args (emitting a warning) when the probe reports the flag is unsupported; keep it otherwise. - Cache probe results keyed by `sandbox:providerKey:environmentId:command` (no lease id) so the probe is reused across ephemeral leases — important for `reuseLease: false` sandbox configs like Daytona, which would otherwise re-probe on every run. - Conservative fallback: if the probe itself can't run/parse, assume the flag is supported (preserves prior behavior). ## Verification - `node_modules/.bin/vitest run src/__tests__/claude-local-execute.test.ts src/__tests__/claude-local-adapter-environment.test.ts` → **2 files, 30 tests passed**. - Regression test issues two `execute()` calls with distinct lease ids and asserts the in-sandbox `--help` probe runs exactly once (cache reuse across leases). - Added tests covering: flag stripped when unsupported, flag retained when supported, warning emitted, and conservative fallback when the probe fails. ## Risks Low risk. The change is additive and gated behind a probe with a conservative default (assume supported on probe failure), so existing environments that support `--effort` are unaffected. Worst case for an environment where the probe is unreliable is the prior behavior (flag passed through). ## Model Used Claude Opus 4.8 (claude-opus-4-8), extended thinking, with tool use / code execution via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI change) - [ ] I have updated relevant documentation to reflect my changes (N/A — no doc-facing behavior change) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
07e98d2b2c |
feat(adapter-utils): add observable sandbox sync progress (#8395)
## Thinking Path > - Paperclip is the control plane for running AI-agent companies, so long-running remote work needs to stay observable to human operators. > - Cloud / sandbox agents are an active roadmap area, and their workspace sync path is part of the runtime substrate every remote coding run depends on. > - In the sandbox and SSH execution-target flows, Paperclip logged that sync had started, then often went silent for the full transfer window. > - That made large remote syncs feel stalled and also hid a real performance problem in the command-managed sandbox upload path. > - The first part of this pull request threads a throttled progress-reporting surface through the adapter execution-target stack so sync and restore work can emit meaningful updates. > - The second part fixes the command-managed sandbox transport itself: it removes the old serial 32KB append bottleneck, but also falls back away from the single-stream path when a provider-backed sandbox runner cannot surface mid-flight stdin progress. > - The result is that sandbox and SSH transfers are both faster and more observable, including the live Daytona-style sandbox case that previously only emitted `0%` and `100%`. ## Linked Issues or Issue Description No public GitHub issue exists for this bug, so it is described inline below following the bug report template. ### What happened - Remote sandbox and SSH workspace syncs could spend a long time transferring data while only logging a start line (`Syncing workspace and runtime assets to sandbox environment`) and, at best, a terminal line. - In the command-managed sandbox path, the original upload implementation also paid a large performance cost by appending base64 data in many small sequential remote writes (thousands of serial 32KB round-trips on a large workspace). - After the initial transport rewrite, live provider-backed sandbox runs still only emitted `0%` and `100%` because the single-stream stdin RPC buffered progress until completion. ### Expected behavior - Long-running sandbox and SSH syncs should periodically report how much of the transfer is complete (a percentage and/or MB transferred) so an operator can tell the run is healthy and making progress rather than stuck. - The main sandbox upload path should not be artificially slow. - A transfer that fails partway should leave an explicit failure marker in the log rather than a dangling intermediate percentage. ### Steps to reproduce 1. Run an agent against a sandbox (command-managed) or SSH (remote-managed) execution target with a non-trivial workspace. 2. Watch the run log during the workspace/runtime asset sync phase. 3. Observe that the log shows the sync start line and then stays silent for the full transfer (live provider-backed sandbox runs only show `0%` then `100%`). ### Paperclip version or commit - Branch `PAPA-825-provide-status-updates-when-syncing-sandboxes` off `master`. ### Deployment mode - Self-hosted / local instance using sandbox (command-managed) and SSH (remote-managed) execution targets, including provider-backed sandbox runners. ## What Changed - Added shared throttled runtime progress reporting and threaded `onProgress` through the adapter execution-target surface and adapter `execute.ts` entrypoints. - Added sync and restore progress reporting for the command-managed sandbox path and the SSH/remote-managed path, including git import/export progress where totals are known. - Reworked command-managed sandbox transfer behavior so uploads use the faster single-stream path when appropriate, but fall back to chunked progress-emitting writes when the runner cannot expose mid-stream stdin progress. - Marked provider-backed environment sandbox runners as not supporting single-stream stdin progress so live sandbox runs emit meaningful intermediate updates instead of only `0%` and `100%`. - Emit an explicit terminal failure marker (`failed at NN% (x/y MB)`) when an SSH/tar transfer rejects, so a failed sync no longer leaves a dangling intermediate percentage in the log. - Run the SSH sync/restore size estimate (local directory walk / remote `du` probe) concurrently with the transfer instead of awaiting it before opening the pipe, so progress instrumentation no longer adds startup latency proportional to workspace file count. - Added and extended focused regression coverage for runtime progress throttling and the new failure marker, command-managed sandbox transfers, sandbox orchestration, SSH transfer progress, and environment execution-target wiring. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/runtime-progress.test.ts packages/adapter-utils/src/ssh-fixture.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` - `pnpm exec vitest run packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/environment-execution-target.test.ts` - `npx tsc --noEmit` for `packages/adapter-utils` ## Risks - The provider-backed sandbox fallback now prefers chunked command-managed writes when progress hooks are active, so small-to-medium uploads may trade some raw throughput for observable intermediate progress on runtimes that cannot surface true mid-stream stdin progress. - Progress percentages on tar-based transfers still depend on estimates in some cases, so operators may briefly see MB-only lines before the estimate resolves, then near-final clamping before the terminal `100%` line. - This PR changes shared execution-target behavior used by multiple adapters, so regressions would most likely appear in remote runtime setup/teardown flows rather than in a single adapter. ## Model Used - Initial implementation: OpenAI GPT-5.4 via Codex local agent (`codex_local`), high reasoning mode. - Observability follow-ups (failure marker, concurrent size estimate, added tests): Claude Opus 4.8 via Claude Code (`claude_local`). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb68198c42 |
feat(adapter-claude-local): emit errorCode claude_refusal on stop_reason refusal (#8314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run through adapters; `@paperclipai/adapter-claude-local` shells out to the Claude CLI and maps its result JSON into Paperclip's run outcome (`errorCode` / `errorFamily`) > - A Fable 5 policy refusal exits the CLI cleanly (`exitCode=0`, `is_error=false`) with `stop_reason: "refusal"`, which the adapter mapped to `errorCode: null` > - Paperclip therefore recorded the run as a silent success, so the agent's heartbeat stalled with no signal for operators and no hook to retry or alert > - This pull request adds a refusal detector to the adapter so a refusal resolves to a distinct `errorCode: "claude_refusal"` (`errorFamily: "model_refusal"`) > - The benefit is that policy refusals become observable: the server persists the code, operators can see it on the run, and downstream retry/alert policy can key off it ## Linked Issues or Issue Description No public GitHub issue exists. Describing the underlying bug in-PR, following `.github/ISSUE_TEMPLATE/bug_report.yml`: **What happened?** When the Claude CLI returns `stop_reason: "refusal"` (a model policy refusal), it still exits cleanly (`exitCode=0`, `is_error=false`). `@paperclipai/adapter-claude-local` keyed refusal detection off the failure flag during error-code resolution, so it returned `errorCode: null` — a refused run was indistinguishable from a successful one, recorded by Paperclip as a silent success with no signal to retry or alert. **Expected behavior** A refusal should resolve to a distinct, non-null error code (`errorCode: "claude_refusal"`, `errorFamily: "model_refusal"`) so the server can surface it to operators and downstream policy can react. **Steps to reproduce** 1. Run any agent on the `claude-local` adapter with a prompt the model refuses on policy grounds. 2. The Claude CLI exits `0` with `is_error=false` and `stop_reason: "refusal"`. 3. Observe the run resolves `errorCode: null` (pre-fix) instead of a refusal-specific code. **Paperclip version or commit** `adapter-claude-local` v2026.609.0 (`dist/server/execute.js` error-code resolution block); reproduced on `master`. **Deployment mode** Local / self-hosted. **Agent adapter(s) involved** `@paperclipai/adapter-claude-local` (Claude Code, local). ## What Changed - **`packages/adapters/claude-local/src/server/parse.ts`** — new `isClaudeRefusalResult()` helper, parallel to `isClaudeMaxTurnsResult()`. Detects `stop_reason` / `stopReason` / `error_code` / `errorCode` == `refusal` and `subtype` == `model_refusal`, case/whitespace tolerant. - **`packages/adapters/claude-local/src/server/execute.ts`** — compute the refusal flag independent of the `failed` flag (a refusal exits cleanly, so keying off `failed` would miss it); resolve `errorCode: "claude_refusal"` and `errorFamily: "model_refusal"`; surface `stopReason: "refusal"` in `resultJson`. - **`packages/adapter-utils/src/types.ts`** — widen `AdapterExecutionErrorFamily` with `"model_refusal"`. - **`packages/adapters/claude-local/src/server/index.ts`** — export the new helper. - **`packages/adapters/claude-local/src/server/parse.test.ts`** — 7 new unit tests. ## Verification ``` npx vitest run packages/adapters/claude-local # 35 passed (7 new) npx tsc --noEmit -p packages/adapter-utils # clean npx tsc --noEmit -p packages/adapters/claude-local # clean ``` Manual: a result JSON with `stop_reason: "refusal"` on a clean exit now resolves `errorCode: "claude_refusal"` and `errorFamily: "model_refusal"`; non-refusal results are unaffected. ## Risks Low risk. Additive only — it introduces a new error code on a path that previously returned `null`; no existing error code or success path changes. `claude_refusal` is deliberately **not** added to the transient-upstream retry set: a refusal is deterministic, so retrying the same prompt yields the same refusal. It surfaces as a distinct non-transient code for operators; downstream retry/alert policy can key off `claude_refusal` / `errorFamily: model_refusal` later. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`), extended-thinking mode, run via Claude Code with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no related PRs found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (N/A — no UI changes) - [x] I have updated relevant documentation to reflect my changes (N/A — internal adapter error-code addition; no user-facing docs) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (this push re-runs the gates; will confirm once green) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (prior run was 4/5 on a PR-description note now addressed; will confirm on re-run) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Full-Stack Engineer <noreply@paperclip.ing> |
||
|
|
323b6220cc |
Fix Gemini headless invocation stalls (#8368)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters turn Paperclip runs into unattended CLI invocations. > - The `gemini_local` adapter depends on Gemini CLI behaving non-interactively in headless worker sessions. > - When Gemini CLI falls back to browser-based auth, the process can stall at startup instead of producing stream-json output. > - The adapter should make headless intent explicit and turn missing auth into a classified, fast failure. > - This pull request hardens the Gemini child-process environment and parser classification around that failure mode. > - The benefit is that operators get actionable `gemini_auth_required` failures instead of silent hung runs. ## Linked Issues or Issue Description No exact public issue exists for this runtime stall. Related: #2344 covers a Gemini CLI adapter environment/auth probe failure. This PR addresses a different runtime path: unattended `gemini_local` executions that can stall when Gemini CLI attempts interactive browser auth in a headless session. Bug summary: - Adapter: `gemini_local` - Symptom: child Gemini CLI process starts but produces no stream-json output when auth requires interactive browser flow - Expected: unattended runs either produce stream-json output or fail quickly with a classified auth error - Actual: the run can hang at invocation until an external watchdog kills it - Scope: Gemini adapter process env, auth-required parsing, and adapter docs/tests Duplicate search performed: - `gh pr list --state all --search "gemini headless invocation stall"` found only this PR. - `gh pr list --state all --search "gemini NO_BROWSER NO_COLOR"` found only this PR. - `gh issue list --state all --search "gemini headless authentication"` found related issue #2344 but no exact duplicate. ## What Changed - Set a headless-safe Gemini invocation env at the final child-process boundary: `TERM=xterm-256color`, `COLORTERM=truecolor`, and `NO_BROWSER=1`. - Delete inherited `NO_COLOR` from the child env so Gemini CLI keeps color-capable terminal behavior when deciding whether it can run non-interactively. - Classify Gemini `FatalAuthenticationError: Manual authorization is required...` failures as `gemini_auth_required`. - Update Gemini adapter docs to describe the non-interactive `--prompt` path and headless env behavior. - Add/extend focused parser and remote execution tests for the auth-required and env invariants. ## Verification - `pnpm exec vitest run packages/adapters/gemini-local/src/server/parse.test.ts packages/adapters/gemini-local/src/server/execute.remote.test.ts` — 23 tests passed. - `pnpm --filter @paperclipai/adapter-gemini-local typecheck` — passed. - Local Gemini CLI probe confirmed `NO_BROWSER=1` turns the browser-auth stall into a fast auth error instead of an interactive wait. ## Risks Low risk. The change is limited to `gemini_local` invocation env, parsing, docs, and tests. It does not touch schemas, API routes, persisted data, or other adapters. Operational risk: environments that intentionally rely on `NO_COLOR` being inherited by Gemini child processes will no longer pass that variable through. That is intentional here because Gemini CLI auth/headless behavior should not depend on inherited color suppression. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex via Codex CLI, with shell/file-editing tool use for repository inspection, code edits, tests, and GitHub PR maintenance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge /cc @codex — please review. |
||
|
|
6f142a60ce |
build(deps-dev): bump @types/node from 22.19.11 to 22.19.21 (#7748)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 22.19.11 to 22.19.21. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
dd92c7d2dd |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.31.4 to 0.47.0 (#7744)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.31.4 to 0.47.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.47.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.46.0...v0.47.0">0.47.0</a> (2026-06-17)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.179 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/783">#783</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/59a098c2b530bbae034e9a2dfbd31f8b4ef2a4d0">59a098c</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Duplicate assistant messages in feed (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/785">#785</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/12d34e64e53564602ac1c38a30127e234c5c25ff">12d34e6</a>)</li> </ul> <h2>v0.46.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.45.1...v0.46.0">0.46.0</a> (2026-06-16)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.178 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/777">#777</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/58549ffe6a8b02ce59894e567407bd4299c11428">58549ff</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Better handle out of turn events (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/780">#780</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/4f273a20d870c9c69f71556b8e0519f1de30f285">4f273a2</a>)</li> <li>Forward option details in elicitation meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/779">#779</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b3640599ae685beecacd93e012d5bbc9dac716f7">b364059</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/764">#764</a></li> </ul> <h2>v0.45.1</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.45.0...v0.45.1">0.45.1</a> (2026-06-16)</h2> <h3>Bug Fixes</h3> <ul> <li>Fix terminal error printing as text instead of terminal output (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/776">#776</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/db6eaaf71484a321e47093ad65bcf8994943cb31">db6eaaf</a>)</li> <li>Scope custom answers per question (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/774">#774</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/d58004a34880e0a76833697319eb2a9efa6a43c7">d58004a</a>)</li> </ul> <h2>v0.45.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.44.0...v0.45.0">0.45.0</a> (2026-06-15)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> bump the minor group with 3 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/763">#763</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7de5e4bcca9bfea70593092060f82bc8abe33e0e">7de5e4b</a>)</li> <li><strong>deps:</strong> bump <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.177 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/771">#771</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1be5ca57ee772fe90e41126365dc4186a18ad257">1be5ca5</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>preserve ANTHROPIC_CUSTOM_MODEL_OPTION when availableModels is set (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/768">#768</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/cc2885f6a9993cf61e759c3c770015f94c218627">cc2885f</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.46.0...v0.47.0">0.47.0</a> (2026-06-17)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.179 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/783">#783</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/59a098c2b530bbae034e9a2dfbd31f8b4ef2a4d0">59a098c</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Duplicate assistant messages in feed (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/785">#785</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/12d34e64e53564602ac1c38a30127e234c5c25ff">12d34e6</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.45.1...v0.46.0">0.46.0</a> (2026-06-16)</h2> <h3>Features</h3> <ul> <li>Update to claude-agent-sdk 0.3.178 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/777">#777</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/58549ffe6a8b02ce59894e567407bd4299c11428">58549ff</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>Better handle out of turn events (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/780">#780</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/4f273a20d870c9c69f71556b8e0519f1de30f285">4f273a2</a>)</li> <li>Forward option details in elicitation meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/779">#779</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b3640599ae685beecacd93e012d5bbc9dac716f7">b364059</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/764">#764</a></li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.45.0...v0.45.1">0.45.1</a> (2026-06-16)</h2> <h3>Bug Fixes</h3> <ul> <li>Fix terminal error printing as text instead of terminal output (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/776">#776</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/db6eaaf71484a321e47093ad65bcf8994943cb31">db6eaaf</a>)</li> <li>Scope custom answers per question (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/774">#774</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/d58004a34880e0a76833697319eb2a9efa6a43c7">d58004a</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.44.0...v0.45.0">0.45.0</a> (2026-06-15)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> bump the minor group with 3 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/763">#763</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7de5e4bcca9bfea70593092060f82bc8abe33e0e">7de5e4b</a>)</li> <li><strong>deps:</strong> bump <code>@anthropic-ai/claude-agent-sdk</code> to 0.3.177 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/771">#771</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1be5ca57ee772fe90e41126365dc4186a18ad257">1be5ca5</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>preserve ANTHROPIC_CUSTOM_MODEL_OPTION when availableModels is set (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/768">#768</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/cc2885f6a9993cf61e759c3c770015f94c218627">cc2885f</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.43.0...v0.44.0">0.44.0</a> (2026-06-09)</h2> <h3>Features</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/794aa846844a2fe8a8574c2539e2c4107e9182d1"><code>794aa84</code></a> chore(main): release 0.47.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/784">#784</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/12d34e64e53564602ac1c38a30127e234c5c25ff"><code>12d34e6</code></a> fix: Duplicate assistant messages in feed (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/785">#785</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/59a098c2b530bbae034e9a2dfbd31f8b4ef2a4d0"><code>59a098c</code></a> feat: Update to claude-agent-sdk 0.3.179 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/783">#783</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/c27e6aec9059a920f9cd768f492e25934653b3ff"><code>c27e6ae</code></a> chore(main): release 0.46.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/778">#778</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/4f273a20d870c9c69f71556b8e0519f1de30f285"><code>4f273a2</code></a> fix: Better handle out of turn events (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/780">#780</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b3640599ae685beecacd93e012d5bbc9dac716f7"><code>b364059</code></a> fix: Forward option details in elicitation meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/779">#779</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/58549ffe6a8b02ce59894e567407bd4299c11428"><code>58549ff</code></a> feat: Update to claude-agent-sdk 0.3.178 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/777">#777</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/116f06eab178fe50857f7a4150deb7f5c4cfee28"><code>116f06e</code></a> chore(main): release 0.45.1 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/775">#775</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/db6eaaf71484a321e47093ad65bcf8994943cb31"><code>db6eaaf</code></a> fix: Fix terminal error printing as text instead of terminal output (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/776">#776</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/d58004a34880e0a76833697319eb2a9efa6a43c7"><code>d58004a</code></a> fix: Scope custom answers per question (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/774">#774</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.31.4...v0.47.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5320a44088 |
Guard codex_local agents from shared OpenAI key (#8272)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The `codex_local` adapter runs local Codex CLI processes and builds their environment from persisted agent config plus host process env. > - A host-level `OPENAI_API_KEY` or shared Codex auth home can silently make new agents spend through shared credentials. > - Existing agents can be repaired manually, but new and updated agents need a persistent guard at the agent configuration boundary. > - This pull request isolates new and updated `codex_local` agents with per-agent `CODEX_HOME` and an empty `OPENAI_API_KEY` override. > - The benefit is that future agent creation or adapter updates cannot silently fall back to shared OpenAI credentials. ## Linked Issues or Issue Description Paperclip work item: [ZOL-5477](/ZOL/issues/ZOL-5477). No matching GitHub issue exists, so the bug is described inline following `.github/ISSUE_TEMPLATE/bug_report.yml`. **Pre-submission checklist** - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am on the latest `master` commit for this PR branch. - [x] I have confirmed the error originates in Paperclip's `codex_local` adapter configuration boundary, not in a provider outage. **What happened?** New or updated `codex_local` agents could inherit a host-level `OPENAI_API_KEY` or use a shared Codex home when their adapter config did not explicitly isolate those values. That made it possible for future agents or manual adapter edits to silently fall back to shared OpenAI credentials. **Expected behavior** Creating, hiring, or updating a `codex_local` agent should either persist isolated per-agent configuration or reject unsafe shared Codex home configuration with a clear 422 response. The guard must not print secret values. **Steps to reproduce** 1. Create or update a `codex_local` agent without an explicit `adapterConfig.env.OPENAI_API_KEY` override. 2. Run it on a host where the Paperclip server process has `OPENAI_API_KEY` set. 3. Observe that the adapter process can inherit the host key unless Paperclip persists a blocking empty override. 4. Set `adapterConfig.env.CODEX_HOME` to a shared path such as `~/.codex` or the company-level `codex-home`. 5. Observe that the old code allowed the shared auth home instead of returning a validation error. **Paperclip version or commit** - Reproduced by inspection against `master` before this PR. **Deployment mode** - Local dev / self-hosted server with `codex_local` agents. **Installation method** - Built from source. **Agent adapter(s) involved** - Codex. **Database mode** - Not database-related. **Access context** - Board and agent configuration paths. **Relevant logs or output** - No secret-bearing logs included. **Relevant config** - Unsafe shape: missing `adapterConfig.env.OPENAI_API_KEY`, or shared `adapterConfig.env.CODEX_HOME`. - Fixed shape: per-agent `CODEX_HOME` plus empty `OPENAI_API_KEY` override. **Additional context** Related PR search for `codex_local OPENAI_API_KEY CODEX_HOME` found: - #3681 `fix: preserve managed Codex auth and repo-root env loading` - #5621 `fix: copy worktree codex auth locally` Those are adjacent auth-handling changes, but they do not add the agent create/update guard implemented here. **Privacy checklist** - [x] I have reviewed all pasted output for PII, usernames, file paths, API keys, tokens, company names, and redacted where necessary. ## What Changed - Added a `codex_local` config guard in agent create, hire, and update routes. - The guard assigns `adapterConfig.env.CODEX_HOME` to `companies/<companyId>/agents/<agentId>/codex-home` when missing. - The guard persists `adapterConfig.env.OPENAI_API_KEY = ""` when missing, preventing host env inheritance. - Shared `CODEX_HOME` values for the company codex-home, host `$CODEX_HOME`, or `~/.codex` now fail with a 422 error. - Added route tests for create, hire, update, and rejected shared host Codex home. - Updated `codex_local` and development docs to describe the per-agent home contract. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/agent-skills-routes.test.ts` - `pnpm typecheck` - `git diff --check upstream/master...HEAD` - `gh pr list --repo paperclipai/paperclip --state all --search "codex_local OPENAI_API_KEY CODEX_HOME" --limit 20 --json number,title,state,url` - `rg -n "codex|OPENAI_API_KEY|CODEX_HOME|adapter" ROADMAP.md` returned no roadmap overlap. ## Risks - Existing legacy `codex_local` agents with shared `CODEX_HOME` will get a clear 422 when their adapter config is updated until the shared path is replaced. This is intentional because silent fallback is the bug being guarded. - Low migration risk: no database migration and no secret values are printed or persisted beyond the empty override. ## Model Used - OpenAI GPT-5.5 Codex, Codex coding-agent session with repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Paperclip - Issue: [ZOL-5477](/ZOL/issues/ZOL-5477) - Owner: Разработчик (`6625498c-66c9-429f-b578-4463ddc3ba16`) - Status: waiting reviewer - Next action: merge after approval and green CI --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b7c42be86 |
fix(openclaw-gateway): complete and stabilize OpenClaw Gateway integration (#2322)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `openclaw_gateway` adapter is how operators wire Paperclip agents to an OpenClaw gateway over WebSocket > - The adapter UI previously only exposed a handful of config fields in edit mode; many timeout / auth / session-routing knobs were unreachable through the form > - The serializer also forgot to inject the configured `authToken` into the `x-openclaw-token` header, and the server-side execute path lacked retries on transient gateway errors and an `OPENCLAW_TOKEN` env fallback > - This pull request exposes the full set of config fields in both create and edit modes, fixes the serializer, hardens the server-side execute path, and pins the existing default request timeouts (120s / 120000ms) — see the dedicated commit and the new unit tests > - The benefit is operators can configure and reconfigure an `openclaw_gateway` agent end-to-end through the UI, with no silent change to the defaults documented in the adapter README and `doc/ONBOARDING_AND_TEST_PLAN.md` ## Linked Issues or Issue Description Closes #414 Closes #1901 Closes #2309 ## What Changed - **UI**: Removed the `!isCreate` guard so all `openclaw_gateway` config fields are visible in both create and edit modes (`authToken`, `agentId`, `sessionKeyStrategy`, `sessionKey`, `timeoutSec`, `waitTimeoutMs`, `disableDeviceAuth`, `autoPairOnFirstConnect`, `role`, `scopes`, `paperclipApiUrl`, `headersJson`, `payloadTemplate`, `runtimeServices`). - **Serialization** (`packages/adapters/openclaw-gateway/src/ui/build-config.ts`): inject `authToken` into headers as `x-openclaw-token`; apply safe defaults on create (`timeoutSec=120`, `waitTimeoutMs=120000`, `sessionKeyStrategy="issue"`, `role="operator"`, `scopes=["operator.admin"]`). - **Backend** (`packages/adapters/openclaw-gateway/src/server/execute.ts`): add `OPENCLAW_TOKEN` env-var fallback for `authToken`, retry logic (max 2 retries with backoff for transient gateway errors), session-key prefix `agent:{agentId}:{sessionId}` when `agentId` is configured. - **Defaults restoration** (dedicated commit): an earlier revision of this PR lowered the default request timeouts to `60` / `30000`. The current branch restores the historical `timeoutSec=120` / `waitTimeoutMs=120000` defaults that match the values documented in `packages/adapters/openclaw-gateway/src/index.ts`, `src/server/execute.ts` on master, and the worked example in `doc/ONBOARDING_AND_TEST_PLAN.md`. - **Tests** (new): `packages/adapters/openclaw-gateway/src/ui/build-config.test.ts` pins the documented timeout and identity defaults so the silent-halve regression cannot recur. ## Verification - `pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck` - `pnpm typecheck` (root) - Manual: create a new `openclaw_gateway` agent — all fields visible, defaults populate as documented. - Manual: edit an existing `openclaw_gateway` agent — every field round-trips correctly and saves. - Manual: unset `authToken` in the form and set `OPENCLAW_TOKEN` env var — adapter picks up the env-var fallback. - Manual: simulate a transient gateway error — execute retries up to 2 times with backoff before failing. ## Risks - Low risk. Surface area is one adapter, behind explicit operator configuration. The defaults change in this PR is a restoration of values that already exist on master and in the adapter docs, so no production agent sees a behavioral shift relative to the prior release. Field exposure in edit mode is purely additive — existing values are preserved on save. ## Model Used - Provider/model: Claude (Anthropic) — `claude-opus-4-7` - Mode: standard tool use, no extended thinking - Capability notes: code execution + repository file edits via Claude Code ## Cross-references and status (maintainer) Closes #414 Closes #1901 Closes #2309 ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip Bot <bot@paperclip.dev> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Devin Foley <devin@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
69a368ed51 |
fix(gemini-local): pre-select gemini-api-key auth in managed-HOME settings.json for headless runs (#7918)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The gemini-local adapter runs gemini-cli headlessly, including on remote/sandboxed execution targets where the adapter manages a dedicated HOME under the runtime root > - gemini-cli hard-refuses headless runs with "Invalid auth method selected." unless `$HOME/.gemini/settings.json` persists an auth selection; setting `GEMINI_DEFAULT_AUTH_TYPE` alone does NOT satisfy it (proven in an isolated pod) > - With a managed HOME the runtime root replaces the image home, so any settings.json baked into the agent image (or the user's real home) is invisible to the CLI, and every sandboxed gemini run dies before doing any work > - This affects any sandbox provider that runs gemini with API-key auth through the managed-HOME path (SSH, E2B, Daytona, Kubernetes, or any other remote execution target); it is a headless-execution bug fix, not gateway- or deployment-specific behavior > - This pull request makes the adapter pre-select the `gemini-api-key` auth type in the managed `$HOME/.gemini/settings.json` whenever a Gemini/Google API key is present, writing both settings schema generations and never touching an existing settings.json > - The benefit is that gemini agents actually run headlessly on remote and sandboxed execution targets without any manual settings provisioning ## Linked Issues or Issue Description No existing issue; describing the bug in-PR (bug template fields): - **What happened:** Headless gemini-local runs on remote/sandboxed execution targets fail immediately with `Invalid auth method selected.` even though `GEMINI_API_KEY` is provided. - **Expected:** Providing the API key should be enough for a headless run to authenticate and proceed. - **Root cause:** gemini-cli requires an auth selection persisted in `$HOME/.gemini/settings.json` for non-interactive runs; the `GEMINI_DEFAULT_AUTH_TYPE` env var does not substitute for it (verified in an isolated pod with only the env var set). The adapter's managed-HOME execution path points HOME at the runtime root, so any pre-existing settings.json (image-baked or user home) is hidden and the CLI finds no auth selection. - **Reproduction:** Run the gemini-local adapter against a remote/sandboxed execution target with `GEMINI_API_KEY` set and no settings.json under the managed HOME; the run aborts with the error above. - Duplicate/related search: no existing PR or issue addresses this; closest related is #7693 (bundles gemini-cli in the Docker image), which makes the CLI available but does not fix headless auth selection. ## What Changed - `packages/adapters/gemini-local/src/server/execute.ts`: after provisioning the managed HOME, when a Gemini/Google API key is present, write `$HOME/.gemini/settings.json` pre-selecting `gemini-api-key` auth. Both settings schema generations are written (legacy top-level `selectedAuthType` and current `security.auth.selectedType`) so old and new gemini-cli versions are covered. - The write is strictly scoped to the managed HOME (the per-run runtime root on sandbox transports). On non-managed remote targets (SSH), where the remote home is the user's real home and existing settings remain visible to the CLI, the adapter creates nothing (review feedback, P1). - The write is guarded by `[ -f ... ] ||` so a user-shipped settings.json (e.g. via workspace) is never overwritten. - The key-presence gate checks the run env AND the host process env (`GEMINI_API_KEY` / `GOOGLE_API_KEY`): in sandboxed paths the key never enters the adapter's run env; it reaches the agent pod via the sandbox provider's per-run secret (env passthrough from the host env), so the host env is the correct signal there. - `packages/adapters/gemini-local/src/server/execute.remote.test.ts`: a new sandbox-transport test asserts the settings.json write lands under the per-run runtime root (path + `gemini-api-key` content), and the SSH test asserts no settings.json is created on a non-managed home. ## Verification - `npx vitest run packages/adapters/gemini-local`: 3 files, 17 tests, all pass. - `pnpm --filter @paperclipai/adapter-gemini-local typecheck` and `build`: clean (test file is covered by the package tsconfig `include`). - Negative control: in an isolated pod, gemini-cli with `GEMINI_API_KEY` + `GEMINI_DEFAULT_AUTH_TYPE` set but no settings.json still fails with `Invalid auth method selected.`; with the settings.json written by this change, the run proceeds. - Verified end-to-end: a gemini agent in a hardened Kubernetes (gVisor) sandbox completed a real task (with `GOOGLE_GEMINI_BASE_URL` pointing at a GenAI-compatible endpoint), producing a billed usage row. That deployment supplies the verification evidence; the fix applies to any sandbox provider running gemini with API-key auth. ## Risks - Low risk. The new write only fires on the managed-HOME path (per-run runtime root) when an API key is present, and only when no settings.json exists yet, so existing setups, real user homes on SSH targets, and user-provided settings are unaffected. - If a future gemini-cli changes the settings schema again, the file may need a third generation key; both current generations are written today. ## Model Used - Claude (Anthropic), Claude Opus 4.8, 1M context, extended thinking, with tool use (code execution / shell) via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (no UI change) - [x] I have updated relevant documentation to reflect my changes (code comments document the behavior; no doc pages cover managed-home auth) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (review requested) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9e750d3e92 |
feat(codex-local): env-driven gateway routing via PAPERCLIP_CODEX_PROVIDERS config.toml (#7919)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex-local` adapter runs the OpenAI Codex CLI; Paperclip already maintains a managed `CODEX_HOME` per company and ships it to remote/sandboxed execution targets > - Deployments increasingly put an OpenAI-compatible LLM gateway between the harness and the model for cost, governance, or data-residency reasons: LiteLLM, OpenRouter, Portkey, Kong, a corporate proxy, self-hosted models (vLLM/Ollama), or region-pinned/sovereign endpoints. But Codex has no CLI flag or env var for a custom endpoint: its only mechanism is `[model_providers.<id>]` tables (with `base_url`, `env_key`, `wire_api`) in `$CODEX_HOME/config.toml`, selected by a root-level `model_provider` key > - Today there is no supported way to get such provider config into the managed `CODEX_HOME`, so gateway routing requires hand-editing files the adapter owns and regenerates > - This pull request adds the codex analogue of #7837's opencode mechanism: a `PAPERCLIP_CODEX_PROVIDERS` JSON env var whose shape maps 1:1 onto codex's TOML schema, merged into the managed `config.toml` so the existing asset-shipping + `env.CODEX_HOME` mechanics deliver it to local and sandboxed runs alike; nothing here is specific to one hosting setup > - The benefit is Codex works behind any OpenAI-compatible gateway with config only; with no env set, behavior is unchanged ## Linked Issues or Issue Description No existing issue; describing in-PR (feature / adapter enhancement). - **Gap:** there is no supported way to register a custom/gateway `[model_providers.*]` endpoint for `codex-local`. Codex's only custom-endpoint mechanism is `config.toml` (`base_url` + `env_key` + `wire_api`, selected via the root `model_provider` key), and the adapter owns/regenerates the managed `CODEX_HOME`, so operators cannot durably hand-edit it. - Related: #7837 (the opencode-local analogue of this change, same env-driven gateway-routing pattern). Searched for duplicate/related PRs: no existing codex-local gateway/provider-routing PR found. > Note on ROADMAP: this is adapter-level, opt-in config (defaults unchanged) that *enables* gateway routing for one harness; it is not the core "Cloud / Sandbox agents" platform work itself. ## What Changed - New `prepareCodexRuntimeConfig()` (`packages/adapters/codex-local/src/server/runtime-config.ts`): reads `PAPERCLIP_CODEX_PROVIDERS` (run env first, then `process.env`), shaped as `{"providers": {"<id>": {base_url, env_key, wire_api, ...}}, "model_provider": "<id>"}`, and merges it into the managed `CODEX_HOME`'s `config.toml`. No-op when unset or empty. - A malformed value (invalid JSON, not a JSON object, no `providers` object, no usable provider entries, or individual entries with empty names or non-object values, which are skipped by name) is never silently dropped: each case surfaces a distinct, user-visible note (via the prepare notes, which flow into command notes + `onLog`) and unusable input leaves `config.toml` untouched. - Merge is marker-delimited and TOML-correct: existing `config.toml` content is preserved between two managed blocks. Root keys (e.g. `model_provider`) are prepended **before the first table header** (TOML root-region rule), `[model_providers.*]` tables are appended. Pre-existing same-name provider sections and root `model_provider` keys are excised so the managed definitions win without duplicate-table parse errors. - `{env:VAR}` placeholders are expanded server-side for literal-credential fields; `env_key` indirection remains the preferred path. - Crash-safe restore: prepare writes a pre-run backup (`config.toml.paperclip-backup`) before the merged file; `cleanup()` restores the original in the execute `finally` and removes the backup. If a run never reaches `cleanup()` (a throw during the setup between prepare and execution, or SIGKILL), the next prepare restores the original from the backup with full fidelity, including user `[model_providers.*]` sections the merge excised (review feedback, P2); plain block-stripping remains the fallback for pre-backup state. - An explicit adapter-config `env.CODEX_HOME` override is treated as user-managed: no merge, surfaced as a command note. - Dependency-free hand-emitted TOML (strings/numbers/booleans, arrays of scalars, plain objects as inline tables); basic strings escape U+0000-U+001F and U+007F per TOML 1.0 (review feedback, P2). Merged output was additionally validated locally with python tomllib during development; the committed tests assert the structural invariants. - `execute.ts` wiring: `prepareCodexRuntimeConfig` runs after `prepareManagedCodexHome` (before the home ships to the remote target), notes surface via `onLog` + command notes, and the `finally` calls `cleanup()`. **Note for reviewers:** current codex removed `wire_api = "chat"` (openai/codex#10157, Feb 2026), so gateway provider configs must use `wire_api = "responses"`, i.e. the gateway must speak `/v1/responses`. The adapter passes the value through verbatim; this is a codex-side constraint worth knowing when configuring it. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local build` and `typecheck`: tsc clean against current `master` - `pnpm exec vitest run packages/adapters/codex-local`: 45 passing (incl. 17 `runtime-config` tests: fresh-merge + cleanup restore, root-region placement, same-name provider override, inline tables/arrays, DEL escaping, `{env:}` expansion from run env + `process.env`, per-case malformed-input notes with `config.toml` untouched, skipped-entry notes alongside a successful merge, silent no-op when unset/empty, explicit-`CODEX_HOME` skip note, backup restore of excised user sections after an interrupted run, backup removal on cleanup, stale-block self-heal, re-run replacement) - Verified end-to-end: a codex agent in a hardened Kubernetes (gVisor) sandbox completed a real task routed through an OpenAI-compatible gateway's `/v1/responses`, with a billed usage row recorded on the gateway. That deployment supplies the verification evidence; the mechanism is gateway-agnostic. ## Risks Low. Entirely env-driven and opt-in; with `PAPERCLIP_CODEX_PROVIDERS` unset the adapter never touches `config.toml` and behavior is byte-identical to before. The merge preserves user content, restores the original file on cleanup, and survives interrupted runs via the pre-run backup; malformed input surfaces a visible note and is ignored without touching `config.toml`. No migration/UI impact. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`, 1M context), extended thinking + tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md (adapter-level opt-in config enabling gateway routing; not the core sandbox-platform work, noted above) - [x] I have searched GitHub for duplicate or related PRs and linked them above (#7837 is the opencode analogue; no codex-local duplicate found) - [x] I have either (a) linked existing issues OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (n/a, no UI) - [ ] I have updated relevant documentation to reflect my changes (env var documented inline; no central doc references the adapter env yet) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (green on the previous head; re-running on the final note-copy polish commit) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (both review P2s are fixed at head: the interrupted-run restore via the pre-run backup and the U+007F escaping; a re-review is requested for the note-copy polish) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6e4aca9c67 |
feat(pi-local): env-driven gateway routing via PAPERCLIP_PI_PROVIDERS models.json (#7920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `pi-local` adapter runs the Pi coding agent, including inside remote/sandboxed execution targets; Pi resolves `--provider P --model M` by an exact (provider, id) match against its model registry, and it has no base-url CLI flag or env var: a `models.json` in its agent config dir (`$PI_CODING_AGENT_DIR`, falling back to `$HOME/.pi/agent`) is its only mechanism for custom or OpenAI/Anthropic-compatible endpoints > - Deployments increasingly put an LLM gateway between the harness and the model for cost, governance, or data-residency reasons: LiteLLM, OpenRouter, Portkey, Kong, a corporate proxy, self-hosted models (vLLM/Ollama), or region-pinned/sovereign endpoints. Today there is no supported way to get such provider config into Pi's registry for orchestrated runs > - The opencode adapter gained the equivalent capability in #7837 and codex in #7919; this pull request is the Pi analogue, so the harness layer stays gateway-agnostic regardless of which CLI an agent uses; nothing here is specific to one hosting setup > - This pull request reads `PAPERCLIP_PI_PROVIDERS` (Pi's `models.json` `providers` shape), materialises a managed `models.json` in a temp agent-config dir, points `PI_CODING_AGENT_DIR` at it, and ships it to remote execution targets with the run > - The benefit is Pi works behind any compatible gateway with config only; with no env set, behavior is unchanged ## Linked Issues or Issue Description No existing issue; describing in-PR (feature / adapter enhancement). - **Gap:** there is no supported way to register custom/gateway providers + models for `pi-local`. Pi's only custom-endpoint mechanism is a `models.json` in its agent config dir, and orchestrated (especially sandboxed) runs have no way to provision one declaratively. - Related: #7837 (the opencode-local analogue, same env-driven gateway-routing pattern) and #7919 (the codex-local analogue). Searched for duplicate or related PRs: no existing pi-local gateway/provider-routing PR found. > Note on ROADMAP: this is adapter-level, opt-in config (defaults unchanged) that *enables* gateway routing for one harness; it is not the core "Cloud / Sandbox agents" platform work itself. ## What Changed - New `packages/adapters/pi-local/src/server/runtime-config.ts`: `preparePiRuntimeConfig()` reads `PAPERCLIP_PI_PROVIDERS` (a JSON object in pi's `models.json` `providers` shape) from the run env, then `process.env`. When set, it expands `{env:VAR}` placeholders (run env first, then process env; unresolvable placeholders left intact), writes `{"providers": ...}` to a managed temp dir as `models.json`, and returns env with `PI_CODING_AGENT_DIR` pointing at it plus a cleanup handle. - `execute.ts`: the prepared dir ships to remote execution targets as the managed-runtime asset `agentConfig` (same mechanism as opencode's `xdgConfig`), and `PI_CODING_AGENT_DIR` is repointed to the in-target path; cleanup runs in `finally`. - Misconfiguration is visible, not silent: a set-but-unusable `PAPERCLIP_PI_PROVIDERS` (invalid JSON, not an object, no provider objects) surfaces an explanatory note instead of proceeding unconfigured into an opaque model-not-found failure later, and provider entries with non-object values are skipped with a note naming them. Unset/empty stays a silent no-op (feature off). - Defaults unchanged: with `PAPERCLIP_PI_PROVIDERS` unset, the adapter behaves byte-for-byte as before, for local runs and for every existing sandbox provider. ## Verification - All pi-local tests green against this base (new: providers written verbatim, `{env:VAR}` expansion from run env/process env/unresolvable, no-op when unset, `PI_CODING_AGENT_DIR` set and shipped, the misconfiguration notes incl. skipped non-object entries, remote asset sync + env repoint). Typecheck and build clean. - Production end-to-end evidence (our deployment, used as verification, not as the scope of the change): a pi agent in a Kubernetes gVisor sandbox resolved a custom provider from the shipped `models.json`, completed an assigned issue through an Anthropic-compatible gateway, and landed a billed usage row. ## Risks Low. The entire feature is opt-in behind one env var; the only behavior change when it is set is the intended one. The managed dir replaces the host agent dir for the run by design (credentials travel inside the provider config or via env-key indirection), which is the correct posture for orchestrated runs. No migration/UI impact. ## Model Used Claude Opus 4.8 (`claude-opus-4-8`, 1M context), extended thinking + tool use, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md (adapter-level opt-in config enabling gateway routing; not the core sandbox-platform work, noted above) - [x] I have searched GitHub for duplicate or related PRs and linked them above (#7837 and #7919 are the opencode/codex analogues; no pi-local duplicate found) - [x] I have either (a) linked existing issues OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (n/a, no UI) - [ ] I have updated relevant documentation to reflect my changes (env var documented inline; no central doc references the adapter env yet) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (green on the previous head; re-running on the final note-copy polish commit) - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (both prior review findings are fixed at head: the indirect notes-based guard is now an explicit `agentConfigDir` handle, and a failed `models.json` write no longer leaks the temp dir; a re-review is requested for the note-copy polish) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1ac1ba5442 |
feat(opencode-local): env-driven gateway routing (custom providers, small/cheap model, remote allow-all) (#7837)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `opencode-local` adapter runs the OpenCode harness; its
model/provider routing assumes built-in providers (anthropic/openai/...)
and their default models
> - Deployments increasingly put an OpenAI/Anthropic-compatible LLM
gateway between the harness and the model for cost, governance, or
data-residency reasons: LiteLLM, OpenRouter, Portkey, Kong, a corporate
proxy, self-hosted models (vLLM/Ollama), or region-pinned/sovereign
endpoints. But OpenCode only resolves `--model provider/model` when the
model is registered in a provider's `models` map, and
`OPENCODE_ALLOW_ALL_MODELS` does NOT bypass its internal `getModel()`
> - Several lanes also fall back to built-in default models the gateway
may not serve: the auxiliary/title model (e.g. `claude-haiku-*`) and the
budget/recovery "cheap" lane (`openai/gpt-5.1-codex-mini`); these abort
runs with "no keys found that support model"
> - This pull request makes the adapter's provider/model wiring
declarative via env, so any such deployment can register gateway models
+ pin the auxiliary/budget lanes without code changes; nothing here is
specific to one hosting setup
> - The benefit is OpenCode works behind any compatible gateway with
config only; with no env set, behavior is unchanged
## Linked Issues or Issue Description
No existing issue; describing in-PR (feature / adapter enhancement).
- **Gap:** there is no supported way to register custom/gateway
providers + models for `opencode-local`, nor to pin the auxiliary
(title-gen) and budget (recovery) model lanes, so routing OpenCode
through a gateway fails at `getModel()` or on the default helper models.
- Related: #5737 (exe.dev sandbox installs for gemini/opencode local),
#5823 (unblock claude_local on remote sandbox providers).
> Note on ROADMAP: this is adapter-level, opt-in config (defaults
unchanged) that *enables* gateway routing for one harness; it is not the
core "Cloud / Sandbox agents" platform work itself. Happy to
redirect/discuss in #dev if preferred.
## What Changed
- `PAPERCLIP_OPENCODE_PROVIDERS`: merge custom/extended providers
(OpenCode `provider` shape) into the runtime `opencode.json`, so gateway
models are registered and `--model provider/model` resolves. `{env:VAR}`
placeholders are expanded server-side (so a key need not depend on the
sandbox run env).
- A malformed `PAPERCLIP_OPENCODE_PROVIDERS` is no longer silently
ignored: invalid JSON, a non-object value, and individual provider
entries with non-object values (which are skipped by name) each append a
visible note to the run notes so the misconfiguration is diagnosable
(addresses both review P1s).
- `PAPERCLIP_OPENCODE_SMALL_MODEL` / `PAPERCLIP_OPENCODE_CHEAP_MODEL`:
pin the auxiliary (title-generation) and budget (recovery-retry) lanes
to gateway-served models; defaults unchanged.
- Honour `OPENCODE_ALLOW_ALL_MODELS` on the **remote** execution path
too (was local-only, a parity gap).
- `PAPERCLIP_OPENCODE_PRINT_LOGS`: optional toggle adding `--print-logs`
so OpenCode logs surface on stderr for diagnosing remote/sandbox runs.
- `buildOpenCodeModelProfiles()` guards its `process.env` default with
`typeof process` so the shared client/server module stays browser-safe
(a bare `process.env` at module load threw ReferenceError in the browser
under Vite dev middleware and broke UI rendering in the e2e lane).
## Verification
- `pnpm --filter @paperclipai/adapter-opencode-local build` and
`typecheck` (tsc clean)
- `pnpm exec vitest run packages/adapters/opencode-local/src` shows 33
passing (incl. new tests for the provider merge, `{env:}` expansion, the
malformed/non-object/skipped-entry provider notes, small/cheap-model
resolution, and the remote allow-all bypass)
- Manually verified end-to-end against a real
OpenAI-/Anthropic-compatible gateway: with the providers + small/cheap
model set, both the title-gen and main task route to the configured
gateway model and the agent completes (a real completion is returned and
billed). That deployment supplies the verification evidence; the
mechanism is gateway-agnostic.
## Risks
Low. Everything is env-driven and opt-in; with no env set the generated
config output is unchanged, and the cheap model profile keeps its model
(the only difference is its updated human-readable description).
Defaults preserved: built-in providers, Codex-mini cheap lane with
`variant: low`, no `--print-logs`. No migration/UI impact.
## Model Used
Claude Opus 4.8 (`claude-opus-4-8`, 1M context), extended thinking +
tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md (adapter-level opt-in config enabling
gateway routing; not the core sandbox-platform work, noted above)
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (#5737, #5823)
- [x] I have either (a) linked existing issues OR (b) described the
issue in-PR following the relevant issue template
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] If this change affects the UI, I have included before/after
screenshots (n/a, no UI)
- [ ] I have updated relevant documentation to reflect my changes (env
vars documented inline via comments; no central doc references the
adapter env yet)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(the P1 about silently dropped malformed providers JSON is addressed in
|
||
|
|
c139d6c025 |
fix(codex-local): omit default model so codex CLI picks per auth mode (#7971)
## Thinking Path > - Paperclip orchestrates AI agents through pluggable local adapters; codex_local wraps OpenAI's `codex` CLI. > - The codex_local adapter declares a hard-coded `DEFAULT_CODEX_LOCAL_MODEL = "gpt-5.3-codex"` and multiple Paperclip consumers (UI build-config, server route, OnboardingWizard, NewAgent form, AgentConfigForm) fall back to it when the operator doesn't pick a model. > - That model — and every `*-codex` model plus the older `gpt-5/5.1/5.2` lines — is API-key-only. Codex CLI rejects them on ChatGPT subscription auth with "The 'gpt-5.3-codex' model is not supported when using Codex with a ChatGPT account." > - Every codex_local agent created through the default onboarding path inherits this pin and breaks on its first heartbeat for any user authed via `codex login` (ChatGPT). > - claude_local already takes the right shape: its build-config only sets `adapterConfig.model` when the operator actually picked one, and falls through to whatever default `claude` CLI uses. > - Codex CLI's own default is auth-mode-aware. ChatGPT-subscription accounts get `gpt-5.5`; API-key accounts get the codex-tuned default. A Paperclip-side pin masks this and downgrades whichever group it wasn't built for. > - This PR makes codex_local match claude_local's shape: omit `adapterConfig.model` when the user picks "default," and let the CLI choose. Subscription users stop breaking; API-key users stop getting downgraded. > - The benefit is auth-mode-correct defaults with no Paperclip-side hard pin, plus future-proofing: when OpenAI bumps the CLI default we inherit it for free. ## What Changed - `packages/adapters/codex-local/src/ui/build-config.ts` — only set `adapterConfig.model` when the operator picked one (parity with `packages/adapters/claude-local/src/ui/build-config.ts`). - `server/src/routes/agents.ts` — drop the codex_local-specific `next.model = DEFAULT_CODEX_LOCAL_MODEL` fallback in `applyCreateDefaultsByAdapterType`. Bypass-sandbox default is left in place (security posture, not a model choice). - `ui/src/pages/NewAgent.tsx`, `ui/src/components/AgentConfigForm.tsx`, `ui/src/components/OnboardingWizard.tsx` — stop pre-populating the model field with `DEFAULT_CODEX_LOCAL_MODEL` when the user selects the Codex adapter. Other adapters' defaults (gemini_local, cursor, opencode_local) are unchanged. - `DEFAULT_CODEX_LOCAL_MODEL` is preserved as an exported constant for downstream consumers / plugin authors who want to opt in to a pin; we just stop forcing it on operators who didn't ask for one. - Test: assert `buildCodexLocalConfig` omits `model` when input is blank. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/ui/build-config.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts server/src/__tests__/adapter-registry.test.ts server/src/__tests__/heartbeat-model-profile.test.ts server/src/__tests__/agent-permissions-routes.test.ts` → 74/74 passing - `pnpm exec vitest run ui/src/lib/duplicate-agent-payload.test.ts ui/src/lib/acpx-model-filter.test.ts` → passing - `pnpm tsc --noEmit -p .` → clean - Live: I separately verified live during initial investigation that on ChatGPT-subscription auth, `gpt-5.3-codex` is rejected and `gpt-5.5` is what Codex CLI picks by default. Omitting model lets the CLI handle that. ## Risks - Telemetry: any sink that reads `adapterConfig.model` for cost attribution will now see the empty/omitted case more often. The CLI emits the actually-used model in its event stream; downstream telemetry should already read from there for accuracy, but worth a check. - Operator UX: "default" now means "whatever the CLI picks" instead of a Paperclip-known model. The selectable catalog still includes `gpt-5.5`, `gpt-5.4`, `gpt-5.3-codex`, etc. for operators who want to pin explicitly. - Existing agents are unaffected — their `adapterConfig.model` is already set; this only changes the *new-agent* default flow. ## Related work - Depends on: an open catalog-add PR adding `gpt-5.5` to the selectable model list and to `CODEX_LOCAL_FAST_MODE_SUPPORTED_MODELS`. Operators who want to switch to `gpt-5.5` explicitly need that PR merged first; this PR is the structural change that makes "default" mean "let the CLI choose." - Closes #5371 — codex_local default model selection persists `gpt-5.3-codex` instead of adapter default (this PR is the exact fix #5371 proposes). - Related: #5132 (opencode-local: hire-time default model fails on ChatGPT-OAuth accounts) — same problem shape on a sibling adapter; not fixed here but worth tracking for a parallel. - Related: #5939 (codex_local adapter hardcodes `gpt-5.3-codex-spark` validation, fails on ChatGPT OAuth accounts regardless of configured model) — separate validation-path bug; not fixed here. ## Model Used Claude (Sonnet-class), running inside Paperclip as a claude_local executor. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |