Files
PaperClipAI/docs/adapters/claude-local.md
T
DottaandPaperclip f447990d77 fix(claude): recognize ACP quota fallback errors (#13831)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Claude adapter reports failures to the recovery system.
> - Recovery must distinguish account quota from other provider limits.
> - The Claude ACP bridge has a default message for an account with no
quota.
> - The current classifier misses that message and returns
`acpx_turn_failed`.
> - This pull request recognizes that exact typed message and uses the
existing quota wait.

## Linked Issues or Issue Description

Refs #13651. This is a narrow follow-up to its typed quota
classification.
Related PRs #13549 and #10276 address broader quota classification and
reset handling. This change covers the bridge's exact default quota
message.

**What happened?**

A typed Claude ACP `limit` failure with the title `The Claude account
has no available quota.` returns `acpx_turn_failed`. The recovery result
has no quota label.

**Expected behavior**

Return `provider_quota`. Use the existing one-hour quota backoff when
the provider gives no reset time. Keep context, turn, rate, and
configured budget limits out of the quota path.

**Steps to reproduce**

1. Use the local ACP fixture to return the exact title above with
category `limit` and severity `error`.
2. Run the Claude adapter in oneshot or persistent mode.
3. Before this change, the new regression cases receive
`acpx_turn_failed` instead of `provider_quota` on ACPX 0.12.0 and
0.13.1.

**Paperclip version or commit**

Reproduced on `a959e4750`. Rebased onto current `master` before
submission.

**Deployment mode**

Local source checkout with isolated ACP child-process fixtures. No live
provider calls or customer-stack changes.

## What Changed

- Recognize the exact Claude bridge quota fallback only for typed
`limit` failures.
- Add real-process regression cases for both pinned ACPX versions and
both session modes.
- Verify that unrelated categories, positive quota wording, and
historical generic limit errors do not imply quota exhaustion.
- Document the fallback and its existing recovery backoff.

## Verification

- Red/green reproduction: all four new fallback cases failed before the
classifier change and passed afterward.
- 108 targeted tests passed across Claude ACP, quota, parser, and server
recovery suites.
- Claude adapter typecheck and build passed.
- Real-process tests verify that provider text stays out of results and
logs.
- Full local `pnpm -r typecheck` and `pnpm build` passed.
- The full local `pnpm test:run` attempt stopped after embedded
PostgreSQL could not load a missing library symlink. The dependency
setup was repaired in the worktree. The isolated database test then
passed. The complete test suites passed in GitHub CI.
- All 54 latest-head checks passed on
`3d4d4e65386c5b6023ba34e6a2abcd94cff9a9bc`. Two optional Storybook jobs
were skipped.
- Greptile: 5/5, with no inline review threads or requested changes.

## Risks

- Low risk. An exact match is required inside an existing typed `limit`
failure.
- If upstream changes this wording, this fallback can stop matching.
Existing quota-message detection remains in place.
- No schema, credential, or permission changes. Historical generic
errors remain ambiguous and are not reclassified.

## Model Used

OpenAI Codex, GPT-6. The session identifies the model family as GPT-6
but does not expose a more specific model ID or context-window size.
Used reasoning, repository inspection, code editing, and terminal test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:01:20 -05:00

154 lines
8.6 KiB
Markdown

---
title: Claude Code
summary: Claude Code local adapter setup and configuration
---
The `claude_local` adapter runs Anthropic's Claude Code CLI locally. It supports session persistence, skills injection, and structured output parsing.
## Quota waits
Claude ACP runs that end with a typed provider-quota error retain the quota
classification and any parsed reset time. Recovery waits until that time, or
uses its existing one-hour quota backoff when no reset time is available.
This includes the Claude bridge's typed “The Claude account has no available
quota.” fallback, which carries no reset timestamp.
The adapter inspects the terminal provider message in memory; the run result
and run log retain only the generic failure message, recovery labels, and reset
timestamp. Context, turn, rate, and configured budget limits are not treated as
subscription quota exhaustion merely because ACP labels them `limit`.
## Prerequisites
- Claude Code CLI installed (`claude` command available)
- Either `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` in adapter or
environment env (or host env), or a Claude Code subscription login
available to the execution target
## Configuration Fields
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `cwd` | string | Yes | Working directory for the agent process (absolute path; created automatically if missing when permissions allow) |
| `model` | string | No | Claude model to use (default: `claude-opus-5`) |
| `promptTemplate` | string | No | Prompt used for all runs |
| `env` | object | No | Environment variables (supports secret refs) |
| `timeoutSec` | number | No | Process timeout (0 = no timeout) |
| `graceSec` | number | No | Grace period before force-kill |
| `maxTurnsPerRun` | number | No | Max agentic turns per heartbeat (defaults to `300`) |
| `dangerouslySkipPermissions` | boolean | No | Skip permission prompts (default: `true`); required for headless runs where interactive approval is impossible |
## Default model
An omitted, empty, or whitespace-only `model` uses Claude Opus 5
(`claude-opus-5`) on both the CLI and ACP engines. This also applies to existing
agents with an unset model, including agents created through the API and agents
running in sandboxes. No database migration is needed. The editor shows the
Paperclip default and leaves the setting unset until you select a model.
An explicit `model` takes precedence over `ANTHROPIC_MODEL`. When only
`ANTHROPIC_MODEL` is configured, the adapter keeps that override. Bedrock and
Vertex configurations without an explicit model keep their provider-specific
default because those providers use different model IDs. Host environment
settings apply only to local targets when resolving the model.
The default does not change explicitly configured agent models or the separate
Paperclip Runner's qualified provider profiles.
## Prompt Templates
Templates support `{{variable}}` substitution:
| Variable | Value |
|----------|-------|
| `{{agentId}}` | Agent's ID |
| `{{companyId}}` | Company ID |
| `{{runId}}` | Current run ID |
| `{{agent.name}}` | Agent's name |
| `{{company.name}}` | Company name |
## Session Persistence
The adapter persists Claude Code session IDs between heartbeats. On the next wake, it resumes the existing conversation so the agent retains full context.
Session resume is cwd-aware: if the agent's working directory changed since the last run, a fresh session starts instead.
If resume fails with an unknown session error, the adapter automatically retries with a fresh session.
### Poisoned `previous_message_id` (recovery)
Symptom in logs / issue thread:
```
API Error: 400 diagnostics.previous_message_id: must be the `id` from a prior /v1/messages response (starts with `msg_`)
```
What it means: the on-disk Claude Code transcript JSONL for that session contains a malformed (non-`msg_`-prefixed) `previous_message_id`. Anthropic's `/v1/messages` rejects every resume attempt against that transcript with a deterministic 400. Without guards, Paperclip would re-persist the same poisoned session id and the issue is stranded permanently — see [RED-976](../../../) / [RED-978](../../../).
What the adapter does automatically:
1. **Auto-rotate on resume.** If a `--resume` attempt returns this 400, the adapter retries once with a fresh session, deletes the poisoned `<session>.jsonl` from the local Claude config dir (best effort), and uses the fresh session id going forward.
2. **Validate-before-persist.** A result that carries this 400 never gets its `session_id` written back to the task session store, even if Claude Code emits one in the result event. The adapter returns `sessionId: null`, `sessionParams: null`, and `errorCode: "claude_poisoned_previous_message_id"`.
3. **Clear-on-error.** The adapter sets `clearSession: true` on the result, which causes the heartbeat service to drop any persisted session row for that issue (`clearTaskSessions`). The next continuation starts from a clean slate.
On-call checklist if you see this in production:
- Confirm `errorCode` is `claude_poisoned_previous_message_id` in the run row — that means the guards fired correctly and the issue auto-recovers on the next heartbeat.
- If the same issue still loops after one heartbeat, check that `agentTaskSessions` for that `(agentId, taskKey)` was cleared. If not, the adapter return value was lost (e.g. a malformed run finalization) — escalate; do **not** manually edit the row, file a child issue with the run id.
- For remote execution targets (sandbox/SSH), the poisoned JSONL is on the remote and the adapter only logs the cleanup intent. The fresh-session retry still succeeds because it uses a new session id, and the server-side `clearSession: true` is authoritative regardless of remote disk state.
## Skills Injection
The adapter creates a temporary directory with symlinks to Paperclip skills and passes it via `--add-dir`. This makes skills discoverable without polluting the agent's working directory.
## Remote credential ownership
When no API key or `CLAUDE_CODE_OAUTH_TOKEN` is configured,
`claude_local` uses a snapshot-owns-auth topology for managed sandbox execution
targets. When the run uses a sandbox execution target and no explicit
`CLAUDE_CONFIG_DIR` is configured, Paperclip creates a remote
`CLAUDE_CONFIG_DIR` under the run's Claude runtime directory. It uploads
sanitized host-side settings such as `settings.json` and `CLAUDE.md`, but the
managed seed does not upload host Claude credential files.
After the seed is copied, the remote materialization command checks the
execution target's own `$HOME/.claude` directory. For each missing credential
file, it copies `.credentials.json` or `credentials.json` from that remote home
into the managed `CLAUDE_CONFIG_DIR`. That means credentials baked into the
sandbox image win for managed remote Claude runs.
Worked example: a sandbox image contains `$HOME/.claude/.credentials.json` from
its own Claude Code login. Paperclip starts a managed remote `claude_local` run,
uploads only the sanitized config seed, and sets `CLAUDE_CONFIG_DIR` to the
remote runtime config path. Because the managed config has no credential file,
the adapter copies the sandbox image's `$HOME/.claude/.credentials.json` into
that path before invoking Claude. The sandbox snapshot owns the credential for
the run.
This differs from [`codex_local`](/adapters/codex-local), where a
Paperclip-managed sandbox run uploads a host-owned `CODEX_HOME/auth.json` and
therefore shadows any Codex login already present inside the sandbox image.
For manual local CLI usage outside heartbeat runs (for example running as `claudecoder` directly), use:
```sh
npx paperclipai agent local-cli claudecoder --company-id <company-id>
```
This installs Paperclip skills in `~/.claude/skills`, creates an agent API key, and prints shell exports to run as that agent.
## Environment Test
Use the "Test Environment" button in the UI to validate the adapter config. It checks:
- Claude CLI is installed and accessible
- Working directory is absolute and available (auto-created if missing and permitted)
- API key/auth mode hints (`ANTHROPIC_API_KEY` vs `CLAUDE_CODE_OAUTH_TOKEN` vs subscription login)
- A live hello probe (`claude --print - --output-format stream-json --verbose` with prompt `Respond with hello.`) to verify CLI readiness
The probe sees the same layered env as a real run: when an environment is
selected, its environment variables (secret refs included) are resolved and
merged under the adapter config's `env`, so environment-level auth is
reflected in the test result. A secret binding that is missing surfaces as
an `environment_env_binding_missing` failure instead of a silently passing
probe.