Files
PaperClipAI/docs/adapters/claude-local.md
DottaandPaperclip f447990d77 fix(claude): recognize ACP quota fallback errors (#13831)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Claude adapter reports failures to the recovery system.
> - Recovery must distinguish account quota from other provider limits.
> - The Claude ACP bridge has a default message for an account with no
quota.
> - The current classifier misses that message and returns
`acpx_turn_failed`.
> - This pull request recognizes that exact typed message and uses the
existing quota wait.

## Linked Issues or Issue Description

Refs #13651. This is a narrow follow-up to its typed quota
classification.
Related PRs #13549 and #10276 address broader quota classification and
reset handling. This change covers the bridge's exact default quota
message.

**What happened?**

A typed Claude ACP `limit` failure with the title `The Claude account
has no available quota.` returns `acpx_turn_failed`. The recovery result
has no quota label.

**Expected behavior**

Return `provider_quota`. Use the existing one-hour quota backoff when
the provider gives no reset time. Keep context, turn, rate, and
configured budget limits out of the quota path.

**Steps to reproduce**

1. Use the local ACP fixture to return the exact title above with
category `limit` and severity `error`.
2. Run the Claude adapter in oneshot or persistent mode.
3. Before this change, the new regression cases receive
`acpx_turn_failed` instead of `provider_quota` on ACPX 0.12.0 and
0.13.1.

**Paperclip version or commit**

Reproduced on `a959e4750`. Rebased onto current `master` before
submission.

**Deployment mode**

Local source checkout with isolated ACP child-process fixtures. No live
provider calls or customer-stack changes.

## What Changed

- Recognize the exact Claude bridge quota fallback only for typed
`limit` failures.
- Add real-process regression cases for both pinned ACPX versions and
both session modes.
- Verify that unrelated categories, positive quota wording, and
historical generic limit errors do not imply quota exhaustion.
- Document the fallback and its existing recovery backoff.

## Verification

- Red/green reproduction: all four new fallback cases failed before the
classifier change and passed afterward.
- 108 targeted tests passed across Claude ACP, quota, parser, and server
recovery suites.
- Claude adapter typecheck and build passed.
- Real-process tests verify that provider text stays out of results and
logs.
- Full local `pnpm -r typecheck` and `pnpm build` passed.
- The full local `pnpm test:run` attempt stopped after embedded
PostgreSQL could not load a missing library symlink. The dependency
setup was repaired in the worktree. The isolated database test then
passed. The complete test suites passed in GitHub CI.
- All 54 latest-head checks passed on
`3d4d4e65386c5b6023ba34e6a2abcd94cff9a9bc`. Two optional Storybook jobs
were skipped.
- Greptile: 5/5, with no inline review threads or requested changes.

## Risks

- Low risk. An exact match is required inside an existing typed `limit`
failure.
- If upstream changes this wording, this fallback can stop matching.
Existing quota-message detection remains in place.
- No schema, credential, or permission changes. Historical generic
errors remain ambiguous and are not reclassified.

## Model Used

OpenAI Codex, GPT-6. The session identifies the model family as GPT-6
but does not expose a more specific model ID or context-window size.
Used reasoning, repository inspection, code editing, and terminal test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:01:20 -05:00

8.6 KiB

title, summary
title summary
Claude Code Claude Code local adapter setup and configuration

The claude_local adapter runs Anthropic's Claude Code CLI locally. It supports session persistence, skills injection, and structured output parsing.

Quota waits

Claude ACP runs that end with a typed provider-quota error retain the quota classification and any parsed reset time. Recovery waits until that time, or uses its existing one-hour quota backoff when no reset time is available. This includes the Claude bridge's typed “The Claude account has no available quota.” fallback, which carries no reset timestamp. The adapter inspects the terminal provider message in memory; the run result and run log retain only the generic failure message, recovery labels, and reset timestamp. Context, turn, rate, and configured budget limits are not treated as subscription quota exhaustion merely because ACP labels them limit.

Prerequisites

  • Claude Code CLI installed (claude command available)
  • Either ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN in adapter or environment env (or host env), or a Claude Code subscription login available to the execution target

Configuration Fields

Field Type Required Description
cwd string Yes Working directory for the agent process (absolute path; created automatically if missing when permissions allow)
model string No Claude model to use (default: claude-opus-5)
promptTemplate string No Prompt used for all runs
env object No Environment variables (supports secret refs)
timeoutSec number No Process timeout (0 = no timeout)
graceSec number No Grace period before force-kill
maxTurnsPerRun number No Max agentic turns per heartbeat (defaults to 300)
dangerouslySkipPermissions boolean No Skip permission prompts (default: true); required for headless runs where interactive approval is impossible

Default model

An omitted, empty, or whitespace-only model uses Claude Opus 5 (claude-opus-5) on both the CLI and ACP engines. This also applies to existing agents with an unset model, including agents created through the API and agents running in sandboxes. No database migration is needed. The editor shows the Paperclip default and leaves the setting unset until you select a model.

An explicit model takes precedence over ANTHROPIC_MODEL. When only ANTHROPIC_MODEL is configured, the adapter keeps that override. Bedrock and Vertex configurations without an explicit model keep their provider-specific default because those providers use different model IDs. Host environment settings apply only to local targets when resolving the model.

The default does not change explicitly configured agent models or the separate Paperclip Runner's qualified provider profiles.

Prompt Templates

Templates support {{variable}} substitution:

Variable Value
{{agentId}} Agent's ID
{{companyId}} Company ID
{{runId}} Current run ID
{{agent.name}} Agent's name
{{company.name}} Company name

Session Persistence

The adapter persists Claude Code session IDs between heartbeats. On the next wake, it resumes the existing conversation so the agent retains full context.

Session resume is cwd-aware: if the agent's working directory changed since the last run, a fresh session starts instead.

If resume fails with an unknown session error, the adapter automatically retries with a fresh session.

Poisoned previous_message_id (recovery)

Symptom in logs / issue thread:

API Error: 400 diagnostics.previous_message_id: must be the `id` from a prior /v1/messages response (starts with `msg_`)

What it means: the on-disk Claude Code transcript JSONL for that session contains a malformed (non-msg_-prefixed) previous_message_id. Anthropic's /v1/messages rejects every resume attempt against that transcript with a deterministic 400. Without guards, Paperclip would re-persist the same poisoned session id and the issue is stranded permanently — see RED-976 / RED-978.

What the adapter does automatically:

  1. Auto-rotate on resume. If a --resume attempt returns this 400, the adapter retries once with a fresh session, deletes the poisoned <session>.jsonl from the local Claude config dir (best effort), and uses the fresh session id going forward.
  2. Validate-before-persist. A result that carries this 400 never gets its session_id written back to the task session store, even if Claude Code emits one in the result event. The adapter returns sessionId: null, sessionParams: null, and errorCode: "claude_poisoned_previous_message_id".
  3. Clear-on-error. The adapter sets clearSession: true on the result, which causes the heartbeat service to drop any persisted session row for that issue (clearTaskSessions). The next continuation starts from a clean slate.

On-call checklist if you see this in production:

  • Confirm errorCode is claude_poisoned_previous_message_id in the run row — that means the guards fired correctly and the issue auto-recovers on the next heartbeat.
  • If the same issue still loops after one heartbeat, check that agentTaskSessions for that (agentId, taskKey) was cleared. If not, the adapter return value was lost (e.g. a malformed run finalization) — escalate; do not manually edit the row, file a child issue with the run id.
  • For remote execution targets (sandbox/SSH), the poisoned JSONL is on the remote and the adapter only logs the cleanup intent. The fresh-session retry still succeeds because it uses a new session id, and the server-side clearSession: true is authoritative regardless of remote disk state.

Skills Injection

The adapter creates a temporary directory with symlinks to Paperclip skills and passes it via --add-dir. This makes skills discoverable without polluting the agent's working directory.

Remote credential ownership

When no API key or CLAUDE_CODE_OAUTH_TOKEN is configured, claude_local uses a snapshot-owns-auth topology for managed sandbox execution targets. When the run uses a sandbox execution target and no explicit CLAUDE_CONFIG_DIR is configured, Paperclip creates a remote CLAUDE_CONFIG_DIR under the run's Claude runtime directory. It uploads sanitized host-side settings such as settings.json and CLAUDE.md, but the managed seed does not upload host Claude credential files.

After the seed is copied, the remote materialization command checks the execution target's own $HOME/.claude directory. For each missing credential file, it copies .credentials.json or credentials.json from that remote home into the managed CLAUDE_CONFIG_DIR. That means credentials baked into the sandbox image win for managed remote Claude runs.

Worked example: a sandbox image contains $HOME/.claude/.credentials.json from its own Claude Code login. Paperclip starts a managed remote claude_local run, uploads only the sanitized config seed, and sets CLAUDE_CONFIG_DIR to the remote runtime config path. Because the managed config has no credential file, the adapter copies the sandbox image's $HOME/.claude/.credentials.json into that path before invoking Claude. The sandbox snapshot owns the credential for the run.

This differs from codex_local, where a Paperclip-managed sandbox run uploads a host-owned CODEX_HOME/auth.json and therefore shadows any Codex login already present inside the sandbox image.

For manual local CLI usage outside heartbeat runs (for example running as claudecoder directly), use:

npx paperclipai agent local-cli claudecoder --company-id <company-id>

This installs Paperclip skills in ~/.claude/skills, creates an agent API key, and prints shell exports to run as that agent.

Environment Test

Use the "Test Environment" button in the UI to validate the adapter config. It checks:

  • Claude CLI is installed and accessible
  • Working directory is absolute and available (auto-created if missing and permitted)
  • API key/auth mode hints (ANTHROPIC_API_KEY vs CLAUDE_CODE_OAUTH_TOKEN vs subscription login)
  • A live hello probe (claude --print - --output-format stream-json --verbose with prompt Respond with hello.) to verify CLI readiness

The probe sees the same layered env as a real run: when an environment is selected, its environment variables (secret refs included) are resolved and merged under the adapter config's env, so environment-level auth is reflected in the test result. A secret binding that is missing surfaces as an environment_env_binding_missing failure instead of a silently passing probe.