Files
PaperClipAI/docs/adapters/overview.md
T
Devin Foley f0b06d2de9 feat(claude): environment-aware test-environment probe and claude-local CI coverage (#10833)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The claude_local adapter runs Claude Code on sandbox execution
targets, and operators verify an agent's configuration with the
test-environment probe before running it
> - Real runs merge the selected environment's env vars (secret refs
included) under the agent's adapter config env, but the probe built its
config from the adapter config alone — so environment-level auth worked
in runs while the Test button reported missing auth, and a dropped
secret binding passed silently
> - The claude env-test hints also did not recognize
`CLAUDE_CODE_OAUTH_TOKEN` even though the CLI accepts it, and a hello
probe that hit the subscription usage limit reported a hard failure
although authentication worked
> - Separately, the claude-local package test suites were absent from
the CI project list, so two suites drifted broken without notice
> - This pull request makes the probe resolve the same layered env as a
real run, adds the missing auth hint, classifies usage-limit probe
results as a warning, repairs the drifted suites, and turns the
claude-local project on in CI
> - The benefit is a Test button that tells the truth about
environment-level configuration, and a test suite that actually gates
the claude-local adapter

## Linked Issues or Issue Description

No public issue exists; related open PRs: Refs #9488 (recognizes
CLAUDE_CODE_OAUTH_TOKEN in environment checks — overlaps with the
auth-hint portion of this PR via a differently named check; it does not
cover the environment-envVars probe merge, the usage-limit
classification, or the CI coverage), Refs #9933 (live credential
validation in environment checks — complementary, no file-level conflict
with the route change).

The underlying problem, following the enhancement template:

**Current behavior**

The test-environment route builds the probe config from the agent's
adapterConfig only. Real runs merge the selected environment's envVars
under the agent env, so environment-level env vars (including auth such
as `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` bound as environment
secrets) work in runs while "Test environment" cannot see them, and a
missing secret binding passes silently. The claude env-test hints do not
recognize `CLAUDE_CODE_OAUTH_TOKEN`. A hello probe that hits the
subscription usage limit reports a hard `claude_hello_probe_failed`. The
claude-local package test suites do not run in CI, and two of them are
stale.

**Proposed behavior**

The probe resolves the selected environment's envVars
(environment-consumer secret bindings included) and merges them under
the agent config env with the run-path precedence; missing bindings
surface as an explicit error check that fails the test. The env-test
emits a `claude_oauth_token_configured` info check when that variable is
set. Usage-limit probe results classify as a
`claude_hello_probe_usage_limited` warning because auth works and only
the usage window is spent. The claude-local suites run in CI. Docs state
the resulting facts.

**Reason and benefit**

The Test button should tell the truth: it previously contradicted run
behavior for environment-level configuration and hid broken secret
bindings. Enabling the package suites in CI prevents further silent
drift — two suites were already broken on master without anyone
noticing.

**Breaking changes**

None. Runs are unchanged. The probe route only adds env layers and
checks; setups without environment envVars behave exactly as before.

## What Changed

- `server/src/routes/agents.ts`: the test-environment route resolves the
selected environment's envVars (forbidden keys stripped,
environment-consumer secret context) and merges them under the agent
adapterConfig env, mirroring `resolveExecutionRunAdapterConfig`
precedence. Missing secret bindings are skipped, reported as an
`environment_env_binding_missing` error check, and fail the test —
matching the `ConfigurationIncompleteFailure` a real dispatch would
raise.
- `packages/adapters/claude-local/src/server/test.ts`: new
`claude_oauth_token_configured` info hint between the API-key warning
and the subscription fallback; hello-probe classification gains a
`claude_hello_probe_usage_limited` warning for provider-quota results
(previously a hard `claude_hello_probe_failed`).
- `scripts/run-vitest-stable.mjs`: add
`@paperclipai/adapter-claude-local` to `nonServerProjects` so CI runs
the package suites.
- `packages/adapters/claude-local/src/server/execute.remote.test.ts`:
assert both runtime asset syncs (skills and mcp-config); the suite
predated the mcp-config asset.
- `packages/adapters/claude-local/src/server/test.probe.test.ts`:
usage-limit fixture now expects the usage-limited warning; new fixture
covers the genuine transient path (529 overloaded); new tests cover the
token hint and API-key precedence.
- `server/src/__tests__/agent-test-environment-routes.test.ts`: new
tests for the env merge (agent wins on conflict, forbidden key
filtered), missing-binding reporting, and the no-execution-target
fallback path.
- `docs/adapters/claude-local.md`, `docs/adapters/overview.md`: state
the auth-input facts (API key or oauth token wins over stored logins;
snapshot-owns-auth applies when neither is configured) and describe the
environment-aware Test behavior.

## Verification

- `npx vitest run --project @paperclipai/adapter-claude-local` — 131
tests pass (both drifted suites repaired; they fail on master today).
- `npx vitest run
server/src/__tests__/agent-test-environment-routes.test.ts` — 7 tests
pass.
- `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` —
passes with the added project.
- `pnpm typecheck` in `server/` and `packages/adapters/claude-local/` —
clean.

## Risks

- Low risk. The run path is untouched; the probe route change is
additive and inert when the environment has no envVars.
- The probe now performs environment-consumer secret resolution at test
time; access is authorized per binding exactly as at run time, and the
audit consumer is the environment (as before for adapter-config
resolution).
- Enabling the claude-local project in CI adds about 2 seconds of vitest
wall time to the general workspaces group and could surface future
regressions in that package — which is the point.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic
tool use via Claude Code (CLI).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-04 10:31:14 -07:00

8.8 KiB

title, summary
title summary
Adapters Overview What adapters are and how they connect agents to Paperclip

Adapters are the bridge between Paperclip's orchestration layer and agent runtimes. Each adapter knows how to invoke a specific type of AI agent and capture its results.

How Adapters Work

When a heartbeat fires, Paperclip:

  1. Looks up the agent's adapterType and adapterConfig
  2. Calls the adapter's execute() function with the execution context
  3. The adapter spawns or calls the agent runtime
  4. The adapter captures stdout, parses usage/cost data, and returns a structured result

Built-in Adapters

Adapter Type Key Description
Claude Code claude_local Runs Claude Code CLI locally, with a native ACP engine when available
Codex codex_local Runs OpenAI Codex CLI locally, with a native ACP engine when available
Gemini CLI gemini_local Runs Gemini CLI locally (experimental — adapter package exists, not yet in stable type enum)
OpenCode opencode_local Runs OpenCode CLI locally (multi-provider provider/model)
Cursor cursor Runs Cursor in background mode
Pi pi_local Runs an embedded Pi agent locally
Hermes hermes_local Runs the local Hermes CLI through @paperclipai/hermes-paperclip-adapter
Hermes Gateway hermes_gateway Calls an already-running Hermes API server through @paperclipai/hermes-paperclip-adapter/gateway
OpenClaw Gateway openclaw_gateway Connects to an OpenClaw gateway endpoint
Process process Executes arbitrary shell commands
HTTP http Sends webhooks to external agents

Credential ownership for sandbox targets

Local CLI adapters can run on the Paperclip host, SSH targets, or managed sandbox targets. The adapter decides which credential home is authoritative before the CLI starts:

Adapter Credential topology Which credential file wins on managed sandbox targets
codex_local Host-owns-auth for Paperclip-managed CODEX_HOME A host-owned auth.json is symlinked into the managed CODEX_HOME and uploaded to the sandbox. If a per-agent OPENAI_API_KEY is configured, Paperclip writes an API-key auth.json instead and that file wins. A login baked into the sandbox image is shadowed because Codex runs with Paperclip's uploaded CODEX_HOME.
claude_local Snapshot-owns-auth for managed remote Claude config A configured ANTHROPIC_API_KEY or CLAUDE_CODE_OAUTH_TOKEN (agent or environment env) wins over any stored login. Otherwise Paperclip uploads only sanitized settings and skill/runtime assets, and when the remote managed config has no Claude credential files it copies .credentials.json or credentials.json from the sandbox image's own $HOME/.claude, so the image's login wins.

Worked examples:

  • Codex sandbox with host ChatGPT login: the host ~/.codex/auth.json is symlinked into the managed home, then uploaded as the sandbox CODEX_HOME. Codex reads that uploaded file and does not use any auth.json already present inside the sandbox image.
  • Claude sandbox with image login: Paperclip materializes a remote CLAUDE_CONFIG_DIR, then fills missing .credentials.json / credentials.json from the sandbox image's own $HOME/.claude. The snapshot's Claude login is the credential source for the run.

Hermes local vs gateway

Use hermes_local when Paperclip should start the local hermes CLI on the same host for each heartbeat. Use hermes_gateway when Hermes is already running as an HTTP/SSE API server and Paperclip should call that server instead of spawning a process. Both type keys are stable built-ins.

The unified Hermes package owns both built-in adapters. The older @paperclipai/adapter-hermes-gateway package remains only as a deprecated compatibility shim that re-exports the gateway entrypoints for one release. New plugin overrides should target @paperclipai/hermes-paperclip-adapter and set the desired type key (hermes_local or hermes_gateway).

External (plugin) adapters

These adapters ship as standalone npm packages and are installed via the plugin system:

Adapter Package Type Key Description
Droid @henkey/droid-paperclip-adapter droid_local Runs Factory Droid locally

External Adapters

You can build and distribute adapters as standalone packages — no changes to Paperclip's source code required. External adapters are loaded at startup via the plugin system.

# Install from npm via API
curl -X POST http://localhost:3102/api/adapters \
  -d '{"packageName": "my-paperclip-adapter"}'

# Or link from a local directory
curl -X POST http://localhost:3102/api/adapters \
  -d '{"localPath": "/home/user/my-adapter"}'

See External Adapters for the full guide.

Adapter Architecture

Each adapter is a package with modules consumed by three registries:

my-adapter/
  src/
    index.ts            # Shared metadata (type, label, models)
    server/
      execute.ts        # Core execution logic
      parse.ts          # Output parsing
      test.ts           # Environment diagnostics
    ui-parser.ts        # Self-contained UI transcript parser (for external adapters)
    cli/
      format-event.ts   # Terminal output for `paperclipai run --watch`
Registry What it does Source
Server Executes agents, captures results createServerAdapter() from package root
UI Renders run transcripts, provides config forms ui-parser.js (dynamic) or static import (built-in)
CLI Formats terminal output for live watching Static import

Choosing an Adapter

  • Need a coding agent? Use claude_local, codex_local, opencode_local, hermes_local, or install droid_local as an external plugin
  • Need the richest live run feedback? Use claude_local, codex_local, or gemini_local with adapterConfig.engine set to acp when the execution environment satisfies the ACP prerequisites — see Feedback granularity
  • Need Hermes on another host or already running as a service? Use hermes_gateway
  • Need to run a script or command? Use process
  • Need to call a custom external service? Use http
  • Need something custom? Create your own adapter or build an external adapter plugin

Feedback Granularity

Adapter choice determines how much structured, live detail a run's transcript can show while the agent is still working. Every adapter's stdout is streamed to the run log and rendered live in the UI — including runs on sandbox execution targets, whose logs are tailed and delivered incrementally — but the granularity of what you see depends on the event stream the adapter emits.

Rough tiers, richest first:

  1. Native ACP engine (claude_local, codex_local, or gemini_local with engine: "acp") — full structured event stream. ACP emits a JSONL event per meaningful runtime moment: acpx.session (agent, mode, session identity), acpx.status (progress text plus context-window usage), acpx.text_delta (assistant/thinking token deltas), acpx.tool_call (tool title, call id, and status updates as the call progresses), acpx.result (stop reason summary), and acpx.error (code, message, retryability). The transcript renders these as live-updating message, thinking, tool, and status blocks, and repeated acpx.tool_call status updates fold into a single tool card instead of stacking duplicates.
  2. CLI wrappers (claude_local, codex_local, cursor, opencode_local, …). These parse each CLI's own streaming JSON output. You get assistant text, tool calls/results, and a final usage/cost summary, but granularity is limited to what the CLI prints — some emit tool progress, others only call/finish pairs.
  3. Generic adapters (process, http). Plain stdout/stderr lines with no structured transcript — you see raw output only.

Recommendation: use the native ACP engine on claude_local, codex_local, or gemini_local when the selected execution environment supports it. Rich ACP status events (including context usage) and incremental tool-call updates give the closest thing to watching the agent work locally.

UI Parser Contract

External adapters can ship a self-contained UI parser that tells the Paperclip web UI how to render their stdout. Without it, the UI uses a generic shell parser. See the UI Parser Contract for details.