Files
PaperClipAI/doc/adapter-model-audit-2026-10-02.md
9fb955e9a2 fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Codex adapter and the native runner run Codex in local hosts and
in managed sandboxes, and the model catalog lists the models an operator
can select
> - The Codex backend accepts a new model with ChatGPT sign-in only from
a recent enough Codex CLI. An older CLI fails every turn with `The
'<model>' model is not supported when using Codex with a ChatGPT
account.`
> - After #14942 added `gpt-6.1-sol`, an operator selected it for an
agent in a managed Daytona sandbox. The sandbox image still shipped
Codex 0.156.0. The connection test and runs failed with that sentence,
which reads like an account problem
> - Paperclip only checked the Codex compatibility window (`>=0.149.0
<0.161.0`), so 0.156.0 passed and the failure surfaced from the backend
without a cause
> - This pull request records the verified Codex CLI floor for each
gated model, compares the installed `codex --version` with that floor in
the environment Test and in the remote runner, and names the backend
rejection when it still happens
> - The benefit is a precise, actionable message ("gpt-6.1-sol requires
Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image
with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe

## Linked Issues or Issue Description

Refs #14942

**Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT
account error.**

**Steps to reproduce**
1. Promote a sandbox image that ships Codex CLI 0.156.0.
2. Create a Codex agent with a ChatGPT sign-in connection, select
`gpt-6.1-sol`, and run the environment Test or a task in that sandbox.

**Expected behavior**
The Test or the run tells the operator that the Codex CLI in the sandbox
is older than the model needs.

**Actual behavior**
The Test reports `codex_hello_probe_failed` with
`{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The
'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT
account."}}`. A run fails with "The selected model is not supported by
the current ChatGPT connection."

**Evidence**
- Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT
sign-in: https://github.com/openai/codex/issues/49396
- Codex 0.159.0 and 0.159.2 are accepted:
https://github.com/openai/codex/issues/49464 and
https://github.com/decolua/9router/issues/4471 (same error, fixed by
raising the client version identity from 0.155.0 to 0.159.0)
- Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model
catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0
- GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu
with ChatGPT sign-in: https://learn.chatgpt.com/docs/models

## What Changed

- `packages/adapters/codex-local/src/index.ts`: add
`minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol`
and `gpt-6-luna` → 0.157.0, with the sources above),
`parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and
`CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor
return `null`, so nothing probes them. The agent configuration doc
describes the floors.
- `packages/adapters/codex-local/src/server/cli-version.ts` (new): run
`codex --version` where the run would execute it (local, SSH, or
sandbox), compare it with the model floor, and build the
`codex_cli_version_compatible` / `codex_cli_version_incompatible`
checks. Map the backend rejection sentence to
`codex_hello_probe_model_rejected` with the detected CLI version and a
hint that separates a stale CLI from a plan that does not include the
model.
- `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test):
run the version check before the hello probe when the model has a floor.
Skip the hello probe when the CLI is too old
(`codex_hello_probe_skipped_cli_version`). When the probe still fails
with the backend sentence, report `codex_hello_probe_model_rejected`
instead of the generic `codex_hello_probe_failed`. The new codes do not
match the auth-failure patterns, so a managed connection is not
invalidated.
- `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for
remote targets, run the same version check against the shared `codex`
the ACP server spawns.
- `server/src/services/native-runtime/native-session-executor.ts`: after
the compatibility-window check, compare the remote Codex with the
configured model's floor and fail with
`runner_remote_provider_artifact_incompatible: <model> requires Codex
<floor> or newer with ChatGPT sign-in, received <version> from the
sandbox image; promote a sandbox image with Codex 0.160.0 or configure
PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex
fails this check and an npm spec is configured, the existing fallback
installs the pinned release.
- Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor,
backend rejection), ACP-lane Test (too old, compatible, no floor), and
the remote runner floor (nine version/model cases).
- `doc/adapter-model-audit-2026-10-02.md`: record the floors and the
reason.

No catalog entry, default model, saved agent configuration, pin,
lockfile, or image definition changes.

## Verification

Local (Node 25.9, pnpm 9.15.4 via corepack):

```sh
pnpm --filter @paperclipai/adapter-codex-local typecheck
pnpm --filter @paperclipai/adapter-codex-local exec vitest run            # 81 passed
pnpm --filter @paperclipai/server exec vitest run \
  src/services/native-runtime/native-session-executor.test.ts \
  src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex   # 84 passed
```

Also run locally: the full `native-session-executor.test.ts` file (533
passed) and the full Codex adapter suite (81 passed).

Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe
ignored the agent's configured env): the probe now receives the
adapter's string-valued `env` entries, the same ones
`buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override
selects the same `codex` for the Test as for the run. New test covers a
`PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on
that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck`
clean; full Codex adapter suite 503 passed (31 files).

Not run here: `pnpm -r typecheck` for `server` (the direct `tsc
--noEmit` was killed by the sandbox memory cap; the adapter package
typecheck passes and CI covers the server), `pnpm build`, browser
suites, and a live sandbox probe. No UI files changed, so token gates do
not apply.

Manual check for a reviewer: set an agent's Codex model to
`gpt-6.1-sol`, point its environment at a sandbox whose `codex
--version` prints `codex-cli 0.156.0`, and run the environment Test. The
result contains `codex_cli_version_incompatible` with `Detected Codex
CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test
contains `codex_cli_version_compatible` and runs the hello probe.

## Risks

- A floor is enforced for all authentication modes, because the runner
does not know the credential method at verification time. An OpenAI API
key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected
before launch. Paperclip pins Codex 0.160.0 everywhere, so this only
affects installs that lag the pin, and the message names the fix.
- The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release
verified to work; 0.157.0 and 0.158.0 were not verified either way. If
OpenAI accepts an older client, the floor can be lowered in one table.
- The `codex --version` probe adds one short process run to the Test
only for models with a floor.
- Rollback: revert this pull request. No data or configuration migrates.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended
thinking, tool use, and web research, operating as a Paperclip agent
through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 10:37:55 -07:00

6.6 KiB

Adapter model and harness audit: October 2, 2026

This audit follows the September 22, 2026 audit. It covers model selection in the coding-agent adapters and the shared harness pins in the runner, the provider pack, and the evaluation image. Catalog entries identify models; provider accounts and installed CLIs determine access. No agent defaults or saved model selections are migrated.

Model changes

Adapter Changes from the audit
Codex and the Codex runner catalog Add GPT-6.1 Sol (gpt-6.1-sol). Expose efforts through Ultra and Fast mode, as documented. Keep gpt-5.6-sol as the default. Follow-up (October 6, 2026): with ChatGPT sign-in the Codex backend accepts gpt-6.1-sol only from Codex CLI 0.159.0 or newer (0.156.1 and older are rejected with "The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."), and gpt-6-sol / gpt-6-luna entered the bundled catalog in 0.157.0. The adapter now records these floors; the environment Test and the remote runner compare the installed codex --version against them, so a sandbox image baked before the pin moved reports the stale CLI instead of an account error.
Claude on Bedrock Add Sonnet 5.5 (us.anthropic.claude-sonnet-5-5), using the catalog's existing US inference-profile convention. Sonnet 5.5 IDs (direct or Bedrock-qualified) get the documented xhigh and max efforts and require Claude Code 2.1.284 or later on the CLI lane.
OpenCode Add openai/gpt-6.1-sol and anthropic/claude-sonnet-5-5 to the static fallback used by remote environments. Both IDs are present in the OpenCode model registry.
Claude Code (direct) No change in this audit. #14993 (merged October 5, 2026, superseding #14816) adds Sonnet 5.5 (claude-sonnet-5-5, Claude Code 2.1.284 or later) and refreshes the Claude runtime.
Grok Build, Gemini CLI, Kimi Code, Cursor No new verified model IDs. Grok 4.7, Gemini 3.8 Flash, and Kimi K3 remain the newest documented models.

Harness changes

Harness Before After
Codex CLI (shared provider pack) 0.156.0 0.160.0
OpenCode (shared provider pack) 1.18.32 1.18.34
Grok CLI (@xai-official/grok) 1.0.41 1.0.46
Gemini CLI 0.60.0 0.62.0
Kimi Code CLI 2.0.2 2.1.1
Cursor CLI 2026.09.18-9a7762b 2026.10.01-e373342
GitHub CLI 2.101.0 2.102.0
Claude Agent SDK / Claude Code 0.3.280 / 2.1.280 0.3.286 / 2.1.286 via #14993 (merged); this branch keeps that pin
Hermes 0.19.0 0.19.0 (current)
ACPX, claude-agent-acp, codex-acp bridges 0.13.1 / 0.73.0 / 1.6.2 Unchanged; separately qualified
Pi (@earendil-works/pi-coding-agent) 0.87.1 (fleet image) Unchanged; Pi 1.0 is qualified upstream in the Pi runner stack

The remote Codex compatibility window moves to >=0.149.0 <0.161.0. The minimum stays fixed. See paperclip-runner-compatibility.md.

Sources and verification

  • OpenAI's Codex model guide lists GPT-6.1 Sol (gpt-6.1-sol) with reasoning efforts from Light to Ultra and Standard and Fast modes at launch. The bundled Codex 0.160.0 model metadata also contains gpt-6.1-sol. The same page records that gpt-5.4 and gpt-5.4-mini retired from Codex with ChatGPT sign-in on August 31, 2026, and that gpt-5.5 retires from Codex with ChatGPT sign-in on October 14, 2026. Neither retirement applies to the OpenAI API, so the picker keeps those IDs for API-key users; see "Deferred items".
  • Codex CLI releases 0.157.0 through 0.160.0 add GPT-6 Sol and Luna metadata, app-server pagination, and authoritative provider catalogs. The Linux x64 executable digest comes from the integrity-verified @openai/codex@0.160.0-linux-x64 archive.
  • Claude Sonnet 5.5 documents the Bedrock ID anthropic.claude-sonnet-5-5 and the September 28, 2026 release. The Claude Code changelog entry for 2.1.284 adds claude-sonnet-5-5, so that version is the CLI minimum. Models overview lists xhigh and max effort support for the 5.5 generation.
  • OpenCode releases 1.18.33 and 1.18.34 contain fixes only. The OpenCode model registry at models.dev lists both added provider-qualified IDs.
  • Grok CLI 1.0.46 is the latest tag on npm. xAI's model list still ends at Grok 4.7. Gemini CLI 0.62.0 and Kimi Code 2.1.1 are the latest tags on npm; Google's model catalog and Kimi's model page list no newer coding models.
  • Cursor CLI 2026.10.01-e373342 is the version the official installer resolves; the pinned archive digest comes from the versioned tarball, and the extracted binary reports that version. Cursor's catalog shows new Grok 4.7 500k and Kimi K3 rows without exact CLI IDs, so no fallback IDs were added.
  • GitHub CLI 2.102.0 contains security fixes for download commands. The pinned digest matches the release checksums.txt.
  • Hermes on PyPI remains at 0.19.0.

Deferred items

  • gpt-5.4 and gpt-5.4-mini stay in the Codex picker because the OpenAI API still serves them. A later audit can remove them once API retirement is published or operators confirm no API-key agents depend on them. gpt-5.5 needs the same review after October 14, 2026.
  • The ACP bridges (ACPX 0.19.4, claude-agent-acp 0.85.1, codex-acp 2.1.1) are newer upstream. They carry reviewed patches and qualified executable digests, so they need a separate qualification campaign rather than a pin bump. ACPX 0.15.1 and 0.17.0 also change message limits and session journal records.
  • Pi 1.0.0 is published, but the runner's Pi profile and the fleet image stay on the qualified 0.84.2 / 0.87.1 pins until the Pi 1.0 runner stack lands.
  • The native Grok runtime (native:grok 1.0.13) is bound by a qualified profile digest and is not refreshed by this audit.
  • Claude Mythos 5.1 remains invitation-only with no public selectable ID.
  • Cursor's exact IDs for Grok 4.7 500k, Kimi K3, and GLM 5.3 are not shown in the documentation and the unauthenticated CLI returns no list.