mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Codex adapter and the native runner run Codex in local hosts and in managed sandboxes, and the model catalog lists the models an operator can select > - The Codex backend accepts a new model with ChatGPT sign-in only from a recent enough Codex CLI. An older CLI fails every turn with `The '<model>' model is not supported when using Codex with a ChatGPT account.` > - After #14942 added `gpt-6.1-sol`, an operator selected it for an agent in a managed Daytona sandbox. The sandbox image still shipped Codex 0.156.0. The connection test and runs failed with that sentence, which reads like an account problem > - Paperclip only checked the Codex compatibility window (`>=0.149.0 <0.161.0`), so 0.156.0 passed and the failure surfaced from the backend without a cause > - This pull request records the verified Codex CLI floor for each gated model, compares the installed `codex --version` with that floor in the environment Test and in the remote runner, and names the backend rejection when it still happens > - The benefit is a precise, actionable message ("gpt-6.1-sol requires Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe ## Linked Issues or Issue Description Refs #14942 **Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT account error.** **Steps to reproduce** 1. Promote a sandbox image that ships Codex CLI 0.156.0. 2. Create a Codex agent with a ChatGPT sign-in connection, select `gpt-6.1-sol`, and run the environment Test or a task in that sandbox. **Expected behavior** The Test or the run tells the operator that the Codex CLI in the sandbox is older than the model needs. **Actual behavior** The Test reports `codex_hello_probe_failed` with `{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."}}`. A run fails with "The selected model is not supported by the current ChatGPT connection." **Evidence** - Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT sign-in: https://github.com/openai/codex/issues/49396 - Codex 0.159.0 and 0.159.2 are accepted: https://github.com/openai/codex/issues/49464 and https://github.com/decolua/9router/issues/4471 (same error, fixed by raising the client version identity from 0.155.0 to 0.159.0) - Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0 - GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu with ChatGPT sign-in: https://learn.chatgpt.com/docs/models ## What Changed - `packages/adapters/codex-local/src/index.ts`: add `minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol` and `gpt-6-luna` → 0.157.0, with the sources above), `parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and `CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor return `null`, so nothing probes them. The agent configuration doc describes the floors. - `packages/adapters/codex-local/src/server/cli-version.ts` (new): run `codex --version` where the run would execute it (local, SSH, or sandbox), compare it with the model floor, and build the `codex_cli_version_compatible` / `codex_cli_version_incompatible` checks. Map the backend rejection sentence to `codex_hello_probe_model_rejected` with the detected CLI version and a hint that separates a stale CLI from a plan that does not include the model. - `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test): run the version check before the hello probe when the model has a floor. Skip the hello probe when the CLI is too old (`codex_hello_probe_skipped_cli_version`). When the probe still fails with the backend sentence, report `codex_hello_probe_model_rejected` instead of the generic `codex_hello_probe_failed`. The new codes do not match the auth-failure patterns, so a managed connection is not invalidated. - `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for remote targets, run the same version check against the shared `codex` the ACP server spawns. - `server/src/services/native-runtime/native-session-executor.ts`: after the compatibility-window check, compare the remote Codex with the configured model's floor and fail with `runner_remote_provider_artifact_incompatible: <model> requires Codex <floor> or newer with ChatGPT sign-in, received <version> from the sandbox image; promote a sandbox image with Codex 0.160.0 or configure PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex fails this check and an npm spec is configured, the existing fallback installs the pinned release. - Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor, backend rejection), ACP-lane Test (too old, compatible, no floor), and the remote runner floor (nine version/model cases). - `doc/adapter-model-audit-2026-10-02.md`: record the floors and the reason. No catalog entry, default model, saved agent configuration, pin, lockfile, or image definition changes. ## Verification Local (Node 25.9, pnpm 9.15.4 via corepack): ```sh pnpm --filter @paperclipai/adapter-codex-local typecheck pnpm --filter @paperclipai/adapter-codex-local exec vitest run # 81 passed pnpm --filter @paperclipai/server exec vitest run \ src/services/native-runtime/native-session-executor.test.ts \ src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex # 84 passed ``` Also run locally: the full `native-session-executor.test.ts` file (533 passed) and the full Codex adapter suite (81 passed). Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe ignored the agent's configured env): the probe now receives the adapter's string-valued `env` entries, the same ones `buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override selects the same `codex` for the Test as for the run. New test covers a `PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck` clean; full Codex adapter suite 503 passed (31 files). Not run here: `pnpm -r typecheck` for `server` (the direct `tsc --noEmit` was killed by the sandbox memory cap; the adapter package typecheck passes and CI covers the server), `pnpm build`, browser suites, and a live sandbox probe. No UI files changed, so token gates do not apply. Manual check for a reviewer: set an agent's Codex model to `gpt-6.1-sol`, point its environment at a sandbox whose `codex --version` prints `codex-cli 0.156.0`, and run the environment Test. The result contains `codex_cli_version_incompatible` with `Detected Codex CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test contains `codex_cli_version_compatible` and runs the hello probe. ## Risks - A floor is enforced for all authentication modes, because the runner does not know the credential method at verification time. An OpenAI API key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected before launch. Paperclip pins Codex 0.160.0 everywhere, so this only affects installs that lag the pin, and the message names the fix. - The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release verified to work; 0.157.0 and 0.158.0 were not verified either way. If OpenAI accepts an older client, the floor can be lowered in one table. - The `codex --version` probe adds one short process run to the Test only for models with a floor. - Rollback: revert this pull request. No data or configuration migrates. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended thinking, tool use, and web research, operating as a Paperclip agent through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Bender (Fable) <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
6.6 KiB
6.6 KiB
Adapter model and harness audit: October 2, 2026
This audit follows the September 22, 2026 audit. It covers model selection in the coding-agent adapters and the shared harness pins in the runner, the provider pack, and the evaluation image. Catalog entries identify models; provider accounts and installed CLIs determine access. No agent defaults or saved model selections are migrated.
Model changes
| Adapter | Changes from the audit |
|---|---|
| Codex and the Codex runner catalog | Add GPT-6.1 Sol (gpt-6.1-sol). Expose efforts through Ultra and Fast mode, as documented. Keep gpt-5.6-sol as the default. Follow-up (October 6, 2026): with ChatGPT sign-in the Codex backend accepts gpt-6.1-sol only from Codex CLI 0.159.0 or newer (0.156.1 and older are rejected with "The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account."), and gpt-6-sol / gpt-6-luna entered the bundled catalog in 0.157.0. The adapter now records these floors; the environment Test and the remote runner compare the installed codex --version against them, so a sandbox image baked before the pin moved reports the stale CLI instead of an account error. |
| Claude on Bedrock | Add Sonnet 5.5 (us.anthropic.claude-sonnet-5-5), using the catalog's existing US inference-profile convention. Sonnet 5.5 IDs (direct or Bedrock-qualified) get the documented xhigh and max efforts and require Claude Code 2.1.284 or later on the CLI lane. |
| OpenCode | Add openai/gpt-6.1-sol and anthropic/claude-sonnet-5-5 to the static fallback used by remote environments. Both IDs are present in the OpenCode model registry. |
| Claude Code (direct) | No change in this audit. #14993 (merged October 5, 2026, superseding #14816) adds Sonnet 5.5 (claude-sonnet-5-5, Claude Code 2.1.284 or later) and refreshes the Claude runtime. |
| Grok Build, Gemini CLI, Kimi Code, Cursor | No new verified model IDs. Grok 4.7, Gemini 3.8 Flash, and Kimi K3 remain the newest documented models. |
Harness changes
| Harness | Before | After |
|---|---|---|
| Codex CLI (shared provider pack) | 0.156.0 | 0.160.0 |
| OpenCode (shared provider pack) | 1.18.32 | 1.18.34 |
Grok CLI (@xai-official/grok) |
1.0.41 | 1.0.46 |
| Gemini CLI | 0.60.0 | 0.62.0 |
| Kimi Code CLI | 2.0.2 | 2.1.1 |
| Cursor CLI | 2026.09.18-9a7762b | 2026.10.01-e373342 |
| GitHub CLI | 2.101.0 | 2.102.0 |
| Claude Agent SDK / Claude Code | 0.3.280 / 2.1.280 | 0.3.286 / 2.1.286 via #14993 (merged); this branch keeps that pin |
| Hermes | 0.19.0 | 0.19.0 (current) |
ACPX, claude-agent-acp, codex-acp bridges |
0.13.1 / 0.73.0 / 1.6.2 | Unchanged; separately qualified |
Pi (@earendil-works/pi-coding-agent) |
0.87.1 (fleet image) | Unchanged; Pi 1.0 is qualified upstream in the Pi runner stack |
The remote Codex compatibility window moves to >=0.149.0 <0.161.0. The
minimum stays fixed. See
paperclip-runner-compatibility.md.
Sources and verification
- OpenAI's Codex model guide lists
GPT-6.1 Sol (
gpt-6.1-sol) with reasoning efforts from Light to Ultra and Standard and Fast modes at launch. The bundled Codex 0.160.0 model metadata also containsgpt-6.1-sol. The same page records thatgpt-5.4andgpt-5.4-miniretired from Codex with ChatGPT sign-in on August 31, 2026, and thatgpt-5.5retires from Codex with ChatGPT sign-in on October 14, 2026. Neither retirement applies to the OpenAI API, so the picker keeps those IDs for API-key users; see "Deferred items". - Codex CLI releases 0.157.0
through 0.160.0 add GPT-6 Sol and Luna metadata, app-server pagination, and
authoritative provider catalogs. The Linux x64 executable digest comes from
the integrity-verified
@openai/codex@0.160.0-linux-x64archive. - Claude Sonnet 5.5
documents the Bedrock ID
anthropic.claude-sonnet-5-5and the September 28, 2026 release. The Claude Code changelog entry for 2.1.284 addsclaude-sonnet-5-5, so that version is the CLI minimum. Models overview listsxhighandmaxeffort support for the 5.5 generation. - OpenCode releases 1.18.33 and 1.18.34 contain fixes only. The OpenCode model registry at models.dev lists both added provider-qualified IDs.
- Grok CLI 1.0.46 is the
latesttag on npm. xAI's model list still ends at Grok 4.7. Gemini CLI 0.62.0 and Kimi Code 2.1.1 are thelatesttags on npm; Google's model catalog and Kimi's model page list no newer coding models. - Cursor CLI 2026.10.01-e373342 is the version the official installer resolves; the pinned archive digest comes from the versioned tarball, and the extracted binary reports that version. Cursor's catalog shows new Grok 4.7 500k and Kimi K3 rows without exact CLI IDs, so no fallback IDs were added.
- GitHub CLI 2.102.0
contains security fixes for download commands. The pinned digest matches the
release
checksums.txt. - Hermes on PyPI remains at 0.19.0.
Deferred items
gpt-5.4andgpt-5.4-ministay in the Codex picker because the OpenAI API still serves them. A later audit can remove them once API retirement is published or operators confirm no API-key agents depend on them.gpt-5.5needs the same review after October 14, 2026.- The ACP bridges (ACPX 0.19.4,
claude-agent-acp0.85.1,codex-acp2.1.1) are newer upstream. They carry reviewed patches and qualified executable digests, so they need a separate qualification campaign rather than a pin bump. ACPX 0.15.1 and 0.17.0 also change message limits and session journal records. - Pi 1.0.0 is published, but the runner's Pi profile and the fleet image stay on the qualified 0.84.2 / 0.87.1 pins until the Pi 1.0 runner stack lands.
- The native Grok runtime (
native:grok1.0.13) is bound by a qualified profile digest and is not refreshed by this audit. - Claude Mythos 5.1 remains invitation-only with no public selectable ID.
- Cursor's exact IDs for Grok 4.7 500k, Kimi K3, and GLM 5.3 are not shown in the documentation and the unauthenticated CLI returns no list.