7 Commits
Author SHA1 Message Date
6c36c07a4f feat(adapters): add GPT-6.1 Sol and refresh shared coding harness pins (#14942)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents run through coding-agent adapters and the native runner. Both
use the same installed provider CLIs, model catalogs, and reasoning
controls.
> - OpenAI released GPT-6.1 Sol (`gpt-6.1-sol`) in Codex. Anthropic
released Claude Sonnet 5.5. The static Codex, Bedrock, and OpenCode
catalogs do not list these IDs.
> - The shared provider pack pins Codex 0.156.0 and OpenCode 1.18.32.
The evaluation image pins older Grok, Gemini, Kimi, Cursor, and GitHub
CLI releases. Codex 0.156.0 has no bundled metadata for GPT-6.1 Sol.
> - A model entry without a current harness, or a harness pin without
its runner integrity checks, fails at run time.
> - This pull request adds the verified model IDs and moves the harness
pins, executable digests, controller checks, and image pins together.
> - The benefit is that operators can select the current models, and the
native and local adapters share one current CLI installation.

## Linked Issues or Issue Description

Refs #13829 and #13838 (the September 22, 2026 model and harness
refresh). Related pull requests: #14993 (merged October 5, 2026,
superseding #14816) added the direct Claude Sonnet 5.5 entry and
refreshed the Claude runtime to Agent SDK 0.3.286 / Claude Code 2.1.286.
This pull request does not change the Claude runtime or the direct
Claude model list; it keeps the #14993 pins and adds only the Bedrock
Sonnet 5.5 ID. After #14993 merged, this branch was rebased onto
`master` (October 5, 2026). The six overlapping pin regions
(`docker/daytona-runner/Dockerfile`, `docker/daytona-runner/README.md`,
`package.json`, `pnpm-workspace.yaml`,
`packages/adapters/claude-local/src/index.test.ts`,
`packages/paperclip-runner/src/backends/native-backend-factory.test.ts`)
were resolved by keeping this pull request's Codex 0.160.0 and OpenCode
1.18.34 pins next to #14993's Claude 0.3.286 / 2.1.286 pins, taking the
union of the Sonnet 5.5 model IDs in the Claude test, and merging both
README paragraphs. The Sonnet 5.5 effort and CLI-gate lines in the
Claude adapter were identical in both pull requests and merged without a
diff. #14917 and #14918 reordered the Claude and Codex model lists
earlier; the new entries sit where those ordering rules put them.

Sources checked on 2026-10-02:

- [OpenAI Codex models](https://learn.chatgpt.com/docs/models): GPT-6.1
Sol uses `gpt-6.1-sol`, supports reasoning efforts from Light to Ultra,
and has Standard and Fast modes at launch. The page also records that
`gpt-5.4` and `gpt-5.4-mini` retired from Codex with ChatGPT sign-in on
August 31, 2026, and that `gpt-5.5` retires on October 14, 2026. Neither
retirement applies to the OpenAI API.
- [Codex CLI releases](https://github.com/openai/codex/releases) 0.157.0
through 0.160.0. The bundled model metadata in the 0.160.0 Linux binary
contains `gpt-6.1-sol`.
- [Claude Sonnet
5.5](https://platform.claude.com/docs/en/models/sonnet-5-5/overview):
Bedrock ID `anthropic.claude-sonnet-5-5`, released September 28, 2026.
- [OpenCode releases](https://github.com/anomalyco/opencode/releases)
1.18.33 and 1.18.34 (fixes only). The OpenCode model registry lists both
added provider-qualified IDs.
- npm `latest` tags for `@xai-official/grok` 1.0.46,
`@google/gemini-cli` 0.62.0, and `@moonshot-ai/kimi-code` 2.1.1.
[xAI](https://docs.x.ai/docs/models),
[Google](https://ai.google.dev/gemini-api/docs/models), and
[Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html) list no
newer coding models.
- Cursor CLI 2026.10.01-e373342 is the version the official installer
resolves. The pinned digest is the SHA-256 of the versioned Linux x64
archive.
- [GitHub CLI 2.102.0](https://github.com/cli/cli/releases/tag/v2.102.0)
(security fixes). The pinned digest matches the release `checksums.txt`.

## What Changed

- Codex adapter: add `gpt-6.1-sol` to the model list, the Fast mode
list, and the Ultra effort set. It is the first entry: #14918 orders the
list newest version first, and its description notes the ChatGPT app
lists GPT-6.1 Sol first. Update the adapter documentation text.
- Claude adapter: add `us.anthropic.claude-sonnet-5-5` (Bedrock Sonnet
5.5) to the Bedrock catalog in the newest-Sonnet slot after Opus 5.5;
`us.anthropic.claude-sonnet-5` moves into the older-Sonnet group,
matching what `sortClaudeModels` from #14917 produces at runtime. Any
Sonnet 5.5 ID (direct or Bedrock-qualified) now gets the documented
`xhigh` and `max` efforts and requires Claude Code 2.1.284 or later on
the CLI lane (the Claude Code changelog entry for 2.1.284 adds
`claude-sonnet-5-5`). These two lines are identical to the ones #14993
merged, so the branch carries no diff for them.
- OpenCode adapter: add `openai/gpt-6.1-sol` and
`anthropic/claude-sonnet-5-5` to the static fallback catalog.
- Codex runtime pin 0.156.0 → 0.160.0 in the root and workspace
overrides, the runner package, the Codex ACP package patch, the
qualified ACPX profiles, the Linux x64 executable digest, the Rust
provider backend and its tests, the provider-pack manifest pins, the
remote controller pins, the sandbox npm install spec, and the opt-in
qualification scripts.
- Remote Codex compatibility window: upper bound 0.157.0 → 0.161.0. The
minimum stays at 0.149.0.
- OpenCode runtime pin 1.18.32 → 1.18.34 in the runner package, the
materialization script, the server and Rust qualified versions, the eval
and live-session labels, fixtures, and the configuration label.
- Evaluation image (`docker/daytona-runner/Dockerfile`): Grok CLI
1.0.46, Gemini CLI 0.62.0, Kimi Code 2.1.1, Cursor CLI
2026.10.01-e373342 with its digest, GitHub CLI 2.102.0 with its digest,
Codex and OpenCode version probes, and the refreshed lockfile digest.
The Claude Code 2.1.286 probe comes from #14993 and is unchanged here.
- `pnpm-lock.yaml` is not part of this pull request. The repository's
pull request gate rejects lockfile edits, and the refresh bot
regenerates the lockfile on master (the same flow #13838 used). The
Dockerfile `PAPERCLIP_RUNNER_LOCK_SHA256` default is the digest of the
lockfile that `pnpm install --resolution-only --ignore-scripts
--no-frozen-lockfile` (the refresh workflow's command) produces for the
combined pins on the rebased branch (`e1856797…`); that lockfile differs
from master only in the `@openai/codex` 0.160.0 platform packages, the
`@anthropic-ai/claude-agent-sdk` 0.3.286 override that #14993 introduced
(the open refresh-bot pull request #14872 carries that part),
`opencode-ai` 1.18.34 with its Linux x64 baseline, and the `codex-acp`
patch hash.
- Documentation: runner README, runner compatibility doc, environment
variable example, and a new `doc/adapter-model-audit-2026-10-02.md` with
sources and deferred items.
- Tests: Codex adapter catalog, server adapter models, Codex
compatibility window, native session executor pins, runner package
contract, OpenCode materialization, and UI effort options.

Unchanged on purpose: Claude Agent SDK 0.3.286 / Claude Code 2.1.286
(already on `master` from #14993), ACP bridges (`acpx` 0.13.1,
`claude-agent-acp` 0.73.0, `codex-acp` 1.6.2; newer upstream releases
need a separate qualification), the native Grok runtime 1.0.13, Pi
0.84.2 / 0.87.1 (the Pi 1.0 runner stack covers it), and Hermes 0.19.0
(current). `gpt-5.4` and `gpt-5.4-mini` stay in the picker because the
OpenAI API still serves them.

## Verification

Run on Linux x64 with Node 25.9.0 and pnpm 9.15.4 after `pnpm install
--no-frozen-lockfile` (the refreshed lockfile stays local; see above).
The results below were re-run on the rebased head (October 5, 2026) for
the suites the conflict resolution touches; the other rows are from the
original run and are covered by CI on every push:

- Rebased head: `packages/adapters/codex-local` 482 passed;
`packages/adapters/claude-local` 340 passed, 4 failed (`execute.remote`,
`test.probe`, `execute.acp-fallback`, `acp` spawn/env-hardening cases
that fail identically on unchanged `master` in this host environment);
`server` adapter-models + codex-runtime-compatibility +
native-session-executor + adapter-registry 607 passed, 1 failed (the
same adapter-registry override-pause case as before, also failing on
`master` here); `packages/paperclip-runner` native-backend-factory +
qualified-profiles 36 passed; `ui` codex-reasoning-effort +
config-fields + model-utils 19 passed. Rust, full typecheck, build, and
the Docker image are left to CI as before.

- `vitest run` in `packages/adapters/codex-local`: 13 passed. `vitest
run` in `packages/adapters/claude-local` (whole package, including the
new Sonnet 5.5 gate and effort tests): see the latest CI run and the
comment below. `vitest run` in `packages/adapters/opencode-local`: 48
passed, 1 failed (`runtime-config.test.ts` reads the host
`PAPERCLIP_OPENCODE_PROVIDERS` variable; it fails the same way on the
unchanged base).
- `vitest run src/__tests__/adapter-models.test.ts
src/services/native-runtime/codex-runtime-compatibility.test.ts
src/__tests__/adapter-registry.test.ts` in `server`: 84 passed, 1 failed
(`adapter-registry.test.ts` override pause test; it fails the same way
on the unchanged base).
- `vitest run` in `ui` for `codex-reasoning-effort`,
`agent-setup-fields`, `config-fields`, and `ComposerRunSettingsPicker`:
25 passed.
- `node --test test/acpx-codex-package-contract.test.mjs
scripts/materialize-opencode-binary.test.mjs
scripts/runner-protocol-eval-campaign.test.mjs` in
`packages/paperclip-runner`: 23 passed. The package contract test
verifies the installed Codex ACP executable digest and the 0.160.0 patch
pin.
- `vitest run src/drivers/acpx src/backends src/drivers/opencode
src/live/live-session.test.ts` in `packages/paperclip-runner`: 626
passed, 5 failed, 1 skipped. The 5 failures
(`installation-integrity.test.ts` `/proc/self/fd` module loading and one
OpenCode answer-selection test) also fail on the unchanged base under
Node 25; Linux CI runs Node 24.
- `pnpm run test:opencode:qualification` in `packages/paperclip-runner`
against the installed OpenCode 1.18.34 executable: passed.
- `codex --version` from the installed pack prints `codex-cli 0.160.0`.
The Linux x64 executable digest `12eb3e81…652aad` was computed from the
`@openai/codex@0.160.0-linux-x64` archive after checking its registry
`dist.integrity`.
- `pnpm check:token-gates`: all gates clean.
- `pnpm run typecheck:typescript` in `packages/paperclip-runner`:
passed. Package typechecks ran one at a time; see the comment below for
the server and UI results.

Not run here, and needed from CI:

- Rust tests and `pnpm -r typecheck` / `pnpm build` for the server (no
`cargo` in this environment; the server typecheck prepares the runner
vendor build).
- The Docker evaluation image build and the real-binary Codex startup
and session-resume probes (no Docker; the probes need the compiled
`paperclip-runnerd`). The trusted CI runner workflow covers them.
- Authenticated inference with any new model. This change is metadata
and startup validation only.

## Risks

- Codex 0.160.0 changes the bundled model catalog and app-server
behaviour (authoritative provider catalogs, incremental running-turn
tracking). The patched `codex-acp` 1.6.2 bridge is unchanged and
declares `^0.148.0`; it worked with 0.156.0 under the same override. If
CI probes show a protocol change, the pin can return to 0.156.0 by
reverting this pull request.
- The compatibility window upper bound moves to `<0.161.0`. Remote
images with Codex 0.157 to 0.160 become accepted. Older images stay
accepted down to 0.149.0.
- Until the refresh bot lands the regenerated lockfile on master, the
Dockerfile lockfile digest default does not match the committed
lockfile. The trusted CI workflow computes the digest from its own
resolution at build time, so this affects only a local build that passes
no digest.
- Existing saved model selections and effort settings are not changed.
Agents on `gpt-5.4` or `gpt-5.5` with ChatGPT sign-in need a model
change before the OpenAI retirement dates; that is documented, not
enforced.
- Rollout order: deploy the controller and runner from this change
before promoting a sandbox image that carries these pins. Older
controllers reject the new provider-pack pins.

## Model Used

- Claude Fable 5.1 (Anthropic, model ID `claude-fable-5-1`), 1M context
window, adaptive thinking, tool use. The model ran as a Paperclip agent
through the Claude Code harness, performed the web research, edited the
code, and ran the tests listed above.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 13:59:08 -07:00
Devin FoleyandPaperclip be6f49a425 feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path

> - Paperclip runs agents through local adapters and the native runner.
> - Both paths must use the same installed provider CLI.
> - New models require current harness releases.
> - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and
OpenCode 1.18.29.
> - Changing the image alone would fail the runner's exact version and
executable checks.
> - This pull request updates those dependencies, integrity checks,
controller checks, and image pins together.
> - Shared installations can then run the current models without a
task-time download.

## Linked Issues or Issue Description

Refs #13829, which updates model choices and reasoning controls.
Searches found no open PR that updates these runtime pins.

**Current behavior**

The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run
Opus 5.5, which requires 2.1.280. Remote controllers reject provider
packs whose versions differ from their declared pins.

**Proposed behavior**

Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and
OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge
patches and one shared CLI installation per provider.

**Reason and benefit**

Current harnesses support the new model IDs while preserving executable
verification and remote provider-pack compatibility checks.

## What Changed

- Update dependency overrides, the Codex ACP package patch, runtime
profiles, and remote controller pins.
- Verify the new Claude Linux x64 and macOS arm64/x64 executables and
Codex Linux x64 executable against integrity-verified npm archives.
- Refresh OpenCode version checks, fixtures, and the runner
configuration label.
- Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI
pins and archive hashes. Hermes remains current at 0.19.0.
- Refresh the build-time lock digest from clean pnpm 9.15.4 resolution.
Leave lockfile commits to repository automation.
- Document model compatibility and the separation between CLI runtimes
and patched ACP bridges.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- Rust workspace release tests passed.
- Package/patch and OpenCode binary-materialization contract tests: 11
passed.
- Real Codex 0.156.0 startup-ownership and paginated session-resume
probes passed with isolated synthetic homes and no model turn.
- Codex app-server `thread/start` preserved `gpt-6-sol` and
`gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in
catalog does not include those account-served entries.
- Installed Claude integrity probes passed for `claude-opus-5-5` and
`claude-fable-5-1`.
- `pnpm --filter @paperclipai/paperclip-runner
test:opencode:qualification` passed with the actual OpenCode 1.18.32
executable under Node 24 and Node 25. The loopback provider exercise
covers health/version, session creation/read/delete, SSE, and a
completed async prompt.
- `pnpm check:token-gates` passed.
- The targeted runner suite passed 130 tests. Three macOS failures in
snapshot module lookup and OpenCode final-message selection also
reproduce on the unchanged base; Linux CI will provide the platform
check.
- [Final Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399):
all gates passed. Four jobs needed one retry after their CI workers
received shutdown signals. The PR has 55 successful checks, two skipped
checks, Greptile 5/5, and no unresolved review threads.
- Changed runner configuration UI tests: 5 passed.
- Full macOS `pnpm test:run` reached 13,094 passing server tests, 84
skipped, and 18 failures before the wrapper stopped. Failures involved
skill-cache publication permissions, missing bundled connector skills in
the worktree, and a conversation-reset timing case. The 10 cache
permission failures reproduce on the unchanged base; both
conversation-reset cases passed on a targeted retry. The wrapper did not
reach its later workspace/serialized groups locally; Linux CI covers
those groups.
- The local Docker daemon did not respond, so no local Docker build was
run. No billable model requests were made.

## Risks

- Deploy the matching controller and provider pack together. Older
controllers enforce their previous exact pins.
- Current upstream CLIs can change behavior. Existing protocol tests and
isolated real Codex probes cover the integration boundaries;
authenticated model inference is not part of these checks.
- ACP bridge package versions and executable digests stay unchanged
because their executable bytes are unchanged. Only the underlying
CLI/SDK dependencies move.
- No schema migration. Revert the runtime and image pins together to
roll back.

## Model Used

OpenAI GPT-6 via Codex, with repository tools, code execution, and web
research. The exact serving model ID and context window were not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass for the changed surfaces
and real-executable probes; full macOS-suite limitations are listed
above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 17:02:29 -07:00
DottaandPaperclip 2991a59b17 fix(adapters): prevent engine fallback and preserve usable runtime defaults (#13105)
## Thinking Path

> - Paperclip manages agents that must write work and report task
outcomes through its API.
> - Local adapters select an execution engine and its permission
settings.
> - A higher ACP Node requirement can make an unchanged installation
lose access to its default engine.
> - The adapter then silently selects CLI, which can change permissions
and block API access.
> - This pull request keeps the engine choice fixed and reports missing
prerequisites before work starts.
> - It also gives explicit Codex CLI runs usable defaults and keeps
managed services on a supported Node runtime.

## Linked Issues or Issue Description

Refs #12215. Related changes: #11792 raised the Node requirement; #13094
addressed separate runner networking behavior. This change fixes the
engine-selection and managed-launcher paths.

**What happened?**

An unchanged agent could switch from ACP to CLI after an upgrade. Codex
CLI then used read-only permissions with networking disabled. The run
could finish without updating its task. Repeated recovery attempts used
the same unavailable setup. Managed updates also skipped the Node check
and did not refresh old launchers.

**Expected behavior**

An unavailable engine must fail with a clear setup error. It must not
silently select another engine. Explicit CLI runs must be able to write
workspace files and call the API unless the operator configures stricter
settings. Managed updates must validate Node and keep child tools on
that runtime.

**Steps to reproduce**

1. Run an ACP-default agent under Node 22 after the ACP minimum rises to
24.11.
2. Leave the engine unset and disable the approval/sandbox bypass.
3. Observe the old adapter select CLI and fail to write task disposition
through the API.
4. Start a managed service with an old launcher and a supervisor PATH
that selects a different Node for child tools.

## What Changed

- Remove automatic engine fallback for Codex, Claude, Gemini, and Kimi.
Check prerequisites for default and explicit ACP selections.
- Return a configuration error with proof that provider work did not
start. Stop automatic continuation retries for this error.
- Enable Codex ACP workspace networking at the actual turn boundary.
Upstream mode presets otherwise force it off even when config.toml
enables it. Preserve explicit network denial and read-only mode.
- Set workspace-write and network access defaults for explicit Codex CLI
runs. Preserve explicit sandbox modes, profiles, and network
restrictions.
- Pin the validated Node directory in managed launcher PATH. Refresh
legacy launchers during installs and npm/Git updates.
- Reject updates on unsupported Node. Keep update checks, dry runs, and
rollback available.
- Synchronize the qualified Codex ACP executable identity across server,
TypeScript runner, Rust runner, and provider-pack launch paths.
- Add regression tests and update engine and installation documentation.

## Verification

- [Full CI passed on the final
head](https://github.com/paperclipai/paperclip/actions/runs/34387099695):
typecheck, build/native runner verification, all general and serialized
test shards, all browser shards, release registry, canary dry run, and
policy checks.
- Greptile: 5/5 on `2c1d6e2815830a5cd39e36c8a082cc0c4441b6c0`, with no
unresolved review findings. Security gates are green.
- Full workspace typecheck and build also passed locally. The final
deployed Linux build passed.
- Full Codex, Claude, Gemini, and Kimi source test suites: 804 passed, 2
skipped. Installer, updater, and launcher tests: 47 passed. Installed
ACP turn-boundary tests: 3 passed. ACP packaging tests: 14 passed.
Focused recovery classification tests also passed.
- Real Linux Codex CLI runs, both fresh and resumed, wrote a workspace
file and reached the control-plane health API with the new defaults.
- Explicit read-only and network-disabled control probes retained those
restrictions.
- A real ACP run on the final deployed Linux build wrote a file and
reached the control-plane API with HTTP 200, without engine fallback.
The same probe failed DNS before the turn-policy patch.
- Executable-identity and installed-policy contracts: 12 passed.
Affected native server tests: 197 passed. Runner factory tests: 21
passed. Rust qualification and native provider integration tests: 11
passed.
- Deployed the production changes to a Linux service on Node 24.20 after
a verified database backup. Health, bootstrap readiness, static UI,
executable/cwd identity, and guarded restart checks passed. The restart
lost no runs.
- Corrected stale Kimi skill-default and Gemini remote-archive fixtures;
both suites pass.

## Risks

- Default or legacy auto engine settings now fail when ACP is
unavailable. Operators who intend to use CLI must select it explicitly.
- Codex CLI now permits workspace writes and networking by default, and
ACP workspace-write turns permit networking by default. Explicit
operator sandbox settings remain authoritative.
- Old managed launchers keep their pinned Node until they are
reinstalled under a supported runtime. An old updater cannot repair
itself; the documentation gives the current installer command.
- Custom service wrappers and global/source installations must configure
their runtime PATH. No database migration is required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
shell execution, and test tools. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 13:27:24 -05:00
DottaandPaperclip 7ed122911b Add end-to-end session goals to Paperclip Runner
Add capability-aware slash-goal controls, durable provider goal state, PRP v2 negotiation, autonomous goal execution, and safe local session recovery. Integrate with current master, preserve provider session identity, and verify the browser goal/chat/replacement/clear workflow and unsupported-agent rejection.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-08 16:18:47 -05:00
DottaandPaperclip f6a211479f fix: share current CLI runtimes across sandbox adapters (#12994)
## Thinking Path

- Paperclip Runner needs its runtime preinstalled for fast sandbox
startup.
- Native and local adapters should launch one current CLI installation
per provider.
- An older global copy can shadow that installation, and exact native
compatibility pins must match it.
- Update the qualified releases and binary digests, expose shared CLI
entrypoints from the provider pack, and prefer the image-owned bin
directory.
- Keep dependency installation in the image build; task startup only
discovers, links, and verifies artifacts.

## Linked Issues or Issue Description

**What happened?**
Remote native startup rejected a stale global Codex, while CLI-only
images lacked runnerd entirely.

**Expected behavior**
An image-baked runtime starts without uploading binaries or installing
packages. All adapters share the same current provider CLI.

**Steps to reproduce**
Start a native remote task with the old global Codex and the updated
runtime available only under `/opt/paperclip-runner/bin`.

**Paperclip version or commit**
Discovery behavior at `54a99d884`.

**Deployment mode**
Docker with a remote sandbox.

## What Changed

- Prefer `/opt/paperclip-runner/bin`, then the user's local bin
directory, then PATH. Existing metadata and version validation remains
in force.
- Qualify Codex 0.153.4, OpenCode 1.18.29, and Claude SDK 0.3.263 / CLI
2.1.263. Update binary digests, TypeScript/Rust checks, registry
defaults, and the displayed OpenCode version together.
- Share Codex and Claude's native executable with the ACP bridges
through exact dependency overrides. Preserve the separately qualified
ACP bridge implementations and their security patches.
- Expose shared provider-pack CLI launchers; fail the pack build if
Codex ACP resolves a separate Codex installation. Update the eval
image's other agent CLIs to current stable releases and remove duplicate
global provider installs.
- Document the single-current-CLI policy in source comments and
development guidance. Latest stable releases are resolved at
review/build preparation and pinned; task startup never auto-updates.

## Verification

- Native-session and adapter-registry suites: 158 tests passed.
- Provider suites: 88 tests passed, 7 Linux-only checks skipped on
macOS. One existing macOS temporary-path alias assertion passed when
rerun with canonical `TMPDIR=/private/tmp`.
- Package-contract and OpenCode materialization tests: 11 passed.
- Full typecheck, build, and token gates passed. Rust
native-provider/recovery tests: 19 passed.
- Broad local suite: 5,974 passed, 23 failed, 41 skipped. Failures are
in unchanged macOS workspace/path/port and connection suites; focused
runtime tests pass. All latest-head Linux PR checks passed, including
the full test shards, typecheck, build, runner verification, browser
suites, and canary dry run.
- The standalone fleet image built with one current provider CLI each
and passed native Codex/Claude binary-integrity checks. A disposable
Daytona sandbox reported ready in 798 ms; its baked runner completed an
API-key `gpt-5.6-luna` turn in 2,430 ms and returned the expected marker
with a usage receipt. No runtime artifacts were uploaded or installed.
- The normal shared `codex exec` entrypoint also completed an API-key
`gpt-5.6-luna` turn in 2,321 ms.
- Both image builds verify the complete generated lockfile against a
reviewed SHA-256 before package installation or lifecycle execution.
Root lockfile changes remain CI-owned. Merge and rollout remain on hold
for operator review.

## Risks

- Updating provider CLIs changes their behavior for all adapters;
version probes and live native smoke testing are required before image
promotion.
- The image-owned directory takes precedence. Its entries must launch
the same shared CLI as the global PATH, not a private older/newer copy.
- Application qualification pins and the deployed image must move
together. No startup fallback installation is added.
- No schema or authentication-policy changes.

## Model Used

OpenAI GPT-6 (Codex). The session does not expose a more specific model
ID or context-window size. Used reasoning, repository inspection, code
execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-07 10:09:29 -05:00
Dotta af3023f1e3 fix(runner): repair paid provider startup paths (#12769)
## Thinking Path

> - Paperclip manages AI agents that perform work.
> - Paperclip Runner connects durable task runs to local provider
processes.
> - The full-stack paid matrix exposed failures after the runner
integrity repair.
> - Verified JavaScript entrypoints lost their relative module graph
when Linux executed them through descriptor paths.
> - Returned provider startup errors also remained pending and became
indeterminate after recovery.
> - Sparse Codex tool lifecycle events lost the `write_document`
identity before task transcript projection.
> - This pull request repairs those three boundaries and makes the
structured-question fixture deterministic.
> - The benefit is repeatable provider startup, exact failure replay,
and correct inline Plan placement.

## Linked Issues or Issue Description

Refs #12721 and #12700.

**What happened?**

The paid runner matrix failed ACPX and OpenCode startup before provider
session creation. The runner journal then replaced the original startup
error with an indeterminate recovery result. Native Codex saved a Plan
but rendered it only as a fallback card. A legacy Claude waiting reply
could also echo the reserved terminal marker before the answer arrived.

**Expected behavior**

Verified JavaScript providers must start from immutable
descriptor-backed artifacts. Returned startup failures must persist as
terminal failed command results. Native tool lifecycle updates must
preserve the `write_document` boundary. Pre-answer fixture output must
not contain the reserved terminal marker.

**Steps to reproduce**

1. Run the local provider cells in the Runner Full-Stack E2E workflow.
2. Observe ACPX and OpenCode fail during `session.open` before provider
execution.
3. Observe recovery report `execution_indeterminate` instead of the
original startup error.
4. Run the native Codex Plan cell and observe the fallback Plan card
after the tool activity row.
5. Run the legacy Claude structured-question resume cell and observe an
early marker echo in waiting prose.

**Paperclip version or commit**

`0f9452101740835ce0b1488a204bf48acd5bafc3`

**Deployment mode**

Local development with the paid GitHub Actions acceptance workflow.

## What Changed

- Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM
entrypoints before hashing and verified descriptor launch.
- Anchor ACPX dynamic provider package resolution at a
controller-derived provider-pack root and keep that root out of the
provider child environment.
- Persist executor-returned startup errors as redacted durable failed
command results while retaining indeterminate recovery for true process
death.
- Coalesce sparse native tool items by stable ID so a late
`write_document` name, input, and result reach the transcript boundary
once.
- Forbid the structured-question fixture from spelling or announcing its
reserved terminal marker before the user answers.

## Verification

- Rust and TypeScript regression tests cover durable failed replay, true
crash ambiguity, bundle closure, package-root derivation, environment
filtering, exact Codex tool lifecycle coalescing, and prompt
determinism.
- Local execution is intentionally limited to formatters and static diff
checks. GitHub Actions will run tests, type checks, builds, and security
checks.
- After ordinary CI is green, scoped paid cells will validate one ACPX
launch, one OpenCode launch, native Codex Plan projection, and legacy
Claude structured resume before a complete matrix rerun.
- Prior failing matrix:
https://github.com/paperclipai/paperclip/actions/runs/33682434315

## Risks

- Bundling changes the bytes covered by provider launch hashes.
Provider-pack generation already hashes the final built files.
- ACPX still loads qualified provider packages dynamically. The
controller supplies a normalized package root, while existing version,
digest, path, and descriptor checks remain active.
- Durable `failed` is terminal. Replays return the same redacted result
and do not execute the provider effect twice.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5 with agentic reasoning, repository
inspection, code editing, Git, parallel subagents, and GitHub Actions
coordination. The exact deployed snapshot and context-window size are
not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked related public work or described the bug in
this PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] No documentation change is required for this runtime repair
- [x] I have considered and documented the risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 07:58:44 -05:00
Dotta 9ca24bba3c feat(runner): pin the Codex ACPX runtime (#12400)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The package-local host boundary is ready for a concrete ACP
implementation, but the first production profile is Codex only.
> - ACPX must not inherit the server process environment or choose an
executable by pathname after admission.
> - Codex must not re-enable ambient apps, memory, skills, MCP
configuration, or instructions inside its isolated home.
> - This pull request pins only the two required production packages and
applies narrowly tested host patches.
> - The benefit is a minimal dependency boundary that follows the
repository's CI-owned lockfile process.

## Linked Issues or Issue Description

**Agent or provider**

Codex through `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`.

**Why this adapter is useful**

The injected runtime host needs a concrete ACP session manager and the
exact reviewed Codex ACP server. Upstream ACPX does not yet expose a
host-owned spawn callback, and upstream Codex ACP does not yet apply
Paperclip's isolated instruction, MCP, app, memory, and skill boundary.
Both behaviors are required before the dependency can execute inside the
runner.

**How the agent is invoked**

The next pull request will adapt these pinned packages to the private
runtime host. ACPX receives a host-owned callback that consumes the
already verified executable lease. Codex receives only the isolated
environment, explicit base instructions, explicit MCP servers, and the
skills rooted in its private `CODEX_HOME`. This pull request alone does
not spawn either package or register an adapter.

**Additional context**

This pull request is stacked on #12399. It adds no Pi, Claude, AWS, SDK,
lab, browser, or UI dependency. It intentionally does not commit
`pnpm-lock.yaml`: the repository policy job regenerates a manifest-only
PR lockfile artifact for downstream frozen installs, and the lockfile
bot updates master separately.

## What Changed

- Pin `acpx` to `0.13.1` and the Codex ACP server to `1.6.2` in the
runner package.
- Register both patches in the pnpm 9 root configuration and newer-pnpm
workspace configuration.
- Preserve the existing embedded-Postgres and ACPX 0.12 patch entries
used by other packages.
- Patch ACPX to evaluate an allowlisted environment at child-spawn time
and keep spawn cwd out of provider-visible session identity.
- Patch ACPX to accept a host-owned spawn callback with the resolved
arguments and options, allowing the verified command lease to own
execution.
- Patch Codex ACP to retain runner-owned MCP server identity in
permission requests.
- Patch Codex ACP to pass explicit Paperclip base instructions on both
start and resume.
- In isolated mode, disable ambient apps, memory, and existing MCP
configuration; load skills only from `CODEX_HOME`; and configure only
requested servers.
- Add a package contract test that enforces exact versions, Codex-only
dependency scope, both pnpm patch registries, and every required patch
hook.

## Verification

- Both patch files dry-apply successfully to fresh published tarballs
for `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`.
- A local no-lockfile install applied both patches; their runtime
markers and exact installed versions were inspected.
- Runner TypeScript typecheck — passed against the patched packages.
- Runner package tests — passed: 16 Node protocol/package tests and 426
Vitest tests.
- `pnpm -r typecheck` — passed for all applicable workspaces.
- `pnpm build` — passed, including runner binary, server, UI, and
workspace packages.
- `git diff --check` — passed.
- The diff contains 6 files and does not change `pnpm-lock.yaml`, a
GitHub workflow, server selection, or UI behavior.

## Risks

The primary risk is drift between published package contents and
checked-in compiled patches. Exact versions are pinned, both patches are
exercised by package-contract gates, and CI performs the authoritative
regenerated-lockfile frozen install. The spawn callback does not grant a
new executable path: the following adapter must consume the opaque
verified command lease. Codex isolation changes activate only when
`PAPERCLIP_ACPX_ISOLATED_CONTEXT=1`, so existing direct Codex adapters
are unaffected.

## Model Used

OpenAI Codex with GPT-5 and repository tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked an existing public item or described the
issue in this PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal task
identifier
- [x] I have run the affected tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have documented the dependency, patch, isolation, and lockfile
boundaries
- [ ] All applicable GitHub Actions are green
- [ ] Greptile is 5/5 with every actionable comment resolved
- [x] I will address all review findings before requesting merge
2026-08-30 18:11:25 -05:00