Files
PaperClipAI/doc/plans/2026-09-09-codex-integration-acceptance.md
DottaandPaperclip 3b550c80fa fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path

> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.

## Linked Issues or Issue Description

**What happened?**

Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.

**Expected behavior**

Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.

**Steps to reproduce**

1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.

**Paperclip version or commit**

Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at 6abeb6733. Related authority work: Refs #13092.
This PR retains its startup cleanup and protocol-integrity checks.

**Deployment mode**

Local source checkout and disposable Daytona sandbox.

## What Changed

- Classify the exact historical resume usage event before the generic
stale-turn warning.
- Persist cumulative usage baselines across recovery of the same run.
- Use excludeTurns on resume and lightweight thread reads.
- Page turn metadata and selected turn items with cursor and identity
validation.
- Reject unsupported or incomplete history instead of guessing that
execution is idle.
- Trust the startup execution root on its host, including Git worktree
trust keys.
- Start Codex in that root and retain the selected sandbox profile on
later turns.
- Keep unrelated isolated configuration and Codex's separate hook trust
policy.
- Add Rust, TypeScript, accounting, native integration, and local
run-log documentation.

## Verification

- Codex and native-transport TypeScript: 333 passed before PR replay.
- Adjacent OpenCode/ACPX driver and accounting tests: 49 passed.
- Rust library, serialized: 226 passed. Native Codex integration: 72
passed, 1 ignored, plus two pagination regressions.
- Repository typecheck and build passed. All repository test groups have
passing coverage after fixture and resource retests; the initial
monolithic command was not clean.
- Fresh real Codex native browser tasks returned correct answers without
the three targeted notices. Answers persisted after refresh and restart.
- Real same-thread TypeScript driver tests passed locally and in
Daytona, including cold resume, configuration, skills, and an approved
harmless hook.
- Local usage summed to 64,607 tokens. Daytona usage summed to 42,737
tokens. Each sum matched its final session total exactly.
- See doc/plans/2026-09-09-codex-integration-acceptance.md for the scope
and limits of the live tests.
- After replay onto current master and review fixes: 334 Codex, backend,
and live-session tests passed, including checkpoint serialization and
real-runner process restart. TypeScript checks passed.
- The native Codex integration run passed 83 tests; the large lineage
test passed separately with the release runner (its debug build exceeded
the test deadline).
- All GitHub checks passed on the final PR head. Greptile is 5/5 with no
unresolved review threads. CI regenerates the lockfile for the added
TOML dependency, per repository policy.
- The first server shard hit a timing-dependent duplicate-key failure in
the unchanged artifact-document concurrency test. Its focused 11-test
suite passed locally. One CI retry on the same head passed all 103 files
and 1,405 tests (2 skipped): [retry
result](https://github.com/paperclipai/paperclip/actions/runs/34398832930/job/102631274667).

## Risks

- Trust applies only to the server-selected startup root and isolated
configuration. Sandbox and tool permissions remain authoritative.
- Codex still requires approval of individual hook hashes. This change
does not bypass that policy.
- Providers without the required history APIs fail explicitly.
- Daytona acceptance used the production TypeScript driver. Remote
Paperclip UI and remote Rust execution were not tested.
- No new public API, database state, recovery policy, or UI control is
included.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`). Used for reasoning, code edits,
tool use,
and test execution. The exact context-window limit is not exposed in
this
session. Real-provider acceptance used Codex CLI 0.153.4 with
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 15:35:18 -05:00

4.5 KiB

Codex integration acceptance — 2026-09-09

Environment and scope

All live tests used the pinned Codex CLI 0.153.4 and gpt-5.6-sol. Each fixture used an isolated Codex home. A repository contained known launch notes, a configuration marker, a skill marker, and a harmless SessionStart hook. The hook only appended a line to a fixture file.

Streaming and feed-display changes remain separate pull requests. These tests used their combined implementation checkout. Each PR is also checked on the current master before merge.

Native browser acceptance

A fresh test-drive instance started without tasks or prior runs. The test agent used paperclip_runner with Codex. Browser actions created a task in the fixture project, requested the notes, and sent two follow-up messages. The test server was restarted between the first answer and the follow-ups.

All three native runs succeeded. They returned the expected notes and markers, then answered a date-change question and recalled the original reference. Answers persisted after refresh. No provider notice appeared. The task composer remained usable.

The task policy opened new provider threads for completed-task follow-ups. This browser test does not prove same-thread resume. The next tests do.

Same-thread local driver acceptance

The production TypeScript driver ran an initial repository read, a follow-up, a provider shutdown, persisted-session recovery, and a second follow-up. All answers used the same provider thread and retained the required context.

  • No provider notices appeared.
  • Configuration and skill markers loaded.
  • The approved hook ran once at startup and once at cold resume.
  • Run usage was 40,445 + 13,624 + 10,538 = 64,607 tokens.
  • The sum equaled the final cumulative session usage. Historical usage was not charged again.
  • Resume usage produced only the bounded local diagnostic.

Daytona driver acceptance

A disposable Daytona sandbox ran the production TypeScript driver bundle. The test installed Codex 0.153.4 because the image had an older version. Startup, warm follow-up, process shutdown, cold resume, and the second follow-up all succeeded on the same provider thread.

  • Configuration and skill markers loaded.
  • The approved hook ran once at startup and once at cold resume.
  • No trust, history-deprecation, or settled-turn usage warning appeared.
  • Run usage was 23,006 + 11,611 + 8,120 = 42,737 tokens.
  • The sum equaled the final cumulative session usage.
  • A separate warning about missing system bubblewrap remained visible. Codex used its bundled copy. The change does not suppress that warning.

The sandbox was deleted after evidence collection. Remote Paperclip UI and remote Rust execution were not tested.

Hook trust and experience limits

Codex reviews hook hashes separately from repository trust. The test queried hooks/list and approved only the harmless fixture hash through config/batchWrite. Product code does not bypass this policy or copy operator configuration.

The task feed worked correctly. Two separate issues remain outside this change: a dashboard preview could retain old running text, and test-drive restart could use a different database when its saved port did not match its actual port. The fixture configuration was corrected before sending further messages. Existing test data was preserved. Transitions were sampled rather than filmed.

Automated verification before PR preparation

  • Codex/native-transport TypeScript: 333 passed.
  • Adjacent OpenCode/ACPX driver, accounting, and recovery tests: 49 passed.
  • Same-run attach and usage baseline regression: 4 passed.
  • Rust library, serialized: 226 passed.
  • Rust Codex integration: 72 passed, 1 ignored; two new pagination tests passed.
  • Production native server integration passed after rebuilding its stale fake provider. The test now asks Cargo to check binary freshness.
  • Repository typecheck and build passed.
  • Repository test groups have passing coverage after environment retests. General-server coverage was 7,217 passed and 30 skipped. Route coverage was 2,175 passed and 4 skipped across 144 files.

The initial monolithic test command did not pass cleanly. Parallel test loads caused timeouts; isolated reruns passed. CLI database tests reached the macOS shared-memory limit; they passed after unused disposable fixture servers were stopped. No production behavior or timeout setting changed for those failures. The default parallel Rust library run also had Claude fixture transport failures; the complete serialized rerun passed. PR checks record validation after replaying these changes onto current master.