mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents and their work > - Heartbeat execution relies on adapters distinguishing agent failures from failures in the harness running beneath the agent > - Codex MCP transport crashes can kill the CLI after the JSONL protocol has started but before it emits a protocol-terminal event > - Those interrupted streams were left unclassified, so the control plane terminalized the heartbeat as `heartbeat_failed` / `agent_failure` with no continuation > - Agent-level failure is already expressible through the JSONL protocol via an `error` event, `turn.failed`, or `turn.completed`, so an interrupted nonzero exit can be classified structurally without inspecting unstable error strings > - This pull request reports that shape as `codex_harness_crash` in the `transient_upstream` family and routes it through Paperclip's existing bounded retry and recovery-continuation paths > - The benefit is that transient Codex harness failures recover safely without misclassifying quoted agent output or depending on transport-specific wording ## Linked Issues or Issue Description - **What happened:** Codex MCP transport failures, including rmcp worker death, could terminate the CLI mid-turn after protocol output began but before any terminal JSONL event. The run then became an unclassified terminal heartbeat failure with `continuationCount: 0`; this occurred in 3 of 44 L3 Codex-lane trials during the associated benchmark investigation. - **Expected behavior:** a nonzero Codex exit after the protocol starts but before an `error`, `turn.failed`, or `turn.completed` event should be treated as a harness/infrastructure crash and enter the existing bounded retry policy. - **Why structural classification:** transport error strings vary, and stdout may quote agent output that merely discusses network failures. The protocol boundary identifies whether the agent itself produced a terminal result without regex matching. - **Recovery behavior:** `codex_harness_crash` maps to `errorFamily: transient_upstream`, using the existing `same_session` → `safer_invocation` → `fresh_session` ladder plus the recovery-continuation transient-infrastructure path. - Supersedes the regex-based approach in #10150, which is closed. ## What Changed - Added protocol-state tracking that identifies a nonzero exit after protocol start and before any protocol-terminal event as `codex_harness_crash`. - Propagated the structural classification as `transient_upstream` through the Codex adapter. - Added parse unit coverage, including a faithful crash-shaped stream, without matching stderr transport strings. - Added adapter execution coverage using a fake Codex process that emits a protocol prefix and then dies with the observed rmcp stderr line. - Added heartbeat bounded-retry coverage, including the `errorCode`-only fallback, and recovery-continuation classification coverage. ## Verification - `parse.test.ts` — 16 passed. - `codex-local-execute.test.ts` — 16 passed. - `heartbeat-retry-scheduling.test.ts` — 30 passed. - `service.pause-durability.test.ts` — 6 passed. - Server and Codex adapter TypeScript checks passed. - The branch commit is unchanged from the tested and pushed `88f5464d40` handoff. ## Risks - Low risk: the classification requires a nonzero exit after protocol start and before any protocol-terminal event, so normal agent-declared failures and completed turns keep their existing behavior. - The change intentionally broadens recovery for structurally interrupted Codex runs; bounded retry limits still prevent indefinite continuation loops. - No schema, migration, public API, UI, lockfile, or workflow changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent. The exact runtime model ID and context-window size were not exposed by the execution environment; capabilities used for the implementation included repository analysis, reasoning, code editing, and terminal-based test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — the pre-existing, already-pushed branch name was explicitly prescribed for this replacement PR - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>