Files
PaperClipAI/packages/paperclip-runner/docs
DottaandPaperclip 94e8dec56b fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner sends authorized tool calls to the server and saves their
results.
> - A provider turn can stop while a server write is still running.
> - The old shutdown path invented a failed result that could conflict
with the real result.
> - Truncated execution input and incomplete recovery records made the
failure harder to diagnose.
> - This pull request preserves exact inputs and actual outcomes through
shutdown and restart.
> - Tests force the race and crash boundaries so safe retries do not
repeat writes.

## Linked Issues or Issue Description

**What happened?**

Stopping a turn during a server tool call could record a false failure,
then reject the actual result as a conflict. The diagnostic input
formatter could truncate instruction content before execution. A crash
during saved-result delivery could leave that delivery permanently
indeterminate. Cleanup could hide the first failure, and a retry could
overwrite earlier run logs.

**Expected behavior**

Keep dispatched tools pending until their actual result is known.
Preserve accepted input bytes. Accept identical result delivery without
failing the task. Reject conflicting results with enough evidence to
diagnose them. Recover saved-result delivery without repeating the
business operation.

**Steps to reproduce**

1. Hold an instruction update at the filesystem commit barrier.
2. Stop its provider turn before the server returns the result.
3. Release the write, deliver its result, and replay the same result.
4. Repeat with a restart before and after the delivery receipt is saved.
5. Check that there is one write and one audit row, and that the exact
result survives.

**Paperclip version or commit**

The change was developed from `44736c9c7` and rebased onto `0e5830887`.

**Deployment mode**

Self-hosted server with the native runner. Tests use local runner
processes, scripted providers, and PostgreSQL.

Related work: #12353 added durable semantic tool receipts; #12384 added
durable Codex tool recovery; #12404 bound semantic tools to ACPX
sessions. #14633 covers separate native-provider cancellation and
qualification work. This PR addresses server semantic-tool outcomes and
their durable delivery. No duplicate fix was found. AgentMail discovery
is outside this PR.

## What Changed

- Close turn admission without inventing results for dispatched tools.
Keep pending calls and accept late actual results.
- Accept identical result replay with a diagnostic warning. Include call
identity and both result hashes in real conflict errors.
- Preserve exact execution arguments. Reject prohibited or oversized
input before dispatch. Keep diagnostic previews redacted and bounded.
- Commit instruction-attempt evidence before the filesystem write. Save
completed mutation receipts so concurrent and restarted duplicates
return the first result. Recheck authorization before replay. An attempt
without a completed result stays unknown and cannot execute again.
Definite pre-write failures save and replay their original error without
another write.
- Recover an interrupted saved-result delivery only for backends with
durable result receipts. Never replay an ordinary business operation
with an unknown outcome.
- Preserve the initiating error when cleanup also fails. Record
incomplete settlement evidence. Propagate typed unknown-outcome errors
through the native tool wrapper without creating a false completed tool
result.
- Append run-log attempts and restore the durable log before appending
after local file loss. Reject incomplete restores. Publish a restored
prefix only if the destination is absent so concurrent attempts cannot
overwrite new lines.
- Add deterministic race, crash, replay, authorization, exact-content,
and log-restoration tests. Document their assertions in
`packages/paperclip-runner/docs/durable-recovery.md`.

## Verification

- Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55
applicable checks pass; four conditional/manual checks are skipped. This
includes build, typecheck, Rust, both runner TypeScript shards, server
and workspace tests, all eight browser shards, isolated runner
compilation, and the clean-install release dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36746101110).
- Greptile reviewed this exact head at 5/5 with zero new findings. All
three earlier review threads are resolved.
- Focused local verification includes 11 instruction integration tests,
23 surrounding authority/tool tests, 26 run-log tests, and 169
controller/driver tests. The post-rebase controller/transport/runtime
selection passed 415 tests. The full Rust release suite passed 617 tests
with two ignored. The real-process SIGKILL recovery test passed three
consecutive runs.
- The fault matrix in
`packages/paperclip-runner/docs/durable-recovery.md` uses explicit
barriers, real PostgreSQL rollback, durable journal reloads, and killed
runner processes. It covers late results, identical and conflicting
replay, exact long content, concurrent log restoration, lost commit
acknowledgements, and definite failure replay after the original CAS
base becomes valid again. No paid model calls are needed.
- Full local recursive typecheck and build passed during implementation.
Server typecheck and the runner TypeScript build passed after the review
fixes. The broad local repository test run was stopped after repeated
database startup timeouts. Four timing/launch failures in an earlier
broad runner run passed focused reruns without changed assertions or
timeouts. These are local verification limitations; the complete
current-head CI suite is green. An earlier CI workspace job received an
infrastructure shutdown signal; its current-head replacement passed.

## Risks

- A stopped turn can remain blocked when a dispatched operation has no
proven result. The system does not guess its outcome or rerun its
effect.
- Conflicting results still fail settlement. Existing failed or
conflicting journals are not repaired automatically.
- Accepted semantic input is limited to 480 KiB of encoded JSON to fit
the encrypted transport. Larger input fails before execution.
- Instruction filesystem writes and database receipts are not one atomic
storage operation. A separately committed attempt and audit record
survive rollback. An attempt without a completed success or definite
pre-write failure receipt remains blocked as an unknown outcome. It is
not replayed or reported as success.
- Run-log restoration now reads the durable object before appending when
the local log is missing. Failed or incomplete reads reject the append.
- No schema migration, dependency change, workflow change, or AgentMail
change is included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; the broad
local run limitation is recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:54:56 -05:00
..
…