Files
PaperClipAI/server
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
..