mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner sends events that keep task state and steering
controls current.
> - Debug trace correlation reread and parsed the full trace for each
pending event.
> - Repeated scans blocked event delivery and left task state minutes
behind the provider.
> - This pull request indexes appended trace records once and reuses
those correlations.
> - Operators get current run state while native trace evidence remains
available.
## Linked Issues or Issue Description
**What happened?**
Native runs appeared live after their provider turn ended. Steering was
unavailable while the board still showed an active run. A CPU profile
attributed 93% of sampled time to trace lookup and file reads.
**Expected behavior**
Debug trace correlation must not delay run events or task state by
minutes.
**Steps to reproduce**
1. Enable native provider trace capture for a Codex run.
2. Produce a long event stream with pending correlations.
3. Compare provider event times with persisted run event times.
**Paperclip version**
Reproduced on source build 6dd48cad43.
**Deployment mode**
Self-hosted server with the native Codex runner.
Related event-delivery work: Refs #12208. This fix addresses synchronous
trace lookup rather than disconnect draining. Searches found no
duplicate trace-index fix.
## What Changed
- Add a transport-owned incremental index for native trace frame
correlations.
- Read at most 1 MiB per lookup and retry incomplete prefixes and
partial records.
- Preserve latest-frame and terminal-status ordering. Invalidate
replaced or truncated traces.
- Clear each index on transport close. Keep active indexes independent.
- Skip correlation records over 64 KiB without buffering or parsing
their full contents. Preserve the original trace file.
- Add regression tests for repeated misses, appended records, partial
UTF-8, large traces, replacement, and cleanup.
- Document the trace lookup performance constraint.
## Verification
- Latest revision: 11 index regressions and the existing real-runner
trace correlation test pass (12 tests).
- Includes 24 interleaved active traces, oversized records across
appends, and incomplete later interpretations.
- `pnpm build` and `pnpm -r typecheck` pass on final head `ddc74e42c`.
- The real-process hard-restart regression passes after rebuilding its
fake-provider binary from the rebased sources.
- The full local test command reports a pre-existing assertion mismatch
in `native-session-resume.test.ts`: the damaged-epoch recovery case
expects the old attach error instead of the new startup-history error.
The same exact failure was reproduced in an isolated checkout of
unchanged base `a20ecce40`. The PR does not change that test or startup
behavior.
- A recorded-trace benchmark reduced 120 lookups from 2.20 seconds to 22
milliseconds.
- The deployed hotfix removed the scan hotspot. The final main-thread
profile was 94% idle. A guarded restart reported no lost runs.
- All GitHub checks pass on `ddc74e42c`, including build, typecheck,
general tests, serialized server tests, end-to-end tests, canary dry
run, and security checks. Greptile is 5/5 on that exact head with no
open threads.
- Two timeout cases from the long local run pass in isolation: exhausted
quota-monitor evidence and native-question expiry. The duplicate full
local run was stopped after the complete CI suite passed.
## Risks
Each active transport retains event-to-frame mappings in memory and
clears them on close. Records over 64 KiB do not enter the correlation
index; the original trace file retains them. Large traces can need
multiple pending retries before a correlation is available. These
records are diagnostic; this change does not alter execution authority,
event acknowledgements, or provider commands.
## Model Used
OpenAI Codex, GPT-6. The runtime does not expose a more specific model
ID or context-window size. Used reasoning, code editing, shell
execution, CPU profiling, and tests.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests; the
full-suite baseline failure is documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>