Files
PaperClipAI/server
Devin FoleyandBender de9414dd8b fix: return empty read instead of past-EOF range when log reader is caught up (#13592)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server streams agent run logs to the UI. It reads the log in
byte ranges from a local file, or from an S3 mirror after a pod restart.
> - The range math in `readS3Range()` clamps the range end up to the
range start. A fully caught-up reader then asks S3 for the range
`bytes=total-total`.
> - S3 rejects a range that starts at the end of the object. It returns
a 416 `InvalidRange` error. The API turns this into a 500 error, and the
log poller repeats it.
> - This pull request removes the clamp. A caught-up or past-EOF reader
now gets an empty read, and the server does not send an invalid range to
S3.
> - The benefit is that log polling after a pod restart does not cause
repeated 500 errors.

## Linked Issues or Issue Description

No public GitHub issue exists for this bug. The description follows the
bug report template.

**What happened?**

The run-log API returns a 500 error when a client polls a run log that
lives on the S3 mirror and the client is fully caught up (`offset ===
total`). The cause is in `server/src/services/run-log-store.ts`. The
function `readS3Range()` computes `end = Math.max(start, Math.min(start
+ limitBytes - 1, total - 1))`. When `offset === total`, the
`Math.max(start, …)` clamp forces `end` up to `start`. The `start > end`
empty-read guard does not operate, and the code sends `Range:
bytes=total-total` to S3. S3 rejects a range that starts at or past the
end of the object with a 416 `InvalidRange` error. The error monitor
records this error many times, only in the staging environment, because
only the S3 fallback path is sensitive to it. The local-file path has
the same math, but Node file streams accept past-EOF reads. The function
`readFileRange()` in
`server/src/services/workspace-operation-log-store.ts` has the same
latent math.

**Expected behavior**

A caught-up reader gets an empty read: `{ content: "", nextOffset:
undefined }`. The server does not send an invalid range request to S3.
The poller sees no contract change.

**Steps to reproduce**

1. Start a run and let it write a run log.
2. Let the log upload to the S3 mirror, and remove the local file (this
occurs when the pod restarts).
3. Poll the run-log read endpoint until the client offset is equal to
the log size.
4. Poll one more time. The server sends `bytes=total-total` to S3, S3
returns 416 `InvalidRange`, and the API returns a 500 error.

**Relevant logs or output**

```
InvalidRange: Invalid range
    at readS3Range (server/src/services/run-log-store.ts)
```

## What Changed

- `server/src/services/run-log-store.ts` — remove the up-clamp in
`readLocalRange()` and `readS3Range()`. A caught-up or past-EOF reader
gets an empty read.
- `server/src/services/workspace-operation-log-store.ts` — apply the
same fix to the shared math in `readFileRange()`.
- `server/src/services/run-log-store.test.ts` — the in-memory S3 mock
now rejects past-EOF ranges with `InvalidRange`, the same as real S3.
Add two regression tests for caught-up readers on the S3 path and on the
local path.

## Verification

- Run `pnpm vitest run server/src/services/run-log-store.test.ts`. All
17 tests pass.
- Revert only the source fix, and the new regression test fails with the
exact caught-up scenario. This shows the test covers the bug.
- Run the suites that use the workspace operation log store
(`workspace-runtime-control-recovery`,
`workspace-operations-reconciliation`). All 15 tests pass.
- Run `tsc --noEmit` on the server package. It reports no errors.

## Risks

- Low risk. The change only affects the empty and caught-up boundary of
range reads. Normal in-range reads give byte-identical results.
- Behavior change: a read with `limitBytes <= 0` now returns an empty
chunk instead of one byte. No caller passes a non-positive limit (the
default is 256000).
- Caught-up local reads keep the `nextOffset: undefined` semantics, so
pollers see no contract change.

## Model Used

- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), with extended
thinking and tool use, run through the Claude Agent SDK.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
2026-09-18 07:21:11 -07:00
..