mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:35:27 +02:00
codex/runner-active-stop-ordering-eval
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
53aad90b9e |
fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path > - Paperclip coordinates agent work through execution adapters. > - Sandbox ACP sessions send ordered input through a remote file queue. > - A temporary provider 502 currently closes the session during an input upload. > - A lost response can occur after the sandbox has consumed the message, so a blind retry can duplicate input. > - This pull request retries gateway failures with the same sequence and drops consumed sequences at the receiver. > - The session can continue through a brief provider failure without repeating a tool call. ## Linked Issues or Issue Description **What happened?** A sandbox ACP run can fail with `ACP agent disconnected during request (connection_close, exit=null, signal=null)` when a provider input upload returns HTTP 502. The bridge destroys its local socket on the first failure and can discard the diagnostic before the proxy reads it. **Expected behavior** A temporary gateway failure should get a bounded retry. A lost response after successful delivery must not duplicate input or reorder later messages. Permanent failures must still close the session. **Steps to reproduce** 1. Run the real sandbox process bridge with an echo child and a local test runner. 2. Inject a provider 502 before preparation, after a chunk upload, or after final publication and consumption. 3. Send the next input message. Before this change, the connection closes instead of delivering it. Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and `502 sandbox`. Related work: #13287 covers shutdown after bridge loss; #13793 covers large launch envelopes. This change covers ordered input delivery within a running legacy ACP session. ## What Changed - Retry input uploads up to three times for recognized Daytona and Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms delays. - Give each upload separate temporary paths and discard already-consumed input sequences, including late publication from an earlier attempt. Clean failed attempts in the background without removing a published message or another attempt’s files. Cleanup cannot delay retries or shutdown. - Keep later input behind the retry. Stop queued input on permanent failure and flush a fixed diagnostic before closing the socket. Neither failure-diagnostic persistence nor shutdown-warning persistence can block teardown. - Add real-process regression tests for lost responses, late publication, retry exhaustion, immediate permanent failure, and diagnostic redaction. - Give accepted run-log file appends up to three seconds to drain before finalization computes the size, hash, and durable copy. Close the run handle to later appends. This waits only for file writes, independently of later DB progress or live-event persistence. If writes remain stalled, return null size/hash metadata and skip the final durable copy so the run can settle. Late writes cannot restart mirroring. - Preserve legacy comment attribution when final log size is unknown by reading existing entries within the unchanged 2 MB scan limit. Storage errors or a three-second read deadline return the evidence already read instead of failing the comment listing; pagination stops at the deadline. The deadline requests cancellation of the underlying local stream or S3 HEAD, GET, and response stream. A separate response timeout returns partial evidence even when filesystem I/O delays cancellation; late reads cannot append evidence or start another page. Each listing retains its existing batches of eight reads, without a shared admission cap that skips readable logs under contention. - Document the retry and log-finalization boundaries in the development guide. ## Verification - Final commit `347daa564b`: [Linux CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2) passed. Greptile Apex review 13 scored this commit 5/5 with no new findings; all 12 review threads are resolved. - The final CI run initially hit a Cursor test timeout and four Discord credential-lock contention failures. All five cases passed in isolation. The two failed shards and their aggregate gate passed on retry without a code change. Those intermittent failures are not claimed fixed by this PR. - `pnpm --filter @paperclipai/adapter-utils typecheck` passed. - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests passed on the final implementation, including 21 new regressions. The original three fault-injection cases failed before the fix. - The regressions cover failed and indefinitely stalled cleanup, Cloudflare gateway responses and retry exhaustion, permanent errors that must not retry, and teardown while failure logging remains indefinitely stalled. Seven Apex regression cases failed before the review fixes. Adapter-utils typecheck and build passed again after the final review change. - `pnpm exec vitest run server/src/services/run-log-store.test.ts server/src/services/run-log-store-cancellation.test.ts`: all 25 tests passed, including four new regressions that failed before the finalization fix. They cover delayed and failed appends, late-write admission, agreement between the local bytes/summary/durable copy, and a stalled append that exhausts the three-second budget. The timeout case verifies unknown metadata, no final upload, and no mirror restart after late completion. New cancellation tests use the real AWS SDK against a local HTTP server. They verify that stalled HEAD, GET, and response-body connections close on abort and that a subsequent read succeeds. Local range and already-aborted read cases also pass. - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t 'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14 targeted tests passed. The null-size reader case, both storage-error cases, the stalled-read case, and the cancellation/concurrent-listing cases failed before their fixes. The new regressions verify that timed-out reads are cancelled, subsequent listings recover, and two concurrent listings both retain their attribution markers. A read that ignores cancellation still returns partial evidence at three seconds and cannot resume pagination when it finishes; this regression failed before the response-timeout fix. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/server build` passed after the response-timeout change. - Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR; the affected packages were rechecked after review fixes. - Full local `pnpm test:run` failed in the general-server group: 511 files passed, 40 failed, and 158 were skipped. Failures include embedded PostgreSQL initialization, read-only cache directory renames, a macOS long-path fixture, and a workspace exposure assertion. The PostgreSQL, cache-permission, and long-path failures also reproduce with both changed implementation files restored to baseline commit `24c58e479a`. The exposure suite passes in isolation both on baseline and the fixed branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later local test groups were not reached. - An earlier CI run hit the Telegram retry-timing failure fixed upstream in #14501. The branch includes that master fix. The selected recovery test passed against a fresh, migrated PostgreSQL 16 database. The embedded PostgreSQL runner is unavailable on this Mac; the isolated database was stopped and removed afterward. - No live agent turn was replayed. The tests use local child processes and injected provider failures. ## Risks Retries are restricted to recognized Daytona SDK and Cloudflare bridge gateway-error messages, which survive plugin RPC serialization. Other errors fail immediately. Temporary upload paths are now unique for all command-managed queue writes. Receiver sequence checks prevent duplicate input; retries do not restart an agent turn. Cleanup and failure logging are nonblocking and best effort; session teardown remains the final cleanup boundary. Log finalization now drains accepted local file writes for at most three seconds and ignores later appends on the closed run handle. A timeout leaves final size/hash unknown and skips the final durable upload; an existing partial mirror may remain available, but it is not claimed as a verified final snapshot. It does not wait for later DB progress or live-event persistence. Optional attribution keeps partial evidence when a read fails or times out. Cancellation closes S3 requests and response streams. Local filesystem I/O may finish after the caller deadline, but a late read cannot change the returned evidence or continue pagination. Later listings can retry after storage recovers. There is no schema, authentication, or permission change. Revert this commit to restore the previous behavior. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; targeted tests pass and full-suite limitations are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
b4e7ba5143 |
feat(run-logs): durable run-log store via object-storage mirror (#8984)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every agent run streams its stdout/stderr/system output into the run-log store (`server/src/services/run-log-store.ts`), and the run-log API serves those logs back for review and debugging > - The only store implementation is `local_file`: logs live on the server pod's filesystem under `PAPERCLIP_HOME` > - In hardened / ephemeral deployments, `PAPERCLIP_HOME` is an `emptyDir` with no persistent volume, so every pod restart wipes the log files while the DB row still references them — the run-log API then returns "Run log not found" for every completed run after any redeploy > - Run logs are the primary audit/debugging trail for agent work; losing them on routine redeploys undermines trust in the platform > - This pull request adds transparent durability: when `RUN_LOG_S3_BUCKET` is set, the store mirrors each completed log to object storage on `finalize` (same `logRef` key) and falls back to it on `read` when the local file is gone; live append/tail stays on the fast pod-local file > - The benefit is that completed run logs survive pod restarts and redeploys with zero changes for existing deployments (unset bucket = today's behaviour) and zero downstream changes (store id stays `local_file`) ## Linked Issues or Issue Description No existing public issue — inline description following the bug report template: **What happened?** After any server pod restart/redeploy, the run-log API returns "Run log not found" for all previously completed runs. The DB still references the log file, but the file is gone because run logs are written only to the pod-local filesystem. **Expected behavior:** Completed run logs remain readable across pod restarts and redeploys. **Steps to reproduce:** 1. Deploy the server with `PAPERCLIP_HOME` on an `emptyDir` (no persistent volume — common in hardened/ephemeral Kubernetes deployments). 2. Complete an agent run and confirm its log is readable via the run-log API. 3. Restart or redeploy the server pod. 4. Request the same run's log — the API throws "Run log not found". **Paperclip version or commit:** reproducible on current `master`. **Deployment mode:** Kubernetes (server pod without persistent volume). **Agent adapter(s) involved:** Not adapter-specific (core bug). Supersedes #8795. ## What Changed - `server/src/services/run-log-store.ts`: the local-file store becomes a durable store with an optional object-storage mirror - `finalize` mirrors the completed NDJSON log to S3-compatible object storage (keyed by the same `logRef`), best-effort so a failed upload can never break run finalization; upload failures are logged via `console.warn` so operators can detect a persistently broken mirror before a pod roll makes logs unreadable - `read` serves the pod-local file when present and falls back to a ranged object-storage read (with correct `nextOffset`) when the local file is gone - Live `append`/tail stays on the pod-local file — fast path unchanged, no per-chunk PUT - Store id stays `local_file`, so nothing downstream changes (feedback pipeline, read casts, fixtures untouched) - New optional config, all read at store construction: `RUN_LOG_S3_BUCKET`, `RUN_LOG_S3_ENDPOINT`, `RUN_LOG_S3_REGION` (default `us-east-1`), `RUN_LOG_S3_PREFIX` (default `run-logs`), `RUN_LOG_S3_FORCE_PATH_STYLE` (default `true`); credentials via the standard AWS env chain; works with any S3-compatible endpoint - Reuses the existing `createS3StorageProvider`; deliberately independent from `PAPERCLIP_STORAGE_PROVIDER` so enabling durable logs does not redirect workspace/file storage - `server/src/services/run-log-store.test.ts` (new): 7 tests with an in-memory `StorageProvider` mock ## Verification - `npx vitest run src/services/run-log-store.test.ts` in `server/` — 7/7 pass locally: - store id stays `local_file` - live read served from the local file (no S3 round-trip) - `finalize` uploads the completed log to the mirror - read falls back to S3 after a simulated pod roll (local file deleted) - ranged S3 read returns correct slice + `nextOffset` - not-found when neither local nor mirror has the log - local-only safe degrade when no bucket is configured - `npx tsc --noEmit -p server` — clean for the touched files - Manual: set `RUN_LOG_S3_*` against any S3-compatible endpoint (e.g. MinIO), complete a run, delete the local `.ndjson` file, and re-request the log via the run-log API — it is served from the mirror ## Risks - Low risk: with `RUN_LOG_S3_BUCKET` unset (the default), behaviour is byte-for-byte today's local-only store - Mirror upload is best-effort by design — a misconfigured bucket loses durability (not correctness) for affected runs; failures are now surfaced via a `console.warn` per failed upload - No DB migration, no API shape change, no change to the persisted `store`/`logRef` handle format ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Fable 5), via Claude Code with extended thinking and tool use (code execution, file editing). Original implementation TDD-authored with the same tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
04a19cbc6e | Address artifact PR review feedback | ||
|
|
f60c1001ec |
refactor: rename packages to @paperclipai and CLI binary to paperclipai
Rename all workspace packages from @paperclip/* to @paperclipai/* and the CLI binary from `paperclip` to `paperclipai` in preparation for npm publishing. Bump CLI version to 0.1.0 and add package metadata (description, keywords, license, repository, files). Update all imports, documentation, user-facing messages, and tests accordingly. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
6d0f58d559 |
fix: storage S3 stream conversion, API client FormData support, and attachment API
Fix S3 provider to use async generator for web stream conversion instead of Readable.fromWeb, add postForm helper and attachment API methods to the UI client, and add local disk storage provider tests. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
fdd2ea6157 |
feat: add storage system with local disk and S3 providers
Introduces a provider-agnostic storage subsystem for file attachments. Includes local disk and S3 backends, asset/attachment DB schemas, issue attachment CRUD routes with multer upload, CLI configure/doctor/env integration, and enriched issue ancestors with project/goal resolution. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |