Files
PaperClipAI/packages/db
Devin Foley e1d2e279a3 fix(db): replay a query whose socket write failed on a recycled pooled connection (#13417)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A hosted instance keeps its data in Postgres, and it reaches that
database through a connection pooler
> - A pooler recycles the server side of a connection that sat idle. The
client still holds the socket and believes it is open
> - The next query fails while the driver writes it to that dead socket,
with `write CONNECTION_CLOSED host:5432`
> - The request that happened to draw the recycled connection fails,
although nothing is wrong with the query or the database
> - Only one code path guards against this today, so the same error
keeps reaching users and error reporting from every other path
> - This pull request replays a query when the write itself failed,
because those bytes never reached the server
> - The benefit is that a recycled connection costs one retry instead of
one failed request

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
Requests fail with `write CONNECTION_CLOSED <host>:5432` (driver code
`CONNECTION_CLOSED`). It is most visible on the first queries after an
idle period, when the pool holds connections the pooler already
recycled.

**Expected behavior**
A connection that the pooler recycled while it was idle is not a
user-visible failure. The client notices the dead socket and uses a live
connection.

**Steps to reproduce**
1. Run the server against a pooled Postgres endpoint (for example a Neon
`-pooler` host).
2. Leave the instance idle until the pooler recycles the server side of
the pooled connections.
3. Issue any request that queries the database.

**Paperclip version or commit**
Present on master.

## What Changed

- `packages/db/src/transient-write-retry.ts` (new) wraps the root
`postgres.js` client. When a query fails while the driver writes it to
the socket, the wrapper runs it again on a fresh connection. Three
attempts, 50 ms then 100 ms backoff.
- `packages/db/src/client.ts` gives Drizzle the wrapped client. The
teardown registry keeps the real client, because shutdown must end the
actual pool.
- The wrapper is deliberately narrow:
- It matches only the write phase (`code === "CONNECTION_CLOSED"` and a
message that starts with `write CONNECTION_CLOSED`). The write failed,
so the server never saw the query, and a replay cannot run anything
twice. That makes it safe for reads and writes alike.
- An error after the write propagates untouched, because the server may
have acted on the query.
- `CONNECTION_ENDED` and `CONNECTION_DESTROYED` propagate untouched,
because they mean a deliberate shutdown.
- Queries inside `db.transaction()` run on the scoped client that
`sql.begin()` returns, which the wrapper does not touch. A transaction
that loses its connection must abort, not replay.
- A query runs once however many handlers attach to it, and it still
executes lazily, like `postgres.js` itself.

## Verification

```
npx vitest run packages/db/src/transient-write-retry.test.ts            # 7 passed
npx vitest run packages/db/src/client-teardown-registry.test.ts \
              packages/db/src/client-options.test.ts                    # existing db suites pass
npx vitest run server/src/__tests__/cloud-tenant-transient-db-retry.test.ts  # 6 passed
cd packages/db && npx tsc --noEmit                                      # clean
```

New tests cover the replay, the `.values()` form Drizzle uses, the
attempt budget, an error that must not replay, and one execution per
pending query. A Drizzle round trip runs through a fake wire-protocol
server, which shows the wrapper is transparent to ordinary queries.

## Risks

Low risk.

- The retry only fires for a failure during the socket write, where the
server never received the query. A replay therefore cannot duplicate an
effect.
- A permanently unreachable database costs two extra attempts and 150 ms
before the same error surfaces.
- Transactions keep exactly their current behavior.
- Existing `retryOnTransientDbConnectionError` in the auth middleware
stays. It wraps a broader set of codes for one path, and it is
unaffected.
- No migration. No configuration change. No API change.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior or configuration changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-14 09:33:48 -07:00
..