Files
PaperClipAI/server/src/services/batch-insert.ts
T
Devin Foley 276ae3a75d Harden company import: durable UI, async jobs, integrity guard, batched inserts (#10523)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Company Import/Export (#10507) moves whole companies between
instances as portability bundles
> - Real-world use on a large company (1,418 issues, ~10.6k comments)
surfaced a cluster of related failures: the import took hours and the
browser connection died while the server kept running, a retry silently
produced a second partial import, the progress/error UI gave no durable
signal, and a cloud-tenant user couldn't even open the companies
afterward
> - Root cause of the slowness: importBundle inserted every issue,
comment, and document as a separate round-trip to a network Postgres —
an N+1-over-network pattern
> - This pull request hardens the whole import path: durable
progress/error UI, an async server-side job so imports survive dropped
connections (with a duplicate-submit guard), a fail-closed guard against
incomplete payloads, and batched inserts that cut a large import from
hours to minutes
> - The benefit is that migrating a real, large company actually
completes, is legible while it runs, and can't half-import twice

## Linked Issues or Issue Description

- Refs #10507 (the Import/Export feature this hardens). Supersedes
#10513 (the progress/error-UI piece, folded in here). No open issue;
problem described above (large-company import: slow, connection-fragile,
silently duplicable, opaque UI).

## What Changed

- **Batched inserts (perf):** importBundle pre-generates entity ids in
JS and inserts in chunked multi-row statements, so children no longer
wait on parents' generated ids. A 1,418-issue import drops from ~15,600
insert statements to **82** (190×); benchmark below. Import semantics —
collision handling, pause-on-import,
label/blocker/monitor/attachment/embedded-asset handling, blob sha
verification — are unchanged (full portability suite green).
- **Async import jobs for board sessions:** the existing cloud-tenant
async job path opens to board sessions with per-actor job keys; the
import page submits, polls, and resumes watching after a reload or
dropped connection instead of holding one fragile request. A
non-terminal job blocks a duplicate submit (409 returns the running
job), preventing the double-import.
- **Fail-closed completeness guard:** an optional `expectedFileCount` on
inline imports; the server rejects (422 `import_payload_incomplete`) a
body carrying fewer files than declared, so a re-framed/short payload
fails loudly instead of half-importing.
- **Durable progress/error UI (was #10513):** persistent progress panels
with size-aware copy, persistent error panels with retry guidance, and
inline explanation when the preview button is disabled;
request-lifecycle guards so stale previews/imports can't publish or
detach.

## Verification

- `pnpm -r` typechecks (shared, server, ui) clean.
- `company-portability.test.ts` (76) +
`company-portability-routes.test.ts` (30) green — the import correctness
net — plus new `CompanyImport.test.tsx` async/resume/409 coverage and a
new batching regression test (a 50-issue import issues <50 issue-insert
statements; rows land unchanged).
- **Batching benchmark (embedded Postgres):** at 1,418 issues × 7
comments × 1 doc — 82 insert statements vs ~15,598 one-per-row (190×),
~1s wall-clock; a row-verifying run at that scale imports all 1,418
issues / 9,926 comments / 1,418 documents with unique identifiers and no
warnings (no rows dropped by chunking). Over a network DB the round-trip
reduction is the hours→minutes lever.
- What is NOT directly measured here: wall-clock against a real network
Postgres (that happens on a staging deploy); the local timing is
network-free.

## Risks

- Batching is the load-bearing change: it rewrites the import write
path. Mitigated by the unchanged 106-test correctness suite, a new
scale/row-integrity test, and per-writer transactions (a failure rolls
back its table group; not a single outer transaction across writers —
noted, correctness preserved).
- Async jobs are in-memory (lost on server restart → pollers 404 and can
resubmit); matches the pre-existing cloud-tenant job semantics.
- `expectedFileCount` is optional (older callers unaffected); over-count
is allowed, only under-count fails closed.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI,
extended thinking + tool use; implementation across Fable 5 subagents
with live diagnosis against a running instance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-30 16:52:58 -07:00

70 lines
2.6 KiB
TypeScript

// Chunked multi-row insert helper.
//
// PostgreSQL caps a single statement at 65535 bind parameters. A multi-row
// insert binds `columnsPerRow * rowCount` parameters, so large imports must
// split their rows into chunks that stay under that ceiling. This helper keeps
// the arithmetic (and the "insert nothing when there is nothing" guard) in one
// place so every import writer batches identically.
// One below the hard 65535 ceiling so the arithmetic never lands exactly on it.
export const POSTGRES_MAX_BIND_PARAMS = 65534;
// A conservative row cap so a single statement stays small even for narrow
// tables. Large enough that the common import tables issue a handful of
// statements, small enough to avoid pathological statement sizes.
export const DEFAULT_INSERT_CHUNK_ROWS = 500;
// drizzle types `insert` as a per-table generic (`<T extends PgTable>`), which a
// table-agnostic helper cannot satisfy; `any` on the boundary is the standard
// escape hatch. Callers pass concretely-typed `db`/`tx` handles and tables.
type InsertExecutor = {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
insert: (table: any) => { values: (rows: any) => unknown };
};
/**
* Insert `rows` into `table` using chunked multi-row statements.
*
* Rows are normalized to a shared column set (the union of keys across the
* batch, with missing/`undefined` values written as `null`) so a single
* multi-row `values()` call stays well-formed even if a caller omitted an
* optional column on some rows. Callers must therefore set every NOT NULL /
* default-backed column they rely on explicitly; only genuinely nullable
* columns should be left absent.
*/
export async function insertRowsInChunks(
executor: InsertExecutor,
table: unknown,
rows: Array<Record<string, unknown>>,
options?: { maxRows?: number },
): Promise<void> {
if (rows.length === 0) return;
const keys = new Set<string>();
for (const row of rows) {
for (const key of Object.keys(row)) keys.add(key);
}
const columns = [...keys];
const normalized = rows.map((row) => {
const out: Record<string, unknown> = {};
for (const key of columns) {
const value = row[key];
out[key] = value === undefined ? null : value;
}
return out;
});
const columnsPerRow = Math.max(1, columns.length);
const chunkSize = Math.max(
1,
Math.min(
options?.maxRows ?? DEFAULT_INSERT_CHUNK_ROWS,
Math.floor(POSTGRES_MAX_BIND_PARAMS / columnsPerRow),
),
);
for (let start = 0; start < normalized.length; start += chunkSize) {
await executor.insert(table).values(normalized.slice(start, start + chunkSize));
}
}