mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `packages/db` module owns all database migrations via a sequential numbering system already validated at build time > - A recent migration introduced an O(n²) batch backfill over a large, unindexed table — it caused the server's `listen` to block for ~5 minutes on databases with millions of rows > - Nothing in the current CI pipeline catches large-table migration risk patterns (DO-loop mutations, batched LIMIT mutations without support indexes, full-table mutations, non-concurrent index creation) before they land > - This PR wires a new static migration-safety checker (`check-migration-safety.ts`) into the existing `check:migrations` gate in `packages/db/package.json`, so risky patterns fail CI before reaching production > - The checker baselines all historical findings already present in the codebase, so the gate fails only on *new* unbaselined risky patterns > - The benefit is that the specific O(n²) backfill shape (and related patterns) will be caught at author time rather than at incident time ## Linked Issues or Issue Description **Feature — static migration safety lint** **Problem or motivation** Migrations against large tables (millions of rows) have caused production startup blocks. The root pattern is a batched `LIMIT`-based backfill iterating via an unindexed column, making each batch a sequential scan — O(n²) overall. No CI gate exists to flag this class of problem before merge. **Proposed solution** A static SQL-level checker that scans new migration files for known dangerous patterns against known-large tables, producing structured findings that are either baselined (suppressed) or fail the build. Patterns detected: `DO $$ loop` mutations on large tables without a same-migration support index, batched `LIMIT` mutations on large tables missing a same-migration support index, unbounded full-table mutations (no `WHERE` clause), and `CREATE INDEX` without `CONCURRENTLY` on large tables. **Alternatives considered** Runtime instrumentation (only catches issues in production), advisory locking in migrations (doesn't prevent the pattern), per-migration code review (doesn't scale consistently). **Roadmap alignment** Defensive infrastructure / operational reliability — keeps migrations from blocking production startups. Not a user-facing feature. ## What Changed - **`packages/db/package.json`** — extended `check:migrations` script to run `check-migration-safety.ts` after the existing numbering check - **`packages/db/src/check-migration-safety.ts`** — new static checker: SQL pattern matching, rule detection for four dangerous patterns, baseline diffing, and structured exit with findings summary - **`packages/db/src/migration-safety-baseline.ts`** — baseline of all existing historical findings (suppressed from failing the gate); new migrations matching these patterns without a baseline entry will fail - **`packages/db/src/table-size-estimates.ts`** — rough table size estimates from the local dev database; drives `isKnownLargeTable()` used by the safety rules - **`packages/db/src/check-migration-safety.test.ts`** — Vitest coverage for the O(n²) backfill failure mode, suppression via baseline, and each rule type ## Verification ```bash # Run the migration safety checker directly cd packages/db tsx src/check-migration-safety.ts # Run tests cd packages/db npx vitest run src/check-migration-safety.test.ts # Run the full migration check gate (numbering + safety) cd packages/db pnpm run check:migrations ``` - Tests cover the core O(n²) backfill pattern (the motivating incident), baseline suppression, and all four rule types - `check:migrations` now exits non-zero for any new unbaselined large-table migration risk pattern ## Risks - **False positives:** Table-size estimates are from a local dev database snapshot — a table small in dev but large in production would be missed. Best-effort heuristic. - **Baseline drift:** If a baselined finding's SQL changes significantly, the baseline ID (content-hash-based) will no longer match and the finding will re-surface. Intentional but may surprise authors doing incremental fixes. - **SQL parsing limitations:** Regex-based pattern matching rather than a full AST parser — complex SQL may not be detected. Acceptable for an initial gate. - **Low risk to existing behavior:** The gate only fails on *new* findings not present in the baseline. All existing migrations are baselined. ## Model Used - **Provider:** Anthropic - **Model ID:** `claude-sonnet-4-6` - **Context window:** 200K - **Tool use:** yes (file reading, bash, git operations) - **Reasoning mode:** standard (no extended thinking) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>