## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip stores agent work in a PostgreSQL database that evolves via numbered sequential migrations > - Migrations that run large unbounded table scans can block the server's `listen()` call during upgrades, causing multi-minute startup stalls on large databases > - Migration 0126 ran a full sequential scan backfill over `issue_comments` for derived attribution columns — O(n²) due to unindexed `LIMIT`/`OFFSET` batching, pegging CPU for ~5 minutes on a 3–4 M row table > - The safe fix (landed in #9108) replaced 0126 with a new forward-only migration using a partial index + keyset pagination backfill > - But the root cause is the absence of author-time guidance: contributors have no documented rules for writing bounded, indexed migration backfills before they land > - This PR adds a migration authoring checklist to `doc/DATABASE.md` so future contributors have those rules at hand before opening a PR > - The benefit is a durable, discoverable guide that prevents the same class of startup-blocking slowness before it reaches production ## Linked Issues or Issue Description This PR is a documentation follow-on to #9108, which landed the fast 0132 migration fix. It adds author-time guidance that captures the root-cause lesson from that incident. No separate public issue exists for the doc addition; the motivation is described above. Refs #9108 ## What Changed - `doc/DATABASE.md`: Added a **Migration authoring checklist** section with rules for indexed, bounded backfill batches — keyset pagination over `LIMIT`/`OFFSET`, mandatory partial index, idempotent guards, and split-phase schema-vs-data changes. The `check:migrations` CI gate is referenced as the enforcement backstop. ## Verification - `git diff --check -- doc/DATABASE.md` passes (no whitespace errors). - No executable code changed; the checklist is an additive documentation section. ## Risks Low. The change is additive text in `doc/DATABASE.md`. No schema, migration, or code changes. No behavioral diff. ## Model Used Claude claude-sonnet-4-6 (Anthropic, 200 K context, tool use) — used to author the migration authoring checklist and coordinate the PR workflow. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with \`Fixes: #\` / \`Closes #\` / \`Refs #\` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub \`#NNN\` / \`github.com/paperclipai/paperclip\` URLs) - [x] My branch name describes the change (e.g. \`docs/...\`, \`fix/...\`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>
10 KiB
Database
Paperclip uses PostgreSQL via Drizzle ORM. There are three ways to run the database, from simplest to most production-ready.
1. Embedded PostgreSQL — zero config
If you don't set DATABASE_URL, the server automatically starts an embedded PostgreSQL instance and manages a local data directory.
pnpm dev
That's it. On first start the server:
- Creates a
~/.paperclip/instances/default/db/directory for storage - Ensures the
paperclipdatabase exists - Runs migrations automatically for empty databases
- Starts serving requests
Data persists across restarts in ~/.paperclip/instances/default/db/. To reset local dev data, delete that directory.
If you need to apply pending migrations manually, run:
pnpm db:migrate
When DATABASE_URL is unset, this command targets the current embedded PostgreSQL instance for your active Paperclip config/instance.
Issue reference mentions follow the normal migration path: the schema migration creates the tracking table, but it does not backfill historical issue titles, descriptions, comments, or documents automatically.
To backfill existing content manually after migrating, run:
pnpm issue-references:backfill
# optional: limit to one company
pnpm issue-references:backfill -- --company <company-id>
Future issue, comment, and document writes sync references automatically without running the backfill command.
This mode is ideal for local development and one-command installs.
Docker note: the Docker quickstart image also uses embedded PostgreSQL by default. Persist /paperclip to keep DB state across container restarts (see doc/DOCKER.md).
2. Local PostgreSQL (Docker)
For a full PostgreSQL server locally, use the included Docker Compose setup:
docker compose up -d
This starts PostgreSQL 17 on localhost:5432. Then set the connection string:
cp .env.example .env
# .env already contains:
# DATABASE_URL=postgres://paperclip:paperclip@localhost:5432/paperclip
Run migrations:
DATABASE_URL=postgres://paperclip:paperclip@localhost:5432/paperclip \
pnpm db:migrate
Start the server:
pnpm dev
3. Hosted PostgreSQL (Supabase)
For production, use a hosted PostgreSQL provider. Supabase is a good option with a free tier.
Setup
- Create a project at database.new
- Go to Project Settings > Database > Connection string
- Copy the URI and replace the password placeholder with your database password
Connection string
Supabase offers two connection modes:
Direct connection (port 5432) — use for migrations and one-off scripts:
postgres://postgres.[PROJECT-REF]:[PASSWORD]@aws-0-[REGION].pooler.supabase.com:5432/postgres
Connection pooling via Supavisor (port 6543) — use for the application:
postgres://postgres.[PROJECT-REF]:[PASSWORD]@aws-0-[REGION].pooler.supabase.com:6543/postgres
Configure
For the application runtime, use a direct PostgreSQL connection unless the database client has explicit prepared-statement configuration for your pooling mode:
DATABASE_URL=postgres://postgres.[PROJECT-REF]:[PASSWORD]@aws-0-[REGION].pooler.supabase.com:5432/postgres
If you later run the app with a pooled runtime URL, set DATABASE_MIGRATION_URL to the direct connection URL. Paperclip uses it for startup schema checks/migrations and plugin namespace migrations, while the app continues to use DATABASE_URL for runtime queries:
DATABASE_URL=postgres://postgres.[PROJECT-REF]:[PASSWORD]@aws-0-[REGION].pooler.supabase.com:6543/postgres
DATABASE_MIGRATION_URL=postgres://postgres.[PROJECT-REF]:[PASSWORD]@aws-0-[REGION].pooler.supabase.com:5432/postgres
If your hosted database requires transaction-pooling-only connections, use a direct or session-pooled connection for Paperclip until runtime pooling support is documented in this guide. Do not edit database client source files as part of deployment setup.
Push the schema
# Use the direct connection (port 5432) for schema changes
DATABASE_URL=postgres://postgres.[PROJECT-REF]:[PASSWORD]@...5432/postgres \
pnpm db:migrate
Free tier limits
- 500 MB database storage
- 200 concurrent connections
- Projects pause after 1 week of inactivity
See Supabase pricing for current details.
Switching between modes
The database mode is controlled by DATABASE_URL:
DATABASE_URL |
Mode |
|---|---|
| Not set | Embedded PostgreSQL (~/.paperclip/instances/default/db/) |
postgres://...localhost... |
Local Docker PostgreSQL |
postgres://...supabase.com... |
Hosted Supabase |
Your Drizzle schema (packages/db/src/schema/) stays the same regardless of mode.
Migration authoring checklist
The 0126 issue comment attribution backfill showed the failure mode this checklist is meant to prevent: each batch looked for the next rows with an unindexed predicate, so PostgreSQL repeatedly scanned the same table and the migration became O(n²) as the table grew.
When authoring migrations or one-time backfills:
- Create the supporting index for the batch predicate before the backfill loop runs.
- Bound batches by an indexed key, such as an id range or keyset pagination cursor. Do not use
OFFSETpagination or a query shape that re-scans already-visited rows each batch. - Avoid unbounded full-table
UPDATEorDELETEstatements. Add a selective predicate and process rows in bounded batches when table size can be large. - Use
CREATE INDEX CONCURRENTLYfor large existing tables when the migration can run outside a transaction and must avoid long write locks. - Split schema changes, index creation, and data backfill into separate phases so each step has clear locking and rollback behavior.
- Treat the
check:migrationsCI gate as the enforcement backstop for these rules. If it flags a migration, rewrite the migration or add a suppression comment with the indexed predicate, batch bound, and reason the remaining scan is safe.
Resource membership tables
Paperclip stores current-user sidebar membership state in:
project_membershipsagent_memberships
These rows are company-scoped and user-scoped. A missing row means the user is joined, so existing users keep seeing projects and agents in the sidebar until they explicitly leave them. Rows only control sidebar visibility; they do not affect project/agent detail access, all-pages, selectors, assignment flows, or existing company permissions.
Both tables use a unique key on (company_id, user_id, resource_id) and keep state as joined or left. Join/leave mutations are idempotent board-user /me operations and write activity entries when the effective state changes.
Plugin database namespaces
The plugin runtime tracks plugin-owned database namespaces and migrations in plugin_database_namespaces and plugin_migrations. Hosted deployments that separate runtime and migration connections should set DATABASE_MIGRATION_URL; plugin namespace migration work uses the migration connection when present.
Backups
Paperclip supports automatic and manual logical database backups. These dumps include
non-system database schemas such as public, the Drizzle migration journal, and
plugin-owned database schemas. See doc/DEVELOPING.md for the current
paperclipai db:backup / pnpm db:backup commands and backup retention
configuration.
Database backups do not include non-database instance files such as local-disk uploads, workspace files, or the local encrypted secrets master key. Back those paths up separately when you need full instance disaster recovery.
Secret storage
Paperclip stores secret metadata and versions in:
user_secret_definitionsuser_secret_declarationscompany_secretscompany_secret_versionscompany_secret_bindingssecret_access_events
Company secrets use company_secrets.scope = 'company' and are bound directly
through company_secret_bindings. User-specific secrets reuse the same provider
and version storage, but each value is a company_secrets.scope = 'user' row
with owner_user_id and user_secret_definition_id set. Definitions describe
the reusable company-level slot, declarations record where user_secret_ref
bindings are required, and the concrete value is selected later for the
responsible user.
Secret-aware env bindings are supported by agents, projects, and routines. Routine env lives in routines.env, is captured in routine_revisions.snapshot, and routine dispatches store routine_runs.routine_revision_id so runtime secret resolution uses the env snapshot that existed when the run was created. Routine secret refs bind with target_type = 'routine', target_id = routines.id, and config_path values under env.*.
For local/default installs, the active provider is local_encrypted:
- Secret material is encrypted at rest with a local master key.
- Default key file:
~/.paperclip/instances/default/secrets/master.key(auto-created if missing). - CLI config location:
~/.paperclip/instances/default/config.jsonundersecrets.localEncrypted.keyFilePath. - Backup/restore requires both the database metadata and the local master key file; either artifact alone is insufficient.
- The server best-effort enforces
0600key file permissions and provider health reports permission warnings. - User-scoped values use the same local encrypted provider path. Database backups preserve definitions, declarations, owner metadata, version metadata, and access events, but restored user-scoped values are decryptable only when the matching local master key is restored with the database.
Optional overrides:
PAPERCLIP_SECRETS_MASTER_KEY(32-byte key as base64, hex, or raw 32-char string)PAPERCLIP_SECRETS_MASTER_KEY_FILE(custom key file path)
Strict mode to block new inline sensitive env values:
PAPERCLIP_SECRETS_STRICT_MODE=true
You can set strict mode and provider defaults via:
pnpm paperclipai configure --section secrets
Inline secret migration command:
pnpm paperclipai secrets migrate-inline-env --company-id <company-id> --apply
# direct database maintenance fallback
pnpm secrets:migrate-inline-env --apply
Hosted AWS provider notes live in SECRETS-AWS-PROVIDER.md.