mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane must keep each active issue on a clear execution or recovery path. > - A missing issue disposition can require more than one bounded repair attempt. > - A server restart could lose that repair path or move source ownership to the recovery owner. > - A parked or expired retry could also make the user interface show a false healthy state. > - Concurrent recovery loops must not schedule the same repair attempt twice. > - This pull request keeps retry state durable, makes scheduling atomic, and keeps source ownership stable. > - The benefit is that recovery continues after a restart and operators see the correct state. ## Linked Issues or Issue Description **What happened?** A run that ended without a valid issue disposition could lose its repair path after a server restart. Manager recovery could also change the source owner. In addition, a parked or expired retry could make the issue look healthy when no active work existed. Concurrent reconciliation could also schedule the same repair attempt twice. **Expected behavior** Paperclip must keep bounded source and manager repair attempts across restarts. Recovery ownership must stay separate from source issue ownership. The server and user interface must report only a live retry as active work. Each repair attempt must be scheduled at most once per company. **Steps to reproduce** 1. Start an agent run on an issue. 2. End the run without a valid issue disposition. 3. Let the first repair attempt schedule a retry. 4. Restart the server, let the retry time pass without a live run, or start two reconciliation loops together. 5. Observe that the repair path can stop, the issue can show a false healthy state, or duplicate retries can be created. **Paperclip version or commit** The problem existed on `master` before candidate head `d8e620fe86bade7df18decac332007f5821ae04f`. **Deployment mode** The problem affects self-hosted servers and local builds that use automatic recovery. ## What Changed - Persist bounded source-owner and manager repair lineages with stable fingerprints and retry limits. - Resume incomplete disposition repairs after a server restart. - Keep recovery ownership separate from source issue ownership and enforce source mutation authority. - Project live retry evidence into issue and blocker summaries. - Show recovery owner, return owner, attempt count, and retry state in the board user interface. - Treat expired or parked retries as attention states unless a queued or running attempt exists. - Atomically deduplicate disposition-repair wake requests with a company-scoped partial unique index. - Reuse the winning run when concurrent reconciliation loses the uniqueness race, without duplicate scheduling activity. - Honor disabled on-demand wake policy before recovery scheduling and again before delayed retry promotion. - Keep the new index migration safe for lagging seeded databases that already contain the index. - Add server and user interface tests for recovery, restart, ownership, retry, concurrency, and blocker states. - Update the implementation and execution semantics documents. ## Verification - Focused server recovery and ownership suites: 282 tests passed on the repaired base candidate. - Focused user interface recovery suites: 128 tests passed on the repaired base candidate. - Atomic-deduplication schema and recovery suites: 111 tests passed on the first Greptile repair. - Recovery and scheduled-retry wake-policy suites: 126 tests passed at `d8e620fe86bade7df18decac332007f5821ae04f`. - The exact lagging-source migration-order test passed after the index migration became idempotent: 1 test passed and 62 unrelated tests were skipped. - `@paperclipai/db` and `@paperclipai/server` typechecks passed at the current head. - Migration generation and migration safety checks passed for migration `0226_tan_colossus.sql`. - `pnpm check:token-gates` passed on the repaired base candidate. - `pnpm -r typecheck` passed on the repaired base candidate. - `pnpm build` passed on the repaired base candidate. - `pnpm test:run` passed 4,540 tests on the repaired base candidate. Four fixed-port cases met listeners that already existed on the host. - The two unchanged fixed-port files passed in an isolated network namespace: 129 tests passed and 27 tests were skipped. - Independent Security and QA reviews approved `63c0423aab54c66f2293a20b0fb3f3b013ee3ba8`; exact-head re-review is required after automated checks settle on `d8e620fe86bade7df18decac332007f5821ae04f`. ## Risks - Recovery orchestration affects issue liveness and ownership. The new paths use bounded attempts, stable fingerprints, row locks, authority checks, and database uniqueness. - A conservative attention state can show more warnings when a scheduled retry has no queued or running attempt. It does not hide stopped work. - Migration `0226_tan_colossus.sql` creates a partial unique index on a known-large table. Migrations run transactionally, so `CONCURRENTLY` is unavailable. The matching disposition-repair key namespace is introduced by this release, so deployed databases have no matching rows before the index is added. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex from the GPT-5 model family used agentic reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window. - Anthropic Claude Opus 5 used a 1M context window, tool use, and code execution for part of the user interface repair, as recorded in the commit history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
@paperclipai/ui
Published static assets for the Paperclip board UI.
What gets published
The npm package contains the production build under dist/. It does not ship the UI source tree or workspace-only dependencies.
Storybook
Storybook config, stories, and fixtures live under ui/storybook/.
pnpm --filter @paperclipai/ui storybook
pnpm --filter @paperclipai/ui build-storybook
Typical use
Install the package, then serve or copy the built files from node_modules/@paperclipai/ui/dist.