mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed deployments can auto-provision bundled sandbox-provider plugins so cloud or remote execution environments appear in the board UI > - In a multi-service deployment, several server processes can share one database and boot concurrently > - A sibling process can create a bundled plugin row while the web process sees it before it reaches `ready` > - The web process correctly avoids clobbering the existing row, but its startup `loadAll()` can miss the plugin and never start that worker locally > - The environments capabilities route then filters out the sandbox provider because the plugin is ready in the database but not running in the web process > - This pull request adds a narrow managed-bundle recovery path that lazily starts the missing worker when the capabilities route sees a ready managed bundled plugin > - The benefit is that the sandbox provider becomes visible after the install finishes, without requiring a web-process restart ## Linked Issues or Issue Description - No public GitHub issue found for this exact deployment race. - Related broad plugin runtime context: Refs #432. Bug description: - What happened: in a managed multi-service deployment with shared database state and bundled plugin auto-install enabled, the API-serving process can skip a plugin row while it is still `installed`, run startup plugin loading before that row becomes `ready`, and then permanently omit the sandbox provider from environment capabilities. - Expected behavior: once the managed bundled plugin row reaches `ready`, the API-serving process should be able to start the plugin worker and include its sandbox provider without a restart. - Steps to reproduce: boot a web process and a sibling worker process concurrently; have the sibling create the bundled plugin row and transition it to `ready` after the web process has already skipped auto-install and run `loadAll()`. - Deployment mode: managed multi-service deployment with shared database state and `plugins.autoInstall` configured. ## What Changed - Added a managed bundled plugin worker recovery helper that single-flights lazy `loadSingle()` starts and only allows configured managed bundled plugin keys. - Passed the managed recovery hook into the environments capabilities route. - Updated `listReadyPluginEnvironmentDrivers()` to attempt bounded recovery for ready managed bundled plugins whose worker is missing in the current process, and only for plugins that actually declare a `sandbox_provider` environment driver. - Made request-time recovery use `loadSingle(id, { markErrorOnFailure: false })` so a local activation failure in one process never transitions the shared plugin row to `error` (a sibling process may be running the plugin successfully). - When error writes are suppressed and activation fails after the worker was spawned, the loader now tears down the partially-registered local runtime (scheduler registration, event subscriptions, agent tools, worker process) instead of leaving a half-activated worker lingering; the teardown steps are factored out of `unloadSingle()` into a shared helper. - A failed recovery attempt now discards the crashed/stopped handle it left registered in the worker manager (a worker that dies during initialize is killed without a scheduled restart), so later capability requests can retry recovery instead of being blocked by the handle-presence gate until a process restart. Handles in starting/running/backoff states are left to the worker manager's own lifecycle; recovery only ever starts when no handle existed, so no pre-existing worker can be affected. - Added a regression test suite covering the installed-to-ready race, allowlist behavior, the driver-kind gate, existing worker handles, concurrent single-flight recovery, bounded slow recovery attempts, suppressed shared error-state writes, partial-runtime teardown on late activation failure, and retry after a dead handle is discarded. ## Verification - `pnpm vitest run src/__tests__/plugin-environment-driver-ready-recovery.test.ts` (in `server/`) passed: 10 tests. - `pnpm --filter @paperclipai/server typecheck` passed. ## Risks - Low risk for self-hosted single-process deployments because lazy recovery is only wired when managed plugin auto-install config is present; with no managed config the capabilities route takes the exact pre-change code path. - The capabilities route can wait briefly while attempting recovery; the attempt is bounded and defaults to 2 seconds. - Failed recovery keeps the prior behavior of omitting the provider until a later successful worker start, and now also cleans up any partially-started local worker so retries begin from a clean slate. ## Model Used - Initial implementation: OpenAI GPT-5 via Codex local coding agent, with repository tool use and command execution. - Review-feedback follow-ups (driver-kind gate, partial-runtime teardown, expanded regression tests): Claude Fable 5 (claude-fable-5) via Claude Code, with repository tool use and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>