mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 06:15:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A host can stop an idle instance to reduce unused compute. > - Zero active runs do not prove that requests, saved work, or cleanup are complete. > - A client can disconnect while its request still writes, and failed cleanup can remain only in memory or on disk. > - This pull request holds new admission and checks accepted work, cleanup, and persisted work under an owned drain. > - The host gets an empty report only while those checks remain valid. ## Linked Issues or Issue Description Refs #13413. That companion PR uses the same owned-hold vocabulary for runtime services. This PR covers HTTP requests, scheduler work, accounting, cleanup, and the instance work inventory. It does not include the preview gateway changes. **What existing behavior does this improve?** The instance task-drain API reports process counters. It does not prove that a host can safely stop the instance for idle sleep. **Current behavior** A quiet run set can coexist with an unfinished request, cleanup after a completed run, future work, or a failed accounting write. **Proposed behavior** Provide a bounded owned idle hold. Block new ingress, track accepted handler promises, inspect durable and local work, and return `none` only when the same hold stays quiet through the checks. Keep normal deployment drains compatible. **Reason and benefit** Hosts can identify eligible idle instances without treating a disconnect or failed cleanup write as completed work. ## What Changed - Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked release to task drain. Existing holds cannot be replaced by another API request. - Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new WebSocket upgrades. Track async handlers in nested Express routers and error middleware until they settle, even after the response or client disconnect. - Count accepted live-event WebSocket authentication through settlement, even after disconnects. Count detached built-in agent, managed-home and runtime-service startup reconciliation after readiness. - Keep health and control mutations tracked. Count control-request authentication separately from the read-only report, including concurrent user/company/membership writes. - Pause new scheduler admissions during idle holds. Count work already in flight, including database backups, and reject scans whose work generation changes. Periodic backups block sleep without a host wake schedule. - Inspect accounting and orphan-cleanup spool directories without skipping temporary or malformed entries. Retain orphan tokens through queue splices, flush failures, and buffer overflow. Keep failed usage capture counted until its database failure fence is written. - Check persisted work across companies in a bounded read-only transaction. Include deferred agent-file cleanup and saved watchdogs whose watched issues are complete. Enabled plugins and unsupported retained work remain blockers. - Require the exact idle owner and completed startup tracking before returning an empty report. Keep reports free of tenant details. - Document the hosting protocol, retry behavior, conservative blockers, and the remaining external provider-stop race. ## Verification - Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper extractions are preserved. GitHub confirms no merge conflicts. - Focused heartbeat renderer/run-log, drain, control-auth, admission, route and PostgreSQL inventory checks: 234 tests passed in 10 suites after the rebase. - The POST task-drain contract includes the expected `409` conflict response. The 84 OpenAPI and instance-settings route tests and server typecheck passed after that final documentation fix. - Final accepted-upgrade/startup regression run: 104 tests passed in five suites, including success and failure after readiness or disconnect. Final server typecheck also passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm install --frozen-lockfile` and the PR diff secret scan passed. No dependency or lockfile changes are added by this PR. - The prior verification covered HTTP disconnects and early responses, async error handlers, concurrent authentication writes, signed bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and failed accounting fences. Those tests passed. - The last full local `pnpm test:run` stopped in the general-server phase with 16,658 passing tests and five skill/connector fixture failures caused by an ancestor workspace skill directory. That full local run preceded this rebase and has not been repeated for the import conflict. Full current-head CI passed: 53 successful checks and two conditional skips, with no failures. - All eight review findings are fixed and their threads resolved, including accepted upgrade authentication, detached startup writes and the POST conflict contract. Current-head Apex review is 5/5 with no new findings; all eight review threads remain resolved. - No live provider stop or production change was performed. ## Risks - The hosting controller must use the owned protocol and hold external admission through its final validation and provider stop. A legacy drain cannot authorize idle sleep. During an idle hold, new requests receive 503 with `Retry-After: 1`; the host must handle queueing or retry before enabling this path. - This is a single-process protocol. An unexpected restart after the last validation can race an external provider stop. The host must serialize deploy/wake/sleep operations and bind the validation to the instance it stops. Multiple replicas need shared fencing. - Some retained state conservatively prevents sleep, including every enabled plugin. This PR does not promise that every inactive instance becomes eligible. - The HTTP adapter uses Express 5 router layers. Real Express tests cover nested routes, errors and disconnects. New routes must register before tracking is installed. Detached work must have durable state or explicit work tracking. - Unrecoverable in-memory cleanup debt keeps the instance awake until reconciliation. This change does not make such debt survive an unplanned process crash. No schema migration or provider configuration change is included. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact deployment variant/model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local fixture limitation and clean full CI result are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>