Files
PaperClipAI/server
Devin FoleyandPaperclip 3367b75ccc fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A host can stop an idle instance to reduce unused compute.
> - Zero active runs do not prove that requests, saved work, or cleanup
are complete.
> - A client can disconnect while its request still writes, and failed
cleanup can remain only in memory or on disk.
> - This pull request holds new admission and checks accepted work,
cleanup, and persisted work under an owned drain.
> - The host gets an empty report only while those checks remain valid.

## Linked Issues or Issue Description

Refs #13413. That companion PR uses the same owned-hold vocabulary for
runtime services. This PR covers HTTP requests, scheduler work,
accounting, cleanup, and the instance work inventory. It does not
include the preview gateway changes.

**What existing behavior does this improve?**

The instance task-drain API reports process counters. It does not prove
that a host can safely stop the instance for idle sleep.

**Current behavior**

A quiet run set can coexist with an unfinished request, cleanup after a
completed run, future work, or a failed accounting write.

**Proposed behavior**

Provide a bounded owned idle hold. Block new ingress, track accepted
handler promises, inspect durable and local work, and return `none` only
when the same hold stays quiet through the checks. Keep normal
deployment drains compatible.

**Reason and benefit**

Hosts can identify eligible idle instances without treating a disconnect
or failed cleanup write as completed work.

## What Changed

- Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked
release to task drain. Existing holds cannot be replaced by another API
request.
- Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new
WebSocket upgrades. Track async handlers in nested Express routers and
error middleware until they settle, even after the response or client
disconnect.
- Count accepted live-event WebSocket authentication through settlement,
even after disconnects. Count detached built-in agent, managed-home and
runtime-service startup reconciliation after readiness.
- Keep health and control mutations tracked. Count control-request
authentication separately from the read-only report, including
concurrent user/company/membership writes.
- Pause new scheduler admissions during idle holds. Count work already
in flight, including database backups, and reject scans whose work
generation changes. Periodic backups block sleep without a host wake
schedule.
- Inspect accounting and orphan-cleanup spool directories without
skipping temporary or malformed entries. Retain orphan tokens through
queue splices, flush failures, and buffer overflow. Keep failed usage
capture counted until its database failure fence is written.
- Check persisted work across companies in a bounded read-only
transaction. Include deferred agent-file cleanup and saved watchdogs
whose watched issues are complete. Enabled plugins and unsupported
retained work remain blockers.
- Require the exact idle owner and completed startup tracking before
returning an empty report. Keep reports free of tenant details.
- Document the hosting protocol, retry behavior, conservative blockers,
and the remaining external provider-stop race.

## Verification

- Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on
master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper
extractions are preserved. GitHub confirms no merge conflicts.
- Focused heartbeat renderer/run-log, drain, control-auth, admission,
route and PostgreSQL inventory checks: 234 tests passed in 10 suites
after the rebase.
- The POST task-drain contract includes the expected `409` conflict
response. The 84 OpenAPI and instance-settings route tests and server
typecheck passed after that final documentation fix.
- Final accepted-upgrade/startup regression run: 104 tests passed in
five suites, including success and failure after readiness or
disconnect. Final server typecheck also passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm install --frozen-lockfile` and the PR diff secret scan passed.
No dependency or lockfile changes are added by this PR.
- The prior verification covered HTTP disconnects and early responses,
async error handlers, concurrent authentication writes, signed
bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and
failed accounting fences. Those tests passed.
- The last full local `pnpm test:run` stopped in the general-server
phase with 16,658 passing tests and five skill/connector fixture
failures caused by an ancestor workspace skill directory. That full
local run preceded this rebase and has not been repeated for the import
conflict. Full current-head CI passed: 53 successful checks and two
conditional skips, with no failures.
- All eight review findings are fixed and their threads resolved,
including accepted upgrade authentication, detached startup writes and
the POST conflict contract. Current-head Apex review is 5/5 with no new
findings; all eight review threads remain resolved.
- No live provider stop or production change was performed.

## Risks

- The hosting controller must use the owned protocol and hold external
admission through its final validation and provider stop. A legacy drain
cannot authorize idle sleep. During an idle hold, new requests receive
503 with `Retry-After: 1`; the host must handle queueing or retry before
enabling this path.
- This is a single-process protocol. An unexpected restart after the
last validation can race an external provider stop. The host must
serialize deploy/wake/sleep operations and bind the validation to the
instance it stops. Multiple replicas need shared fencing.
- Some retained state conservatively prevents sleep, including every
enabled plugin. This PR does not promise that every inactive instance
becomes eligible.
- The HTTP adapter uses Express 5 router layers. Real Express tests
cover nested routes, errors and disconnects. New routes must register
before tracking is installed. Detached work must have durable state or
explicit work tracking.
- Unrecoverable in-memory cleanup debt keeps the instance awake until
reconciliation. This change does not make such debt survive an unplanned
process crash. No schema migration or provider configuration change is
included.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use
and code execution. The exact deployment variant/model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; the full local
fixture limitation and clean full CI result are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:35:45 -07:00
..