## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents discover governed tools through the MCP gateway. > - A listing repeated policy and full run-row reads for each catalog tool. > - Parallel listings multiplied those allocations during run startup. > - One 900-tool baseline listing used 11,489 queries and about 1.7 GiB of extra heap in a fixture. > - This pull request shares reads within a listing and bounds whole listings across the process. > - The benefit is lower discovery memory use while execution still checks current policy. ## Linked Issues or Issue Description Refs #13115. Its on-demand target change affects the same listing loop. **What happened?** MCP discovery repeated roughly 13 reads per tool. Full run snapshots and repeated connection configurations caused large allocations. Per-listing bounds alone did not limit concurrent listings across gateways. **Expected behavior** Discovery reads shared inputs once per listing. The process bounds active and queued listings. Catalog payload size and policy evaluation still grow with the catalog. Tool execution checks current access rules. **Steps to reproduce** Create a company with a remote MCP connection, 900 catalog tools, large schemas, and a large run snapshot. Send concurrent tools/list requests using a run-bound gateway token. Run the committed benchmark for a deterministic reproduction. **Paperclip version or commit** Baseline:f2e0f19630. Measured on Node 26.4.0 on macOS. ## What Changed - Keep Michael Nguyen's per-listing policy cache, scalar policy reads, 16-decision bound, and catalog/connection split under repeatable read. - Admit two whole listings per process and queue at most 32. Return 503 with tool_discovery_busy when full. - Propagate disconnects to discovery. Stop scheduling new reads and drain started reads before freeing the slot. - Project context IDs in gateway authentication and GitHub runtime discovery. Avoid loading task descriptions and results. - Keep discovery audit counts and a SHA-256 digest instead of the full name list. Retain access and activity audit records. - Sweep expired named gateway tokens at startup and on the existing scheduler. Delete at most 500 per pass. Add an idempotent expiry-index migration. Resolve audit token references atomically so cleanup cannot break admitted requests. - Return GET 405 with Allow: POST on stateless MCP gateway and runtime-tools endpoints. Keep runtime-tools authentication. - Add red-green regressions, HTTP concurrency and policy-revocation coverage, and browser approval assertions. - Commit the benchmark harness, raw results, and resource-bound documentation. ## Verification - Baseline listing regressions failed with 722 queries for 50 tools and 6,422 for 500. The connection-row duplication regression also failed. - Follow-up regressions failed before the fixes for full snapshot reads, abandoned listings, full name-list audits, and expired tokens. - Listing and scheduler tests: 14 pass. Existing gateway/policy suites passed after preserving the cleanup return contract. - Two HTTP journeys pass: initialize, GET/SSE rejection, 16 concurrent 500-tool listings, provider call, policy revocation, and denied retry. Token cleanup during provider dispatch also completes successfully and blocks subsequent requests. - pnpm test:e2e tests/e2e/mcp-user-stories.spec.ts --grep '@mcp-runnable': 8 pass. The approval journey clicks Allow once in the browser. Review screenshots wait for loaded content. - pnpm -r typecheck: passes. pnpm build: passes. - The general server group completed with 14,755 passes and 18 failures. The 17 startup mock failures were fixed; all 21 startup tests pass on rerun. The one Discord timing failure passed on the unchanged baseline and on rerun (74 tests). UI and CLI groups pass 7,632 tests. Shared and skill groups pass 853 tests. The remaining database and adapter groups pass 3,010 tests with one worker after a macOS shared-memory limit interrupted a parallel run. All 149 serialized route files pass (2,762 tests). - At 900 tools, one listing falls from 11,489 to 36 queries and from about 1.7 GiB to 37 MiB of extra heap. Sixteen concurrent listings used 169–187 MiB of extra heap. Four connections used 39 queries per listing. - Latest-head verification: 54 successful checks and two expected skips on5c0793090c. Fresh Greptile review: 5/5 with no unresolved threads. - Reproduce with server/scripts/benchmark-tool-gateway-listing.ts. See doc/mcp-discovery-performance.md and doc/benchmarks/2026-10-01-mcp-discovery.json. ## Risks - The process-wide FIFO queue can increase discovery latency. Excess callers must retry 503 responses. - Cancellation applies to discovery. Started database reads finish before their slot is released. - Audit consumers must use visibleToolCount and visibleToolsHash instead of visibleTools. - The expiry index can briefly lock the token table during migration. Sweeps preserve unexpired and non-expiring tokens. - Measurements use isolated fixtures and deterministic providers. They do not establish a production heap limit or affected installation count. Rate-limit reads remain uncached. ## Model Used - Original listing optimization: Anthropic Claude Opus 5.5 (claude-opus-5-5), Claude Code, extended thinking and tool use, as recorded by the original author. - Follow-up fixes and verification: OpenAI Codex, GPT-6, with shell execution, file edits, database fixtures, and browser tests. The exact serving model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
4.3 KiB
MCP discovery resource bounds
Tool discovery shares policy inputs for one listing and projects context IDs in SQL. It does not load task descriptions or run results to decide access. The cache ends with the listing. Tool execution reads current policy and current rate counters.
The control-plane process admits two whole listings at a time and queues at most
32 more. Excess requests receive HTTP 503 with tool_discovery_busy. Disconnected
listing requests leave the queue or stop scheduling new decisions. Already-started
reads drain before the admission slot is released. This applies only to read-only
discovery; a disconnected tool invocation is not automatically cancelled.
Discovery audit rows retain the visible tool count and a SHA-256 digest of sorted names. They do not retain the full name list. Named gateway tokens expire normally; startup and scheduler sweeps delete at most 500 expired tokens per pass using the expiry index. Tokens with no expiry and unexpired tokens remain available. The separate access and activity audit records remain available. Audit inserts resolve the token reference atomically and lock a surviving token row for that statement. If cleanup already removed it, the reference is null and the original token ID remains in audit details. An admitted request can complete; later requests still fail authentication after expiry.
The stateless gateway and runtime-tools MCP endpoints return HTTP 405 with
Allow: POST for GET instead of returning JSON as if it were an SSE stream.
The runtime-tools endpoint still validates its caller before that response.
See the MCP transport contract.
Reproduce the measurements
Build workspace dependencies, then run from server/:
node --expose-gc --max-old-space-size=4096 --import tsx scripts/benchmark-tool-gateway-listing.ts --tools 50,300,900 --parallel 1,16 --repeat 3
node --expose-gc --max-old-space-size=4096 --import tsx scripts/benchmark-tool-gateway-listing.ts --tools 900 --connections 4 --parallel 16 --repeat 3
The harness uses a throwaway embedded PostgreSQL database. It makes no model or
provider calls. Each fixture has a 400 KB run snapshot, 140 KB result, approximately
5 KB input schemas, 15 KB connection configurations, one selector per catalog tool,
and 373 policies. Use --implementation /absolute/path/to/tool-gateway.ts to replay
an earlier implementation, with its corresponding policy-service import.
Verification on 2026-10-01
Node 26.4.0 on macOS; forced GC before each measurement; heap/RSS sampled every
2 ms. Baseline: f2e0f196308629ef05c7f65782243e1713b208f0.
Raw measurements retain every sample.
| Scenario | Queries | Peak extra heap | Elapsed time |
|---|---|---|---|
| Baseline, 900 tools, one listing | 11,489 | 1,680–1,698 MiB | 3.54–3.63 s |
| Fixed, 900 tools, one listing | 36 | 37–37 MiB | 0.24–0.24 s |
| Fixed, 900 tools, 16 listings | 576 total | 169–187 MiB | 3.63–3.65 s |
| Fixed, 900 tools, four connections, 16 listings | 624 total | 146–174 MiB | 3.69–3.72 s |
Before the listing optimization, the 50-versus-500-tool regression test failed with 722 versus 6,422 statements. The connection-row duplication test also failed. The follow-up regressions failed for full run-row discovery reads, abandoned requests, full name-list audit records, and expired named tokens. They pass with the fixes. HTTP coverage exercises initialize, GET/SSE rejection, 16 concurrent 500-tool listings, an allowed remote call, policy revocation after discovery, and the denied subsequent call. The remote provider response is deterministic.
The eight runnable MCP browser stories pass against a throwaway production UI and server, including governed execution and approval journeys.
These are fixture measurements, not a replay of the incident database. RSS includes module/runtime overhead and can stay high after V8 frees objects. Query counts depend on connection/profile sets; uncached rate-limit checks add queries. Catalog payload memory and policy-matching CPU still grow with the catalog. The measurements do not establish a safe production heap limit or the share of active installations affected. Verify restart bursts and sustained traffic on the actual deployment before reducing its temporary heap allowance.