mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents discover governed tools through the MCP gateway. > - A listing repeated policy and full run-row reads for each catalog tool. > - Parallel listings multiplied those allocations during run startup. > - One 900-tool baseline listing used 11,489 queries and about 1.7 GiB of extra heap in a fixture. > - This pull request shares reads within a listing and bounds whole listings across the process. > - The benefit is lower discovery memory use while execution still checks current policy. ## Linked Issues or Issue Description Refs #13115. Its on-demand target change affects the same listing loop. **What happened?** MCP discovery repeated roughly 13 reads per tool. Full run snapshots and repeated connection configurations caused large allocations. Per-listing bounds alone did not limit concurrent listings across gateways. **Expected behavior** Discovery reads shared inputs once per listing. The process bounds active and queued listings. Catalog payload size and policy evaluation still grow with the catalog. Tool execution checks current access rules. **Steps to reproduce** Create a company with a remote MCP connection, 900 catalog tools, large schemas, and a large run snapshot. Send concurrent tools/list requests using a run-bound gateway token. Run the committed benchmark for a deterministic reproduction. **Paperclip version or commit** Baseline:f2e0f19630. Measured on Node 26.4.0 on macOS. ## What Changed - Keep Michael Nguyen's per-listing policy cache, scalar policy reads, 16-decision bound, and catalog/connection split under repeatable read. - Admit two whole listings per process and queue at most 32. Return 503 with tool_discovery_busy when full. - Propagate disconnects to discovery. Stop scheduling new reads and drain started reads before freeing the slot. - Project context IDs in gateway authentication and GitHub runtime discovery. Avoid loading task descriptions and results. - Keep discovery audit counts and a SHA-256 digest instead of the full name list. Retain access and activity audit records. - Sweep expired named gateway tokens at startup and on the existing scheduler. Delete at most 500 per pass. Add an idempotent expiry-index migration. Resolve audit token references atomically so cleanup cannot break admitted requests. - Return GET 405 with Allow: POST on stateless MCP gateway and runtime-tools endpoints. Keep runtime-tools authentication. - Add red-green regressions, HTTP concurrency and policy-revocation coverage, and browser approval assertions. - Commit the benchmark harness, raw results, and resource-bound documentation. ## Verification - Baseline listing regressions failed with 722 queries for 50 tools and 6,422 for 500. The connection-row duplication regression also failed. - Follow-up regressions failed before the fixes for full snapshot reads, abandoned listings, full name-list audits, and expired tokens. - Listing and scheduler tests: 14 pass. Existing gateway/policy suites passed after preserving the cleanup return contract. - Two HTTP journeys pass: initialize, GET/SSE rejection, 16 concurrent 500-tool listings, provider call, policy revocation, and denied retry. Token cleanup during provider dispatch also completes successfully and blocks subsequent requests. - pnpm test:e2e tests/e2e/mcp-user-stories.spec.ts --grep '@mcp-runnable': 8 pass. The approval journey clicks Allow once in the browser. Review screenshots wait for loaded content. - pnpm -r typecheck: passes. pnpm build: passes. - The general server group completed with 14,755 passes and 18 failures. The 17 startup mock failures were fixed; all 21 startup tests pass on rerun. The one Discord timing failure passed on the unchanged baseline and on rerun (74 tests). UI and CLI groups pass 7,632 tests. Shared and skill groups pass 853 tests. The remaining database and adapter groups pass 3,010 tests with one worker after a macOS shared-memory limit interrupted a parallel run. All 149 serialized route files pass (2,762 tests). - At 900 tools, one listing falls from 11,489 to 36 queries and from about 1.7 GiB to 37 MiB of extra heap. Sixteen concurrent listings used 169–187 MiB of extra heap. Four connections used 39 queries per listing. - Latest-head verification: 54 successful checks and two expected skips on5c0793090c. Fresh Greptile review: 5/5 with no unresolved threads. - Reproduce with server/scripts/benchmark-tool-gateway-listing.ts. See doc/mcp-discovery-performance.md and doc/benchmarks/2026-10-01-mcp-discovery.json. ## Risks - The process-wide FIFO queue can increase discovery latency. Excess callers must retry 503 responses. - Cancellation applies to discovery. Started database reads finish before their slot is released. - Audit consumers must use visibleToolCount and visibleToolsHash instead of visibleTools. - The expiry index can briefly lock the token table during migration. Sweeps preserve unexpired and non-expiring tokens. - Measurements use isolated fixtures and deterministic providers. They do not establish a production heap limit or affected installation count. Rate-limit reads remain uncached. ## Model Used - Original listing optimization: Anthropic Claude Opus 5.5 (claude-opus-5-5), Claude Code, extended thinking and tool use, as recorded by the original author. - Follow-up fixes and verification: OpenAI Codex, GPT-6, with shell execution, file edits, database fixtures, and browser tests. The exact serving model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
96 lines
4.8 KiB
TypeScript
96 lines
4.8 KiB
TypeScript
/**
|
|
* Reproducible discovery load probe. Uses a throwaway embedded PostgreSQL DB;
|
|
* no provider calls, production data, or model inference.
|
|
* node --expose-gc --import tsx scripts/benchmark-tool-gateway-listing.ts
|
|
* --tools 50,300,900 --parallel 1,16 --repeat 3
|
|
* --implementation accepts a local tool-gateway module for before/after replay.
|
|
*/
|
|
import { randomUUID } from "node:crypto";
|
|
import { pathToFileURL } from "node:url";
|
|
import { eq } from "drizzle-orm";
|
|
import {
|
|
createDb, heartbeatRuns, toolCatalogEntries, toolConnections, toolPolicies,
|
|
toolProfileEntries, startEmbeddedPostgresTestDatabase,
|
|
} from "@paperclipai/db";
|
|
import { createListingFixture, recordingDb } from "../src/__tests__/helpers/tool-gateway-listing-fixture.js";
|
|
import type { createToolGatewayService } from "../src/services/tool-gateway.js";
|
|
|
|
function option(name: string, fallback: string) {
|
|
const index = process.argv.indexOf(name);
|
|
return index < 0 ? fallback : process.argv[index + 1] ?? fallback;
|
|
}
|
|
function counts(value: string) {
|
|
const result = value.split(",").map(Number);
|
|
if (result.some((n) => !Number.isInteger(n) || n < 1)) throw new Error("Expected positive integer counts");
|
|
return result;
|
|
}
|
|
const toolCounts = counts(option("--tools", "50,300,900"));
|
|
const parallels = counts(option("--parallel", "1,16"));
|
|
const repeats = counts(option("--repeat", "3"))[0]!;
|
|
const connections = counts(option("--connections", "1"))[0]!;
|
|
const implementation = option("--implementation", "");
|
|
const moduleUrl = implementation ? pathToFileURL(implementation) : new URL("../src/services/tool-gateway.js", import.meta.url);
|
|
const factory = (await import(moduleUrl.href)).createToolGatewayService as typeof createToolGatewayService;
|
|
const temp = await startEmbeddedPostgresTestDatabase("paperclip-listing-benchmark-");
|
|
const db = createDb(temp.connectionString);
|
|
try {
|
|
for (const tools of toolCounts) {
|
|
const fixture = await createListingFixture(db, tools, { connectionCount: connections });
|
|
await db.update(heartbeatRuns).set({
|
|
contextSnapshot: { issueId: fixture.issue.id, projectId: fixture.project.id, taskMarkdown: "x".repeat(400_000) },
|
|
resultJson: { summary: "x".repeat(140_000) },
|
|
}).where(eq(heartbeatRuns.id, fixture.run.id));
|
|
await db.update(toolConnections).set({ config: { url: "https://8.8.8.8/mcp", notes: "x".repeat(15_000) } })
|
|
.where(eq(toolConnections.companyId, fixture.company.id));
|
|
await db.update(toolCatalogEntries).set({ inputSchema: {
|
|
type: "object", properties: { query: { type: "string", description: "x".repeat(5_000) } },
|
|
} }).where(eq(toolCatalogEntries.companyId, fixture.company.id));
|
|
await db.insert(toolProfileEntries).values(fixture.entries.map((entry) => ({
|
|
companyId: fixture.company.id, profileId: fixture.namedGateway.profileId,
|
|
selectorType: "catalog_entry" as const, effect: "include" as const, catalogEntryId: entry.id,
|
|
})));
|
|
await db.insert(toolPolicies).values(Array.from({ length: 371 }, (_, i) => ({
|
|
companyId: fixture.company.id, name: `Unmatched fixture policy ${i}`, policyType: "block" as const,
|
|
selectors: { catalogEntryId: randomUUID() }, priority: 100 + i,
|
|
})));
|
|
const input = { gatewayId: fixture.namedGateway.id, bearerToken: fixture.token.token };
|
|
await factory(db).listToolsForNamedGateway(input);
|
|
for (const parallel of parallels) {
|
|
for (let iteration = 1; iteration <= repeats; iteration += 1) {
|
|
global.gc?.();
|
|
const before = process.memoryUsage();
|
|
let peakHeap = before.heapUsed;
|
|
let peakRss = before.rss;
|
|
const sample = () => {
|
|
const memory = process.memoryUsage();
|
|
peakHeap = Math.max(peakHeap, memory.heapUsed);
|
|
peakRss = Math.max(peakRss, memory.rss);
|
|
};
|
|
const timer = setInterval(sample, 2);
|
|
const recorder = recordingDb(db);
|
|
const started = performance.now();
|
|
let visible = 0;
|
|
try {
|
|
const listings = await Promise.all(Array.from({ length: parallel }, () =>
|
|
factory(recorder.db).listToolsForNamedGateway(input)));
|
|
visible = listings[0]!.length;
|
|
sample();
|
|
} finally {
|
|
clearInterval(timer);
|
|
}
|
|
const elapsedMs = performance.now() - started;
|
|
global.gc?.();
|
|
const after = process.memoryUsage();
|
|
console.log(JSON.stringify({
|
|
tools, connections, parallel, iteration, visible, queries: recorder.statements.length,
|
|
elapsedMs: Math.round(elapsedMs), peakHeapIncreaseMiB: Math.round((peakHeap - before.heapUsed) / 1_048_576),
|
|
peakRssMiB: Math.round(peakRss / 1_048_576), retainedHeapIncreaseMiB: Math.round((after.heapUsed - before.heapUsed) / 1_048_576),
|
|
forcedGc: Boolean(global.gc), fixture: { snapshotBytes: 400_000, schemaBytes: 5_000, policyCount: 373 },
|
|
}));
|
|
}
|
|
}
|
|
}
|
|
} finally {
|
|
await temp.cleanup();
|
|
}
|