## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
8.1 KiB
MCP Runtime Operations
This runbook covers Paperclip Tools & Access runtime slots for MCP connections. It is written for board and CloudOps operators responding to stuck local stdio slots, degraded remote HTTP connections, capacity deferrals, restart storms, and secret-resolution failures.
Do not print raw bearer tokens, gateway session tokens, credential headers, environment variables, or secret values while following this runbook. The APIs below return redacted state and audit metadata; keep shell tracing disabled when exporting credentials.
Tool action approvals require PAPERCLIP_TOOL_ACTION_SIGNING_SECRET to be set independently from auth/JWT secrets. Rotate it deliberately: changing it invalidates outstanding signed tool-action approvals, so drain or reject pending approvals before rotation.
Support Matrix
| Transport | Local trusted | Hosted cloud / public authenticated | Notes |
|---|---|---|---|
remote_http |
Supported | Supported | Preferred production path. Paperclip proxies calls through the gateway with policy, audit, timeout, and redaction controls. |
local_stdio |
Supported through approved templates and supervised runtime slots | Supported only when an explicitly trusted MCP runtime worker/host is configured | Set PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST or PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST only for a worker that is allowed to supervise local processes. Do not enable arbitrary agent-supplied commands. |
Metrics
The board runtime health API summarizes one-hour event windows plus current durable slot state:
curl -fsS \
-H "Authorization: Bearer $BOARD_API_KEY" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq .
Metrics surfaced there include:
- Current slot counts: active, starting, running, idle, failed, stopped.
- Stuck slot counts: starting/running slots without progress for 5 minutes.
- Runtime events: capacity deferrals, restart attempts, restart suppression, idle evictions.
- Tool-call health: call count, timeout count/rate, failure count/rate, average latency, p95 latency.
- Connection health: active, disabled, degraded,
remote_http, andlocal_stdioconnection counts. - Secret failures: missing-secret failures in the last hour.
- Audit write failures: durable
audit_write_failedcounter increments whenever MCP audit-event persistence fails.
Alerts
| Alert | Severity | Suggested threshold | First responder action |
|---|---|---|---|
mcp_runtime_stuck_starting_slot |
Critical | Any starting slot older than 5 minutes | Inspect slot health/logs, stop the slot, restart it once, then disable the connection if it sticks again. |
mcp_runtime_stuck_running_slot |
Critical | Any running slot with no progress for 5 minutes | Inspect recent audit events and active calls; restart only after confirming no healthy call is still in progress. |
mcp_runtime_high_timeout_rate |
Warning/Critical | Warning at >=3 timeouts and >=10% in 1 hour; critical at >=10 timeouts or >=25% | Check upstream MCP health, runtime capacity, and gateway audit failures before retrying workloads. |
mcp_runtime_high_error_rate |
Warning/Critical | Warning at >=5 failures and >=10% in 1 hour; critical at >=10 failures or >=25% | Group audit failures by reasonCode, then fix credentials/config or disable the affected connection. |
mcp_runtime_capacity_deferrals_repeated |
Warning/Critical | Warning at >=3 capacity deferrals in 1 hour; critical at >=10 | Stop idle/stale slots, reduce noisy workloads, or raise slot caps only after confirming host capacity. |
mcp_runtime_restart_storm |
Warning/Critical | Warning at >=3 restarts in 1 hour; critical on any restart suppression | Stop the slot, inspect stderr/audit reason codes, and keep the connection disabled until the template/upstream is fixed. |
mcp_runtime_connection_health_degraded |
Warning/Critical | Any active enabled connection with degraded/failed/missing-secret health, or any disabled enabled-path connection | Run health check, refresh catalog after recovery, or keep the connection disabled and route agents to alternatives. |
mcp_runtime_missing_secret_failures |
Warning/Critical | Warning on any missing-secret failure; critical at >=3 in 1 hour | Check secret bindings and provider health without revealing secret values; rotate or rebind missing secrets. |
mcp_runtime_audit_write_failures |
Critical | Any audit write failure | Treat as a control-plane incident; restore DB/audit durability before retrying tool workloads. |
Diagnose A Stuck Slot
-
Read the health summary and note firing alert names:
curl -fsS \ -H "Authorization: Bearer $BOARD_API_KEY" \ "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}' -
List durable runtime slots:
curl -fsS \ -H "Authorization: Bearer $BOARD_API_KEY" \ "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots" | jq . -
Inspect recent gateway audit events without printing secrets:
curl -fsS \ -H "Authorization: Bearer $BOARD_API_KEY" \ "$PAPERCLIP_URL/api/tool-gateway/audit?companyId=$COMPANY_ID&limit=100" \ | jq '[.[] | {createdAt, action, entityType, entityId, reasonCode: .details.reasonCode, tool: .details.tool, durationMs: .details.durationMs}]' -
Identify the affected
slotId,connectionId,reasonCode, and whether the slot isstarting,running,idle,failed, orstopped.
Clear A Stuck Slot
Stop the slot first when it is stale, idle, failed, or confirmed not to be serving a healthy active call:
curl -fsS -X POST \
-H "Authorization: Bearer $BOARD_API_KEY" \
-H "Content-Type: application/json" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/stop" \
-d '{}' | jq .
Restart once when the template/config is expected to recover:
curl -fsS -X POST \
-H "Authorization: Bearer $BOARD_API_KEY" \
-H "Content-Type: application/json" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/restart" \
-d '{}' | jq .
If restart suppression fires, do not keep retrying. Disable the connection:
curl -fsS -X PATCH \
-H "Authorization: Bearer $BOARD_API_KEY" \
-H "Content-Type: application/json" \
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID" \
-d '{"enabled":false,"status":"disabled"}' | jq '{id, name, enabled, status, healthStatus}'
Verify Recovery
-
Run the connection health check:
curl -fsS -X POST \ -H "Authorization: Bearer $BOARD_API_KEY" \ -H "Content-Type: application/json" \ "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/health-check" \ -d '{}' | jq '{connection: {id: .connection.id, healthStatus: .connection.healthStatus, healthMessage: .connection.healthMessage}, runtimeSlot}' -
Refresh the catalog after a remote endpoint or stdio template recovers:
curl -fsS -X POST \ -H "Authorization: Bearer $BOARD_API_KEY" \ -H "Content-Type: application/json" \ "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/catalog/refresh" \ -d '{}' | jq '{discoveredCount, quarantinedCount}' -
Re-read runtime health:
curl -fsS \ -H "Authorization: Bearer $BOARD_API_KEY" \ "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
Recovery is complete when stuck-slot alerts clear, timeout/error rates return below threshold, the connection is healthy or intentionally disabled, and audit events show no new restart suppression or capacity deferrals.
Verification Coverage
Automated coverage includes:
- A synthetic degraded runtime-health scenario in
server/src/__tests__/tool-access-service.test.tsthat creates a stale running slot, degraded connection, timeout event, capacity deferral, and restart suppression. - A durable audit-write failure scenario in
server/src/__tests__/tool-access-service.test.tsthat verifiesmcp_runtime_audit_write_failuresfires from the counter path. - A gateway runtime recovery scenario in
server/src/__tests__/tool-gateway.test.tsthat recovers a stuck local stdio slot before reuse.