Files
PaperClipAI/doc/MCP-RUNTIME-OPERATIONS.md
T
DottaandPaperclip 3db2e6bdd2 feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 8/8 and focuses on end-to-end coverage,
operator docs, evals, and release notes
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack

## Linked Issues or Issue Description

- Related parity reference: #9534
- Problem: The complete stack needs discoverable browser scenarios,
operator guidance, threat modeling, eval coverage, and a parity proof
before merge.
- Proposed solution: Adds MCP user-story and Smoke Lab e2e suites,
docs/evals/release notes, the skill update, and the root e2e driver
script registration.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/07-ui-apps-activation`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: QA for flag audit and e2e/browser acceptance;
Greptile on every PR.

## What Changed

- Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release
notes, the skill update, and the root e2e driver script registration.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.

## Verification

- `pnpm typecheck`
- `node --check scripts/e2e-mcp-user-stories.mjs`
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
--list` — 43 tests discovered
- `git diff pap10341-split/08-e2e-docs
6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes)

## Risks

- Browser suites depend on runtime services and environment setup; this
PR validates discovery locally while QA owns full flag-on/flag-off
execution.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Stack Coordination

- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-14 15:48:57 -05:00

8.1 KiB

MCP Runtime Operations

This runbook covers Paperclip Tools & Access runtime slots for MCP connections. It is written for board and CloudOps operators responding to stuck local stdio slots, degraded remote HTTP connections, capacity deferrals, restart storms, and secret-resolution failures.

Do not print raw bearer tokens, gateway session tokens, credential headers, environment variables, or secret values while following this runbook. The APIs below return redacted state and audit metadata; keep shell tracing disabled when exporting credentials.

Tool action approvals require PAPERCLIP_TOOL_ACTION_SIGNING_SECRET to be set independently from auth/JWT secrets. Rotate it deliberately: changing it invalidates outstanding signed tool-action approvals, so drain or reject pending approvals before rotation.

Support Matrix

Transport Local trusted Hosted cloud / public authenticated Notes
remote_http Supported Supported Preferred production path. Paperclip proxies calls through the gateway with policy, audit, timeout, and redaction controls.
local_stdio Supported through approved templates and supervised runtime slots Supported only when an explicitly trusted MCP runtime worker/host is configured Set PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST or PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST only for a worker that is allowed to supervise local processes. Do not enable arbitrary agent-supplied commands.

Metrics

The board runtime health API summarizes one-hour event windows plus current durable slot state:

curl -fsS \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq .

Metrics surfaced there include:

  • Current slot counts: active, starting, running, idle, failed, stopped.
  • Stuck slot counts: starting/running slots without progress for 5 minutes.
  • Runtime events: capacity deferrals, restart attempts, restart suppression, idle evictions.
  • Tool-call health: call count, timeout count/rate, failure count/rate, average latency, p95 latency.
  • Connection health: active, disabled, degraded, remote_http, and local_stdio connection counts.
  • Secret failures: missing-secret failures in the last hour.
  • Audit write failures: durable audit_write_failed counter increments whenever MCP audit-event persistence fails.

Alerts

Alert Severity Suggested threshold First responder action
mcp_runtime_stuck_starting_slot Critical Any starting slot older than 5 minutes Inspect slot health/logs, stop the slot, restart it once, then disable the connection if it sticks again.
mcp_runtime_stuck_running_slot Critical Any running slot with no progress for 5 minutes Inspect recent audit events and active calls; restart only after confirming no healthy call is still in progress.
mcp_runtime_high_timeout_rate Warning/Critical Warning at >=3 timeouts and >=10% in 1 hour; critical at >=10 timeouts or >=25% Check upstream MCP health, runtime capacity, and gateway audit failures before retrying workloads.
mcp_runtime_high_error_rate Warning/Critical Warning at >=5 failures and >=10% in 1 hour; critical at >=10 failures or >=25% Group audit failures by reasonCode, then fix credentials/config or disable the affected connection.
mcp_runtime_capacity_deferrals_repeated Warning/Critical Warning at >=3 capacity deferrals in 1 hour; critical at >=10 Stop idle/stale slots, reduce noisy workloads, or raise slot caps only after confirming host capacity.
mcp_runtime_restart_storm Warning/Critical Warning at >=3 restarts in 1 hour; critical on any restart suppression Stop the slot, inspect stderr/audit reason codes, and keep the connection disabled until the template/upstream is fixed.
mcp_runtime_connection_health_degraded Warning/Critical Any active enabled connection with degraded/failed/missing-secret health, or any disabled enabled-path connection Run health check, refresh catalog after recovery, or keep the connection disabled and route agents to alternatives.
mcp_runtime_missing_secret_failures Warning/Critical Warning on any missing-secret failure; critical at >=3 in 1 hour Check secret bindings and provider health without revealing secret values; rotate or rebind missing secrets.
mcp_runtime_audit_write_failures Critical Any audit write failure Treat as a control-plane incident; restore DB/audit durability before retrying tool workloads.

Diagnose A Stuck Slot

  1. Read the health summary and note firing alert names:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
    
  2. List durable runtime slots:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots" | jq .
    
  3. Inspect recent gateway audit events without printing secrets:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/tool-gateway/audit?companyId=$COMPANY_ID&limit=100" \
      | jq '[.[] | {createdAt, action, entityType, entityId, reasonCode: .details.reasonCode, tool: .details.tool, durationMs: .details.durationMs}]'
    
  4. Identify the affected slotId, connectionId, reasonCode, and whether the slot is starting, running, idle, failed, or stopped.

Clear A Stuck Slot

Stop the slot first when it is stale, idle, failed, or confirmed not to be serving a healthy active call:

curl -fsS -X POST \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/stop" \
  -d '{}' | jq .

Restart once when the template/config is expected to recover:

curl -fsS -X POST \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/restart" \
  -d '{}' | jq .

If restart suppression fires, do not keep retrying. Disable the connection:

curl -fsS -X PATCH \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID" \
  -d '{"enabled":false,"status":"disabled"}' | jq '{id, name, enabled, status, healthStatus}'

Verify Recovery

  1. Run the connection health check:

    curl -fsS -X POST \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      -H "Content-Type: application/json" \
      "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/health-check" \
      -d '{}' | jq '{connection: {id: .connection.id, healthStatus: .connection.healthStatus, healthMessage: .connection.healthMessage}, runtimeSlot}'
    
  2. Refresh the catalog after a remote endpoint or stdio template recovers:

    curl -fsS -X POST \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      -H "Content-Type: application/json" \
      "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/catalog/refresh" \
      -d '{}' | jq '{discoveredCount, quarantinedCount}'
    
  3. Re-read runtime health:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
    

Recovery is complete when stuck-slot alerts clear, timeout/error rates return below threshold, the connection is healthy or intentionally disabled, and audit events show no new restart suppression or capacity deferrals.

Verification Coverage

Automated coverage includes:

  • A synthetic degraded runtime-health scenario in server/src/__tests__/tool-access-service.test.ts that creates a stale running slot, degraded connection, timeout event, capacity deferral, and restart suppression.
  • A durable audit-write failure scenario in server/src/__tests__/tool-access-service.test.ts that verifies mcp_runtime_audit_write_failures fires from the counter path.
  • A gateway runtime recovery scenario in server/src/__tests__/tool-gateway.test.ts that recovers a stuck local stdio slot before reuse.