Files
PaperClipAI/doc/MCP-RUNTIME-OPERATIONS.md
DottaandPaperclip c46e41e81c fix(heartbeat): validate native MCP gateway ownership (#14914)
## Thinking Path

> - Paperclip lets people manage agents and govern their tool access.
> - Native runs receive an immutable MCP tool assignment for one agent.
> - The gateway must enforce that owner when it authenticates a run
token.
> - Older gateway rows stored the owner only in metadata.
> - This change validates the gateway and profile, binds new rows, and
repairs valid older rows on reuse.
> - It also delivers each native assignment once.
> - Other agents cannot use the assignment, and explicit shared gateways
keep their configured scope.

## Linked Issues or Issue Description

Builds on #14012 by @busla (Jón Levy). That PR adds agent binding and
seven regressions. This PR carries that fix onto current master and adds
legacy authentication, profile validation, and duplicate-delivery
coverage. Related: #14864 improves discovery memory use.

**What happened?**
Native gateway creation stored an owner in metadata but left `agentId`
null. Authentication could therefore accept another agent's run token.
Managed discovery could also deliver historical native assignments
again.

**Expected behavior**
A native assignment accepts only its owner's run token. The gateway and
profile must refer to the same immutable assignment. The current
assignment enters the run configuration once.

**Steps to reproduce**
1. Run the database fixtures in `heartbeat-runtime-mcp-servers.test.ts`
on the baseline.
2. Create a native assignment and inspect its stored gateway owner.
3. Authenticate with another agent's run token, then inspect legacy
reuse and managed delivery.
4. The baseline fails six ownership and delivery cases. The fix passes
all twelve cases.

**Paperclip version or commit**
The red baseline is `f2e0f1963`. This PR is based on `cad26c6bf`, which
includes the merged discovery fix.

**Deployment mode**
Native Paperclip Runner execution and managed Codex MCP delivery.
Reproduction uses isolated database and HTTP fixtures.

## What Changed

- Store the agent owner and agent context on new native gateways.
- Validate profile and gateway assignment metadata before reuse or token
creation.
- Bind valid legacy rows with a company-scoped, null-owner update and
validate the result.
- Reject mismatched run tokens before legacy repair.
- Identify native assignments by gateway metadata, the reserved profile
key, or profile source. Reject missing or malformed provenance,
including JSON null.
- Exclude historical native assignments from managed gateway delivery.
Keep their rows for existing runs.
- Add twelve database and HTTP regressions and document the runtime
contract.

## Verification

- Red baseline: six regressions fail and four controls pass before the
initial fix. Two additional regressions reproduce metadata-loss
admission and a JSON-null TypeError before the review fix.
- All twelve ownership regressions pass on the final code, including
owner admission, cross-agent rejection, metadata loss, JSON-null HTTP
401, and explicit shared-gateway admission. Policy, listing-memory, and
discovery HTTP coverage also passes.
- Full workspace typecheck and build pass locally. Server typecheck and
compilation pass again after the review fix. The final ownership and
grant patches pass 42 combined database and HTTP regressions.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37011383657)
passes for `626a08ae66361cf586105877e24d806b1a7a9c20`: all 54 checks
succeed; two optional Storybook checks skip. This includes full
typecheck, build, all test shards, all eight E2E shards, runner
verification, and the canary dry run.
- Greptile scores that exact head 5/5. No review threads remain
unresolved.

## Risks

Invalid historical native gateway or profile metadata now rejects
authentication. Valid unbound rows are repaired only when their owner
reuses the assignment. Conflicting owners are never overwritten.
Historical rows are retained for existing runs. Explicit shared gateways
use ordinary profiles and keep their configured scopes. The reserved
native profile namespace remains agent-owned even when gateway metadata
is cleared. No schema or dependency changes are included.

## Model Used

Original fix and seven regressions in #14012: Anthropic Claude Opus 5.5,
`claude-opus-5-5`, 1M context, as reported by @busla. Extensions and
verification: OpenAI Codex (GPT-6), with reasoning, repository
inspection, code execution, and tests. This session does not expose the
exact serving model identifier or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:22:33 -05:00

9.5 KiB

MCP Runtime Operations

This runbook covers Paperclip Tools & Access runtime slots for MCP connections. It is written for board and CloudOps operators responding to stuck local stdio slots, degraded remote HTTP connections, capacity deferrals, restart storms, and secret-resolution failures.

Do not print raw bearer tokens, gateway session tokens, credential headers, environment variables, or secret values while following this runbook. The APIs below return redacted state and audit metadata; keep shell tracing disabled when exporting credentials.

Tool action approvals require PAPERCLIP_TOOL_ACTION_SIGNING_SECRET to be set independently from auth/JWT secrets. paperclipai onboard generates it for local instances, and worktree setup propagates or generates an independent value in the worktree .env; operator-managed deployments must set it explicitly. Rotate it deliberately: changing it invalidates outstanding signed tool-action approvals, so drain or reject pending approvals before rotation.

Support Matrix

Transport Local trusted Hosted cloud / public authenticated Notes
remote_http Supported Supported Preferred production path. Paperclip proxies calls through the gateway with policy, audit, timeout, and redaction controls.
local_stdio Supported through approved templates and supervised runtime slots Supported only when an explicitly trusted MCP runtime worker/host is configured Set PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST or PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST only for a worker that is allowed to supervise local processes. Do not enable arbitrary agent-supplied commands.

Native runtime gateway ownership

Paperclip Runner creates an immutable MCP assignment for one agent. Its gateway stores the owner in agentId and the agent context scope. Authentication checks that the run belongs to that owner and that the gateway and profile metadata refer to the same assignment. Invalid ownership or assignment metadata rejects the token with gateway_token_run_context_invalid.

The reserved native: profile key also identifies a native assignment if its gateway metadata is cleared. Clearing metadata cannot turn the assignment into a shared gateway. Invalid JSON metadata, including JSON null, returns the same authentication rejection. Use an ordinary profile for an explicit shared gateway.

Older native gateways can have a null agentId. Authentication still checks their metadata owner. When the owner reuses the assignment, Paperclip validates the gateway and profile before it binds the gateway to that agent. A conflicting owner is never overwritten.

Native assignments enter the run configuration once through the runtime MCP delivery path. Managed gateway discovery excludes historical native assignments. Their stored rows remain available for existing runs. Explicit company gateways continue to use their configured scopes.

Metrics

The board runtime health API summarizes one-hour event windows plus current durable slot state:

curl -fsS \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq .

Metrics surfaced there include:

  • Current slot counts: active, starting, running, idle, failed, stopped.
  • Stuck slot counts: starting/running slots without progress for 5 minutes.
  • Runtime events: capacity deferrals, restart attempts, restart suppression, idle evictions.
  • Tool-call health: call count, timeout count/rate, failure count/rate, average latency, p95 latency.
  • Connection health: active, disabled, degraded, remote_http, and local_stdio connection counts.
  • Secret failures: missing-secret failures in the last hour.
  • Audit write failures: durable audit_write_failed counter increments whenever MCP audit-event persistence fails.

Alerts

Alert Severity Suggested threshold First responder action
mcp_runtime_stuck_starting_slot Critical Any starting slot older than 5 minutes Inspect slot health/logs, stop the slot, restart it once, then disable the connection if it sticks again.
mcp_runtime_stuck_running_slot Critical Any running slot with no progress for 5 minutes Inspect recent audit events and active calls; restart only after confirming no healthy call is still in progress.
mcp_runtime_high_timeout_rate Warning/Critical Warning at >=3 timeouts and >=10% in 1 hour; critical at >=10 timeouts or >=25% Check upstream MCP health, runtime capacity, and gateway audit failures before retrying workloads.
mcp_runtime_high_error_rate Warning/Critical Warning at >=5 failures and >=10% in 1 hour; critical at >=10 failures or >=25% Group audit failures by reasonCode, then fix credentials/config or disable the affected connection.
mcp_runtime_capacity_deferrals_repeated Warning/Critical Warning at >=3 capacity deferrals in 1 hour; critical at >=10 Stop idle/stale slots, reduce noisy workloads, or raise slot caps only after confirming host capacity.
mcp_runtime_restart_storm Warning/Critical Warning at >=3 restarts in 1 hour; critical on any restart suppression Stop the slot, inspect stderr/audit reason codes, and keep the connection disabled until the template/upstream is fixed.
mcp_runtime_connection_health_degraded Warning/Critical Any active enabled connection with degraded/failed/missing-secret health, or any disabled enabled-path connection Run health check, refresh catalog after recovery, or keep the connection disabled and route agents to alternatives.
mcp_runtime_missing_secret_failures Warning/Critical Warning on any missing-secret failure; critical at >=3 in 1 hour Check secret bindings and provider health without revealing secret values; rotate or rebind missing secrets.
mcp_runtime_audit_write_failures Critical Any audit write failure Treat as a control-plane incident; restore DB/audit durability before retrying tool workloads.

Diagnose A Stuck Slot

  1. Read the health summary and note firing alert names:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
    
  2. List durable runtime slots:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots" | jq .
    
  3. Inspect recent gateway audit events without printing secrets:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/tool-gateway/audit?companyId=$COMPANY_ID&limit=100" \
      | jq '[.[] | {createdAt, action, entityType, entityId, reasonCode: .details.reasonCode, tool: .details.tool, durationMs: .details.durationMs}]'
    
  4. Identify the affected slotId, connectionId, reasonCode, and whether the slot is starting, running, idle, failed, or stopped.

Clear A Stuck Slot

Stop the slot first when it is stale, idle, failed, or confirmed not to be serving a healthy active call:

curl -fsS -X POST \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/stop" \
  -d '{}' | jq .

Restart once when the template/config is expected to recover:

curl -fsS -X POST \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/restart" \
  -d '{}' | jq .

If restart suppression fires, do not keep retrying. Disable the connection:

curl -fsS -X PATCH \
  -H "Authorization: Bearer $BOARD_API_KEY" \
  -H "Content-Type: application/json" \
  "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID" \
  -d '{"enabled":false,"status":"disabled"}' | jq '{id, name, enabled, status, healthStatus}'

Verify Recovery

  1. Run the connection health check:

    curl -fsS -X POST \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      -H "Content-Type: application/json" \
      "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/health-check" \
      -d '{}' | jq '{connection: {id: .connection.id, healthStatus: .connection.healthStatus, healthMessage: .connection.healthMessage}, runtimeSlot}'
    
  2. Refresh the catalog after a remote endpoint or stdio template recovers:

    curl -fsS -X POST \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      -H "Content-Type: application/json" \
      "$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/catalog/refresh" \
      -d '{}' | jq '{discoveredCount, quarantinedCount}'
    
  3. Re-read runtime health:

    curl -fsS \
      -H "Authorization: Bearer $BOARD_API_KEY" \
      "$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
    

Recovery is complete when stuck-slot alerts clear, timeout/error rates return below threshold, the connection is healthy or intentionally disabled, and audit events show no new restart suppression or capacity deferrals.

Verification Coverage

Automated coverage includes:

  • A synthetic degraded runtime-health scenario in server/src/__tests__/tool-access-service.test.ts that creates a stale running slot, degraded connection, timeout event, capacity deferral, and restart suppression.
  • A durable audit-write failure scenario in server/src/__tests__/tool-access-service.test.ts that verifies mcp_runtime_audit_write_failures fires from the counter path.
  • A gateway runtime recovery scenario in server/src/__tests__/tool-gateway.test.ts that recovers a stuck local stdio slot before reuse.