mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip lets people manage agents and govern their tool access. > - Native runs receive an immutable MCP tool assignment for one agent. > - The gateway must enforce that owner when it authenticates a run token. > - Older gateway rows stored the owner only in metadata. > - This change validates the gateway and profile, binds new rows, and repairs valid older rows on reuse. > - It also delivers each native assignment once. > - Other agents cannot use the assignment, and explicit shared gateways keep their configured scope. ## Linked Issues or Issue Description Builds on #14012 by @busla (Jón Levy). That PR adds agent binding and seven regressions. This PR carries that fix onto current master and adds legacy authentication, profile validation, and duplicate-delivery coverage. Related: #14864 improves discovery memory use. **What happened?** Native gateway creation stored an owner in metadata but left `agentId` null. Authentication could therefore accept another agent's run token. Managed discovery could also deliver historical native assignments again. **Expected behavior** A native assignment accepts only its owner's run token. The gateway and profile must refer to the same immutable assignment. The current assignment enters the run configuration once. **Steps to reproduce** 1. Run the database fixtures in `heartbeat-runtime-mcp-servers.test.ts` on the baseline. 2. Create a native assignment and inspect its stored gateway owner. 3. Authenticate with another agent's run token, then inspect legacy reuse and managed delivery. 4. The baseline fails six ownership and delivery cases. The fix passes all twelve cases. **Paperclip version or commit** The red baseline is `f2e0f1963`. This PR is based on `cad26c6bf`, which includes the merged discovery fix. **Deployment mode** Native Paperclip Runner execution and managed Codex MCP delivery. Reproduction uses isolated database and HTTP fixtures. ## What Changed - Store the agent owner and agent context on new native gateways. - Validate profile and gateway assignment metadata before reuse or token creation. - Bind valid legacy rows with a company-scoped, null-owner update and validate the result. - Reject mismatched run tokens before legacy repair. - Identify native assignments by gateway metadata, the reserved profile key, or profile source. Reject missing or malformed provenance, including JSON null. - Exclude historical native assignments from managed gateway delivery. Keep their rows for existing runs. - Add twelve database and HTTP regressions and document the runtime contract. ## Verification - Red baseline: six regressions fail and four controls pass before the initial fix. Two additional regressions reproduce metadata-loss admission and a JSON-null TypeError before the review fix. - All twelve ownership regressions pass on the final code, including owner admission, cross-agent rejection, metadata loss, JSON-null HTTP 401, and explicit shared-gateway admission. Policy, listing-memory, and discovery HTTP coverage also passes. - Full workspace typecheck and build pass locally. Server typecheck and compilation pass again after the review fix. The final ownership and grant patches pass 42 combined database and HTTP regressions. - [Full CI](https://github.com/paperclipai/paperclip/actions/runs/37011383657) passes for `626a08ae66361cf586105877e24d806b1a7a9c20`: all 54 checks succeed; two optional Storybook checks skip. This includes full typecheck, build, all test shards, all eight E2E shards, runner verification, and the canary dry run. - Greptile scores that exact head 5/5. No review threads remain unresolved. ## Risks Invalid historical native gateway or profile metadata now rejects authentication. Valid unbound rows are repaired only when their owner reuses the assignment. Conflicting owners are never overwritten. Historical rows are retained for existing runs. Explicit shared gateways use ordinary profiles and keep their configured scopes. The reserved native profile namespace remains agent-owned even when gateway metadata is cleared. No schema or dependency changes are included. ## Model Used Original fix and seven regressions in #14012: Anthropic Claude Opus 5.5, `claude-opus-5-5`, 1M context, as reported by @busla. Extensions and verification: OpenAI Codex (GPT-6), with reasoning, repository inspection, code execution, and tests. This session does not expose the exact serving model identifier or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
173 lines
9.5 KiB
Markdown
173 lines
9.5 KiB
Markdown
# MCP Runtime Operations
|
|
|
|
This runbook covers Paperclip Tools & Access runtime slots for MCP connections. It is written for board and CloudOps operators responding to stuck local stdio slots, degraded remote HTTP connections, capacity deferrals, restart storms, and secret-resolution failures.
|
|
|
|
Do not print raw bearer tokens, gateway session tokens, credential headers, environment variables, or secret values while following this runbook. The APIs below return redacted state and audit metadata; keep shell tracing disabled when exporting credentials.
|
|
|
|
Tool action approvals require `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` to be set independently from auth/JWT secrets. `paperclipai onboard` generates it for local instances, and worktree setup propagates or generates an independent value in the worktree `.env`; operator-managed deployments must set it explicitly. Rotate it deliberately: changing it invalidates outstanding signed tool-action approvals, so drain or reject pending approvals before rotation.
|
|
|
|
## Support Matrix
|
|
|
|
| Transport | Local trusted | Hosted cloud / public authenticated | Notes |
|
|
| --- | --- | --- | --- |
|
|
| `remote_http` | Supported | Supported | Preferred production path. Paperclip proxies calls through the gateway with policy, audit, timeout, and redaction controls. |
|
|
| `local_stdio` | Supported through approved templates and supervised runtime slots | Supported only when an explicitly trusted MCP runtime worker/host is configured | Set `PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST` or `PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST` only for a worker that is allowed to supervise local processes. Do not enable arbitrary agent-supplied commands. |
|
|
|
|
## Native runtime gateway ownership
|
|
|
|
Paperclip Runner creates an immutable MCP assignment for one agent. Its gateway
|
|
stores the owner in `agentId` and the agent context scope. Authentication checks
|
|
that the run belongs to that owner and that the gateway and profile metadata
|
|
refer to the same assignment. Invalid ownership or assignment metadata rejects
|
|
the token with `gateway_token_run_context_invalid`.
|
|
|
|
The reserved `native:` profile key also identifies a native assignment if its
|
|
gateway metadata is cleared. Clearing metadata cannot turn the assignment into
|
|
a shared gateway. Invalid JSON metadata, including JSON `null`, returns the same
|
|
authentication rejection. Use an ordinary profile for an explicit shared gateway.
|
|
|
|
Older native gateways can have a null `agentId`. Authentication still checks
|
|
their metadata owner. When the owner reuses the assignment, Paperclip validates
|
|
the gateway and profile before it binds the gateway to that agent. A conflicting
|
|
owner is never overwritten.
|
|
|
|
Native assignments enter the run configuration once through the runtime MCP
|
|
delivery path. Managed gateway discovery excludes historical native assignments.
|
|
Their stored rows remain available for existing runs. Explicit company gateways
|
|
continue to use their configured scopes.
|
|
|
|
## Metrics
|
|
|
|
The board runtime health API summarizes one-hour event windows plus current durable slot state:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq .
|
|
```
|
|
|
|
Metrics surfaced there include:
|
|
|
|
- Current slot counts: active, starting, running, idle, failed, stopped.
|
|
- Stuck slot counts: starting/running slots without progress for 5 minutes.
|
|
- Runtime events: capacity deferrals, restart attempts, restart suppression, idle evictions.
|
|
- Tool-call health: call count, timeout count/rate, failure count/rate, average latency, p95 latency.
|
|
- Connection health: active, disabled, degraded, `remote_http`, and `local_stdio` connection counts.
|
|
- Secret failures: missing-secret failures in the last hour.
|
|
- Audit write failures: durable `audit_write_failed` counter increments whenever MCP audit-event persistence fails.
|
|
|
|
## Alerts
|
|
|
|
| Alert | Severity | Suggested threshold | First responder action |
|
|
| --- | --- | --- | --- |
|
|
| `mcp_runtime_stuck_starting_slot` | Critical | Any starting slot older than 5 minutes | Inspect slot health/logs, stop the slot, restart it once, then disable the connection if it sticks again. |
|
|
| `mcp_runtime_stuck_running_slot` | Critical | Any running slot with no progress for 5 minutes | Inspect recent audit events and active calls; restart only after confirming no healthy call is still in progress. |
|
|
| `mcp_runtime_high_timeout_rate` | Warning/Critical | Warning at >=3 timeouts and >=10% in 1 hour; critical at >=10 timeouts or >=25% | Check upstream MCP health, runtime capacity, and gateway audit failures before retrying workloads. |
|
|
| `mcp_runtime_high_error_rate` | Warning/Critical | Warning at >=5 failures and >=10% in 1 hour; critical at >=10 failures or >=25% | Group audit failures by `reasonCode`, then fix credentials/config or disable the affected connection. |
|
|
| `mcp_runtime_capacity_deferrals_repeated` | Warning/Critical | Warning at >=3 capacity deferrals in 1 hour; critical at >=10 | Stop idle/stale slots, reduce noisy workloads, or raise slot caps only after confirming host capacity. |
|
|
| `mcp_runtime_restart_storm` | Warning/Critical | Warning at >=3 restarts in 1 hour; critical on any restart suppression | Stop the slot, inspect stderr/audit reason codes, and keep the connection disabled until the template/upstream is fixed. |
|
|
| `mcp_runtime_connection_health_degraded` | Warning/Critical | Any active enabled connection with degraded/failed/missing-secret health, or any disabled enabled-path connection | Run health check, refresh catalog after recovery, or keep the connection disabled and route agents to alternatives. |
|
|
| `mcp_runtime_missing_secret_failures` | Warning/Critical | Warning on any missing-secret failure; critical at >=3 in 1 hour | Check secret bindings and provider health without revealing secret values; rotate or rebind missing secrets. |
|
|
| `mcp_runtime_audit_write_failures` | Critical | Any audit write failure | Treat as a control-plane incident; restore DB/audit durability before retrying tool workloads. |
|
|
|
|
## Diagnose A Stuck Slot
|
|
|
|
1. Read the health summary and note firing alert names:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
|
|
```
|
|
|
|
2. List durable runtime slots:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots" | jq .
|
|
```
|
|
|
|
3. Inspect recent gateway audit events without printing secrets:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/tool-gateway/audit?companyId=$COMPANY_ID&limit=100" \
|
|
| jq '[.[] | {createdAt, action, entityType, entityId, reasonCode: .details.reasonCode, tool: .details.tool, durationMs: .details.durationMs}]'
|
|
```
|
|
|
|
4. Identify the affected `slotId`, `connectionId`, `reasonCode`, and whether the slot is `starting`, `running`, `idle`, `failed`, or `stopped`.
|
|
|
|
## Clear A Stuck Slot
|
|
|
|
Stop the slot first when it is stale, idle, failed, or confirmed not to be serving a healthy active call:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/stop" \
|
|
-d '{}' | jq .
|
|
```
|
|
|
|
Restart once when the template/config is expected to recover:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/restart" \
|
|
-d '{}' | jq .
|
|
```
|
|
|
|
If restart suppression fires, do not keep retrying. Disable the connection:
|
|
|
|
```sh
|
|
curl -fsS -X PATCH \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID" \
|
|
-d '{"enabled":false,"status":"disabled"}' | jq '{id, name, enabled, status, healthStatus}'
|
|
```
|
|
|
|
## Verify Recovery
|
|
|
|
1. Run the connection health check:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/health-check" \
|
|
-d '{}' | jq '{connection: {id: .connection.id, healthStatus: .connection.healthStatus, healthMessage: .connection.healthMessage}, runtimeSlot}'
|
|
```
|
|
|
|
2. Refresh the catalog after a remote endpoint or stdio template recovers:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/catalog/refresh" \
|
|
-d '{}' | jq '{discoveredCount, quarantinedCount}'
|
|
```
|
|
|
|
3. Re-read runtime health:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
|
|
```
|
|
|
|
Recovery is complete when stuck-slot alerts clear, timeout/error rates return below threshold, the connection is healthy or intentionally disabled, and audit events show no new restart suppression or capacity deferrals.
|
|
|
|
## Verification Coverage
|
|
|
|
Automated coverage includes:
|
|
|
|
- A synthetic degraded runtime-health scenario in `server/src/__tests__/tool-access-service.test.ts` that creates a stale running slot, degraded connection, timeout event, capacity deferral, and restart suppression.
|
|
- A durable audit-write failure scenario in `server/src/__tests__/tool-access-service.test.ts` that verifies `mcp_runtime_audit_write_failures` fires from the counter path.
|
|
- A gateway runtime recovery scenario in `server/src/__tests__/tool-gateway.test.ts` that recovers a stuck local stdio slot before reuse.
|