mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
150 lines
8.1 KiB
Markdown
150 lines
8.1 KiB
Markdown
# MCP Runtime Operations
|
|
|
|
This runbook covers Paperclip Tools & Access runtime slots for MCP connections. It is written for board and CloudOps operators responding to stuck local stdio slots, degraded remote HTTP connections, capacity deferrals, restart storms, and secret-resolution failures.
|
|
|
|
Do not print raw bearer tokens, gateway session tokens, credential headers, environment variables, or secret values while following this runbook. The APIs below return redacted state and audit metadata; keep shell tracing disabled when exporting credentials.
|
|
|
|
Tool action approvals require `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` to be set independently from auth/JWT secrets. Rotate it deliberately: changing it invalidates outstanding signed tool-action approvals, so drain or reject pending approvals before rotation.
|
|
|
|
## Support Matrix
|
|
|
|
| Transport | Local trusted | Hosted cloud / public authenticated | Notes |
|
|
| --- | --- | --- | --- |
|
|
| `remote_http` | Supported | Supported | Preferred production path. Paperclip proxies calls through the gateway with policy, audit, timeout, and redaction controls. |
|
|
| `local_stdio` | Supported through approved templates and supervised runtime slots | Supported only when an explicitly trusted MCP runtime worker/host is configured | Set `PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST` or `PAPERCLIP_TOOL_RUNTIME_TRUSTED_HOST` only for a worker that is allowed to supervise local processes. Do not enable arbitrary agent-supplied commands. |
|
|
|
|
## Metrics
|
|
|
|
The board runtime health API summarizes one-hour event windows plus current durable slot state:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq .
|
|
```
|
|
|
|
Metrics surfaced there include:
|
|
|
|
- Current slot counts: active, starting, running, idle, failed, stopped.
|
|
- Stuck slot counts: starting/running slots without progress for 5 minutes.
|
|
- Runtime events: capacity deferrals, restart attempts, restart suppression, idle evictions.
|
|
- Tool-call health: call count, timeout count/rate, failure count/rate, average latency, p95 latency.
|
|
- Connection health: active, disabled, degraded, `remote_http`, and `local_stdio` connection counts.
|
|
- Secret failures: missing-secret failures in the last hour.
|
|
- Audit write failures: durable `audit_write_failed` counter increments whenever MCP audit-event persistence fails.
|
|
|
|
## Alerts
|
|
|
|
| Alert | Severity | Suggested threshold | First responder action |
|
|
| --- | --- | --- | --- |
|
|
| `mcp_runtime_stuck_starting_slot` | Critical | Any starting slot older than 5 minutes | Inspect slot health/logs, stop the slot, restart it once, then disable the connection if it sticks again. |
|
|
| `mcp_runtime_stuck_running_slot` | Critical | Any running slot with no progress for 5 minutes | Inspect recent audit events and active calls; restart only after confirming no healthy call is still in progress. |
|
|
| `mcp_runtime_high_timeout_rate` | Warning/Critical | Warning at >=3 timeouts and >=10% in 1 hour; critical at >=10 timeouts or >=25% | Check upstream MCP health, runtime capacity, and gateway audit failures before retrying workloads. |
|
|
| `mcp_runtime_high_error_rate` | Warning/Critical | Warning at >=5 failures and >=10% in 1 hour; critical at >=10 failures or >=25% | Group audit failures by `reasonCode`, then fix credentials/config or disable the affected connection. |
|
|
| `mcp_runtime_capacity_deferrals_repeated` | Warning/Critical | Warning at >=3 capacity deferrals in 1 hour; critical at >=10 | Stop idle/stale slots, reduce noisy workloads, or raise slot caps only after confirming host capacity. |
|
|
| `mcp_runtime_restart_storm` | Warning/Critical | Warning at >=3 restarts in 1 hour; critical on any restart suppression | Stop the slot, inspect stderr/audit reason codes, and keep the connection disabled until the template/upstream is fixed. |
|
|
| `mcp_runtime_connection_health_degraded` | Warning/Critical | Any active enabled connection with degraded/failed/missing-secret health, or any disabled enabled-path connection | Run health check, refresh catalog after recovery, or keep the connection disabled and route agents to alternatives. |
|
|
| `mcp_runtime_missing_secret_failures` | Warning/Critical | Warning on any missing-secret failure; critical at >=3 in 1 hour | Check secret bindings and provider health without revealing secret values; rotate or rebind missing secrets. |
|
|
| `mcp_runtime_audit_write_failures` | Critical | Any audit write failure | Treat as a control-plane incident; restore DB/audit durability before retrying tool workloads. |
|
|
|
|
## Diagnose A Stuck Slot
|
|
|
|
1. Read the health summary and note firing alert names:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
|
|
```
|
|
|
|
2. List durable runtime slots:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots" | jq .
|
|
```
|
|
|
|
3. Inspect recent gateway audit events without printing secrets:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/tool-gateway/audit?companyId=$COMPANY_ID&limit=100" \
|
|
| jq '[.[] | {createdAt, action, entityType, entityId, reasonCode: .details.reasonCode, tool: .details.tool, durationMs: .details.durationMs}]'
|
|
```
|
|
|
|
4. Identify the affected `slotId`, `connectionId`, `reasonCode`, and whether the slot is `starting`, `running`, `idle`, `failed`, or `stopped`.
|
|
|
|
## Clear A Stuck Slot
|
|
|
|
Stop the slot first when it is stale, idle, failed, or confirmed not to be serving a healthy active call:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/stop" \
|
|
-d '{}' | jq .
|
|
```
|
|
|
|
Restart once when the template/config is expected to recover:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-slots/$SLOT_ID/restart" \
|
|
-d '{}' | jq .
|
|
```
|
|
|
|
If restart suppression fires, do not keep retrying. Disable the connection:
|
|
|
|
```sh
|
|
curl -fsS -X PATCH \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID" \
|
|
-d '{"enabled":false,"status":"disabled"}' | jq '{id, name, enabled, status, healthStatus}'
|
|
```
|
|
|
|
## Verify Recovery
|
|
|
|
1. Run the connection health check:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/health-check" \
|
|
-d '{}' | jq '{connection: {id: .connection.id, healthStatus: .connection.healthStatus, healthMessage: .connection.healthMessage}, runtimeSlot}'
|
|
```
|
|
|
|
2. Refresh the catalog after a remote endpoint or stdio template recovers:
|
|
|
|
```sh
|
|
curl -fsS -X POST \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
-H "Content-Type: application/json" \
|
|
"$PAPERCLIP_URL/api/tool-connections/$CONNECTION_ID/catalog/refresh" \
|
|
-d '{}' | jq '{discoveredCount, quarantinedCount}'
|
|
```
|
|
|
|
3. Re-read runtime health:
|
|
|
|
```sh
|
|
curl -fsS \
|
|
-H "Authorization: Bearer $BOARD_API_KEY" \
|
|
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, metrics, alerts}'
|
|
```
|
|
|
|
Recovery is complete when stuck-slot alerts clear, timeout/error rates return below threshold, the connection is healthy or intentionally disabled, and audit events show no new restart suppression or capacity deferrals.
|
|
|
|
## Verification Coverage
|
|
|
|
Automated coverage includes:
|
|
|
|
- A synthetic degraded runtime-health scenario in `server/src/__tests__/tool-access-service.test.ts` that creates a stale running slot, degraded connection, timeout event, capacity deferral, and restart suppression.
|
|
- A durable audit-write failure scenario in `server/src/__tests__/tool-access-service.test.ts` that verifies `mcp_runtime_audit_write_failures` fires from the counter path.
|
|
- A gateway runtime recovery scenario in `server/src/__tests__/tool-gateway.test.ts` that recovers a stuck local stdio slot before reuse.
|