## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
8.4 KiB
Release Notes — MCP Access Governance (v1)
Draft. Source issue: PAP-10397. Parent: PAP-10341.
Status: ready for CTO review. Final voice/copy pass pending sign-off.
TL;DR
Paperclip now governs every MCP tool call an agent makes. Operators install managed connections, define profiles and policies, approve high-risk actions, and read an append-only audit log. Default-deny on unknown tools; quarantine on schema drift; approval required by default for writes; trust rules to lift human-in-the-loop on safe repeats.
Why this exists
Agents that can call arbitrary MCP tools can also leak data, modify accounts, or run destructive operations the operator never approved. Previous Paperclip releases relied on the agent's runtime trusting whatever MCP server it was pointed at. That is the wrong default for production. MCP Access Governance moves the trust boundary to Paperclip itself: every call is selected, evaluated, audited, and (when needed) gated on a human.
What's new for operators
- Tools & Access UI — A new settings surface (
/<prefix>/companies/<companyId>/tools) covering Applications, Connections, Profiles, Policies, Runtime slots, Audit, and bundled Examples. Built for board users and CloudOps. - Managed connections — Two transports:
remote_http(preferred, hosted SaaS MCP) andlocal_stdio(approved-template-only, gated by deployment trust). Operators never paste raw stdio commands; templates ship in the build. - Catalog with risk classification — Every discovered tool gets a
read/write/destructiverisk level inferred from MCP annotations. Destructive tools and unexpected new write tools are auto-quarantined. - Profiles + bindings — Named bundles of include/exclude entries over the catalog, bound to a
company,agent,project,routine, orissue. Narrowest binding wins. - Policies —
allow,block,require_approval,rate_limit, andtrust_rule. Deny beats allow. Policies stack with profiles and run in priority order. - Approval flow — High-risk calls open an action request with signed arguments and an expiry. Human approver decides; agent resumes on approval, fails cleanly on rejection or expiry.
- Trust rules — Promote an approval into a scoped allow rule tied to the canonical argument hash and the catalog schema hash. When schemas drift, trust rules stop applying and the gateway falls back to approval. Revocations are first-class and audited.
- Audit ledger — Append-only call event log with decision, matched policy IDs, reason code, redaction plan, latency, and outcome. Per-run timelines available via
…/runs/:runId/decisions. - Runtime supervisor — Stdio runtime slots have a real lifecycle (
starting,running,idle,failed), idle eviction, restart suppression on storms, and a board health endpoint with alert recommendations. - Bundled examples —
safe-read-only-todo-kvinstalls an application, a connection against a synthetic local fixture, and a read-only profile in one call. A bundled smoke check exercises read / denied write / audit visibility, so operators can confirm the gateway is healthy without an upstream MCP dependency.
What's new for agents
- Agents speak MCP only to the Paperclip gateway. The gateway returns the agent's effective tool list, validates each call against profile + policies, and records the result.
- When a call resolves to
require_approval, the agent's call blocks until a human decides. The agent does not see why a call was denied — that detail is in the audit log for the operator. This is deliberate: agents must not learn to route around denials. - Tool call arguments and results are subject to a redaction plan recorded on each call event.
Default posture
- Unknown tool: deny.
- Catalog drift (new write or destructive tool seen on a refresh): quarantine.
- Write tool, no policy match, no trust rule: requires approval if the profile's default-action allows writes; otherwise denied.
- Destructive tool: denied until an operator explicitly un-quarantines it.
- Local stdio in
authenticated/public: fails closed unlessPAPERCLIP_TRUSTED_MCP_RUNTIME_HOSTis set on a designated trusted worker. - Agent-supplied stdio commands: rejected. Always.
Upgrade and migration
This release introduces new tables (tool_applications, tool_connections, tool_catalog_entries, tool_profiles, tool_profile_entries, tool_profile_bindings, tool_policies, tool_runtime_slots, tool_gateway_sessions, tool_invocations, tool_action_requests, tool_call_events, tool_rate_limit_counters, tool_access_audit_events) and supporting migrations through 0098_tool_gateway_sessions.sql. No data migration is required for existing companies — the tool access stack is opt-in and inert until an operator installs the first connection or example.
Upgrade steps for existing deployments:
- Apply DB migrations as usual (
pnpm paperclipai migrate). - Confirm the Tools & Access tab appears in the UI for board users.
- From Examples, install
safe-read-only-todo-kvand run the bundled smoke. Expectok: trueacross all three checks (allow_read_tool,deny_write_tool,audit_written). - For each existing agent runtime that previously called MCP servers directly: replace direct MCP wiring with a managed connection. Until you do, those agents have no governed tool access on this release.
- If you run
authenticated/public, decide whether you want a trusted runtime worker for local stdio. If yes, setPAPERCLIP_TRUSTED_MCP_RUNTIME_HOSTon that worker and only that worker. If no, leave it unset —remote_httpconnections continue to work.
There is no downgrade path that preserves audit history. If you must roll back, archive any installed connections first so future audits do not surface orphan IDs.
Known limitations
(See MCP-ACCESS-GOVERNANCE.md#known-limitations for the canonical list. The deltas worth calling out in the release note:)
- Audit-write failures use a durable runtime metric counter. Treat a firing
mcp_runtime_audit_write_failuresalert as a control-plane incident until audit durability is restored. - No CLI surface for tool access in v1. Use the UI or REST API.
- No bulk catalog review — each quarantined entry is reviewed one at a time.
- Trust rules match exact argument shapes only. Pattern-based trust rules are post-v1.
- Rate limits are per-policy, not cross-policy aggregates.
- Action request expiry is fixed by policy; approvers cannot extend from the UI.
- Endpoint mode (Paperclip's own
/mcpsurface) is not subject to the profile/policy stack. - Local stdio runtime slots do not migrate across workers; capacity is per-worker, not per-cluster.
Verification commands
After upgrading, run these to confirm the stack is healthy:
# Sanity-check runtime health
curl -fsS -H "Authorization: Bearer $BOARD_API_KEY" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/runtime-health" | jq '{status, alerts}'
# Install the bundled example + run the smoke
curl -fsS -X POST -H "Authorization: Bearer $BOARD_API_KEY" -H "Content-Type: application/json" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/examples/safe-read-only-todo-kv/install" \
-d '{}' | jq .
curl -fsS -X POST -H "Authorization: Bearer $BOARD_API_KEY" -H "Content-Type: application/json" \
"$PAPERCLIP_URL/api/companies/$COMPANY_ID/tools/examples/safe-read-only-todo-kv/smoke" \
-d '{}' | jq '{ok, checks: [.checks[] | {name, ok, decision, reasonCode}]}'
Expected: runtime-health.status is "ok" (no firing alerts on a clean install); smoke.ok is true with three green checks (allow_read_tool, deny_write_tool, audit_written).
Documentation
- Operator guide: doc/MCP-ACCESS-GOVERNANCE.md
- Runtime runbook: doc/MCP-RUNTIME-OPERATIONS.md
- Demo script: doc/MCP-DEMO-SCRIPT.md
- Deployment modes (auth/exposure/bind): doc/DEPLOYMENT-MODES.md
Acknowledgements
Built across Phase 2–8 of the MCP Access Governance program. Implementation phases: PAP-10385, PAP-10386, PAP-10387, PAP-10388, PAP-10389, PAP-10390. Phase 9 readiness rollup: PAP-10392.