mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cloud-managed instances authenticate tenant users through a trusted-header path (`resolveCloudTenantActor`) that deliberately never grants `instance_admin`, so every tenant user is company-scoped > - On a dedicated (single-owner) managed instance that leaves the paying owner unable to administer their own instance: instance settings, the environments admin surface, and the custom sandbox image flow are all instance-admin gated (the environments UI can't even show the provider/image of the platform sandbox because the restricted read view blanks `config` entirely) > - Re-granting the old blanket `instance_user_roles` row would repeat the mistake the shared-pool hardening fixed: DB rows go stale, resurrect via restores, and elevate through every auth path > - This pull request elevates only the stack `owner`, computed per request at the trusted-header boundary behind a new managed-tier feature flag, and ships that elevation together with code floors on the platform-owned surfaces an instance admin must not control on a managed instance > - The benefit is that dedicated-stack owners can administer their own instance while platform credentials, execution policy, backups, and runtime code-install stay platform-owned, and self-hosted behavior is unchanged ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described here following the feature-request template. ### Problem or motivation - On cloud-managed instances, tenant users resolved from trusted headers are always company-scoped. For dedicated instances with a single paying owner, the owner cannot reach any instance-admin surface of their own instance (instance settings, environments administration, custom image setup), and the restricted environment read view hides even structural fields like the sandbox provider and image. - The previous hardening intentionally removed blanket elevation (and purges stale `instance_user_roles` rows on every trusted-header authentication). That protection must not regress for shared multi-tenant pools. ### Proposed solution - Owner-only, computed, flag-gated elevation plus code floors on platform-owned surfaces, in one PR so the elevation can never ship without the floors. ### Alternatives considered - Re-inserting an `instance_user_roles` row for owners (the pre-hardening model): rejected — DB rows go stale, survive restores, and elevate through every auth path; #7525 removed exactly this. - Widening only the environments read view without any elevation: rejected — it fixes one screen but still leaves a dedicated-instance owner unable to administer instance settings or custom images. - Elevating additional stack roles (`member`/`admin`/`support`): rejected — only the owner has an ownership claim over the whole instance; other roles stay company-scoped. ### Roadmap alignment - Extends the shipped "Cloud deployments" roadmap work (multi-tenant isolation, company-scoped cloud tenants, managed-instance bootstrap) without overlapping planned core items, and leaves self-hosted behavior unchanged. ## What Changed - **New feature key** `enableOwnerInstanceAdmin` (`packages/shared`): boolean flag in `instanceExperimentalSettingsSchema`, catalog tier `managed`, `cloudDefault: true`, `selfHostedDefault: false`. Inert on self-hosted instances — the elevation path only exists behind the cloud tenant trust token. - **Computed elevation** (`server/src/middleware/auth.ts`): `resolveCloudTenantActor` now returns `isInstanceAdmin: true` only when the trusted-header stack role is `owner` **and** the flag is enabled. The flag is resolved through the instance-settings service so the managed-config overlay applies (the control plane can disable elevation fleet-wide without touching tenant databases; a DB row edit or restore cannot resurrect it). Resolution fails closed on settings read errors. The `instance_user_roles` never-insert and the per-request stale-row purge are byte-identical. `member`/`admin`/`support` stack roles stay company-scoped. - **Authorization guard split** (`server/src/services/authorization.ts`): the blanket-allow now trusts the actor's *computed* `isInstanceAdmin` flag (only the attested resolver can set it for `cloud_tenant` actors) while keeping the `instance_user_roles` DB lookup excluded for `cloud_tenant` — a stale or hand-inserted role row still elevates nothing. - **Floor F1 — platform environment credentials** (`server/src/routes/environments.ts`): on cloud-managed instances, platform-provisioned environment rows (`managedByPaperclip` marker, plus the legacy managed-Kubernetes marker) use a single floored view for every reader on all environment routes (list, get, create, update, delete responses): `envVars` are never echoed and credential-shaped `config` keys (reusing the managed-config `SECRET_LIKE_CONFIG_KEY_PATTERN`) are dropped — for **all** actors including instance admins — while structural config (provider, image, template, region, …) and the managed markers stay visible. This also fixes the environments UI for managed sandboxes, which previously lost the provider/image entirely in the restricted view. The floor also covers writes: `PATCH /environments/:id` and `DELETE /environments/:id` on a platform-provisioned row are rejected (403, `environment_platform_managed`) for every actor including instance admins, and the guard binds to the persisted row's markers so a patch cannot strip the managed marker to lift the floor. The one recovery path is a metadata-only PATCH that solely clears the marker keys (null/false), for rows stamped through the old unrestricted API before the markers became reserved — and it never applies to a row whose slot markers are live platform state: the single local row (`environments_local_driver_idx`), which `ensureLocalEnvironment` adopts and stamps on cloud-managed instances from every caller (company creation, the heartbeat, run orchestration), and the single marked sandbox row (`environments_managed_sandbox_idx`) while a managed-sandbox bootstrap path is configured (managed-config `environments` section or `PAPERCLIP_EXECUTION_MODE=kubernetes`) and the provisioner therefore adopts and refreshes it on every boot. Clearing a live slot row's markers would let the next write reclassify it as tenant-managed and bypass the floor; conversely, when no sandbox provisioning path is configured the platform holds no claim on any sandbox row, so a platform marker there is stale by definition and the recovery patch applies. Every marker outside a live slot is clearable, so no legacy row is ever locked permanently. Custom-image setup and probes on the platform sandbox stay available to instance admins — those are the owner-facing flows this elevation exists for. The marker keys themselves are reserved: client create/update payloads that set `managedByPaperclip` or `managedKubernetesSandbox` are rejected (422, `environment_platform_marker_reserved`) on cloud-managed instances, so a tenant row can never be stamped platform-provisioned through the API and self-locked behind the write floor (the provisioner writes markers at the service layer, not through these routes). Tenant-created environments are otherwise unaffected. - **Floor F2 — executionMode** (`server/src/routes/instance-settings.ts`): on cloud-managed instances, `PATCH /instance/settings/general` rejects writes that would change `executionMode` (403, `execution_mode_platform_managed`). Same-value echoes pass so settings forms that submit the full general-settings object keep working. The boot-time execution-policy bootstrap path is untouched (it calls the service directly). - **Floor F3 — manual database backups** (`server/src/routes/instance-database-backups.ts`): floored off on cloud-managed instances (403, `database_backups_platform_managed`); backups are platform-owned there, and the result would also echo a server-side filesystem path. - **Floor F4 — adapter code install** (`server/src/routes/adapters.ts`): `POST /adapters/install` and `POST /adapters/:type/reinstall` are floored off on cloud-managed instances (403, `adapter_install_platform_managed`). Adapter packages execute in the server process, so a runtime install would let an instance admin read the platform trust anchors out of the process environment. This mirrors the existing bundled-only plugin install floor; adapter code on managed instances comes bundled with the platform image. ## Instance-admin surface audit Before widening who can hold `isInstanceAdmin`, every instance-admin-gated surface in `server/src` was enumerated and reviewed for whether its response or side effects could echo process environment values or platform credentials (tenant trust token, JWT signing keys, database connection strings, provider API keys): 29 distinct gate definitions covering ~90+ call sites, in four groups — sole instance-admin gates (12), instance-admin-or-company-permission gates (10), response-shaping/scope-widening sites (6), and the central `allow_instance_admin` short-circuit in the authorization service (58 `decide()` call sites). Findings and dispositions: - **Environment read/write responses** exposed platform sandbox `envVars`/credential-shaped config to instance admins → closed by floor F1. - **Manual backup trigger** echoed a server filesystem path and triggers a platform-owned operation → closed by floor F3. - **Adapter install/reinstall** loads externally fetched code into the server process (indirect, complete env exposure) → closed by floor F4. The sibling plugin-install path already had a bundled-only floor on managed instances and needed no change. - **Token-minting surfaces** (gateway tokens, custom-image terminal/connection tokens) mint credentials scoped to the instance's own resources, not platform trust anchors → acceptable for an owner-admin of a dedicated instance; unchanged. - All remaining gated surfaces return ordinary instance-scoped business data; none echo `process.env` or platform secrets directly. OAuth client secrets are referenced by env-var *name* only; SSH private keys are stored as secret refs before persistence and are not echoed. Operational note for managed platforms: this model assumes the process environment of a managed instance holds only that instance's own credentials. Platform operators should keep provider credentials per-instance (never fleet-shared) since an instance admin ultimately controls in-process code on their own instance. ## Verification - `pnpm vitest run server/src/middleware/cloud-tenant-actor.test.ts` — resolver matrix: owner × flag on/off, flag via managed overlay (on-over-DB-off and off-over-DB-on), member/admin/support × flag on, no-token self-hosted, fail-closed settings read, purge still runs and no role row is ever inserted (14 tests). - `pnpm vitest run server/src/__tests__/authorization-service.test.ts` — computed flag elevates a `cloud_tenant` actor; a stale `instance_user_roles` row still never does; `session` actors unchanged (full suite, embedded Postgres). - `pnpm vitest run server/src/__tests__/environment-routes.test.ts` — F1: no secret echo to admins on get/list, structural config visible to restricted readers, platform-row PATCH/DELETE rejected for admins (including a marker-stripping patch), marker-clear recovery allowed for stale legacy rows and for a marked sandbox row when no provisioning path is configured, but refused on the managed local row and on the sandbox slot row under a managed-config `environments` entry or the forced kubernetes execution mode, client marker-stamping creates/patches rejected, tenant rows still readable and writable, self-hosted read+write regression (60 tests). - `pnpm vitest run server/src/__tests__/environment-service.test.ts` — `ensureLocalEnvironment` adopts a pre-existing local row on cloud-managed instances (marker stamped, other metadata preserved, idempotent — no rewrite on re-ensure) and leaves self-hosted rows untouched (22 tests, embedded Postgres). - `pnpm vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/instance-database-backups-routes.test.ts` — F2 change-vs-echo matrix incl. self-hosted regression; F3 floor for both admin shapes (32 tests). - `pnpm vitest run server/src/__tests__/adapter-routes-authz.test.ts` — F4 floor; self-hosted install/reinstall behavior unchanged (existing cases). - `pnpm vitest run server/src/__tests__/first-admin-claim.test.ts server/src/__tests__/bootstrap-claim-routes.test.ts server/src/__tests__/managed-config.test.ts server/src/__tests__/health.test.ts server/src/__tests__/instance-settings-managed-overlay.test.ts server/src/services/managed-environments.test.ts server/src/services/execution-policy-bootstrap.test.ts` — first-admin bootstrap gate and managed-config behavior unchanged (91 tests). - `pnpm vitest run packages/shared/src/feature-catalog.test.ts` — catalog/schema sync tests cover the new key (selfHostedDefault must equal the schema default). - `pnpm run typecheck` — all 31 workspace projects clean. ## Risks - Self-hosted behavior is unchanged: every floor binds to `isCloudManagedInstance()` (tenant trust token present), the new flag defaults off with no elevation path, and regression tests pin the self-hosted branches. - The elevation is fail-closed and stateless: turning the flag off (managed overlay or DB) de-elevates on the next request; there is no role row to clean up and restores cannot resurrect elevation. - On a cloud-managed instance a pre-existing unmarked local row is adopted (stamped `managedByPaperclip`) by the next ensure and becomes platform-owned — the intended managed-product semantic: the platform owns the single local slot. Self-hosted instances are untouched. - F1 widens restricted readers' view of platform-provisioned rows from fully blanked `config`/`metadata` to structural-only `config` plus markers. Platform-delivered config is guaranteed secret-free by the managed-config contract (secret-shaped keys fail startup), and the floor re-drops secret-shaped keys defensively. - One extra instance-settings read per trusted-header request for owner-role actors (the resolver already performs several queries per request). ## Model Used Claude Fable 5 (Anthropic) — model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code; read-only explore subagents on the same model were used for the surface audit sweep. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>