Files
PaperClipAI/doc/DEPLOYMENT-MODES.md
T
DottaandPaperclip d9d2147171 fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud owns human identity and passes a verified identity to each
tenant.
> - The tenant can report no session while the Cloud session is still
valid.
> - The access gate and direct `/auth` route then show the instance
password form.
> - This pull request sends those users through the configured Cloud
entry endpoint.
> - Cloud can renew the tenant session or show its login page, then
return to the original task.

## Linked Issues or Issue Description

**What happened?**

A Cloud tenant can display the self-hosted email/password form after an
instance session check returns no session. This gives Cloud users the
wrong login method.

**Expected behavior**

An active Cloud session renews tenant access automatically. A signed-out
user signs in through Cloud. Staging and production use their own
configured Cloud origins. Self-hosted instances keep their instance
login form.

**Steps to reproduce**

1. Open a Cloud tenant task or an `/auth?next=...` link.
2. Keep the Cloud session active but make the instance session check
return 401.
3. Observe the instance password form instead of Cloud session recovery.

**Deployment mode**

Cloud-managed authenticated instances. No database or server API
changes.

Searched related authentication PRs. Native self-hosted OIDC support in
#10411 is a separate feature; this change uses the existing Cloud entry
contract.

## What Changed

- Wait for deployment metadata before showing an instance login form.
- Use the health response's Cloud origin and stack slug for session
recovery.
- Preserve the tenant path, query, and fragment. Reject external and
recursive login return targets.
- Limit automatic recovery per tab. Show a manual Cloud retry after
failed recovery. Show service failures as errors.
- Keep self-hosted login and local trusted access. Add focused tests,
browser regressions, deployment documentation, and an unavailable-state
design example.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- UI typecheck and `pnpm check:token-gates` passed after the final UI
edits.
- All UI tests passed: 639 files, 6,777 tests.
- 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud
links, and recovery coordination.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant
UI and database with a simulated Cloud HTTP endpoint. They cover both
Cloud origins, direct auth/task links, no password-form flash, preserved
URLs, reload, and self-hosted login.
- Hands-on browser test used the real Cloud gateway and a fresh tenant
build with disposable local data. Active Cloud session plus a forced
missing instance session returned to the task. An expired tenant cookie
also renewed automatically and returned to the task. Removing both
sessions reached the real Cloud email/social login UI. Persistent
failure stopped at the retry screen; retry succeeded after removing the
injected fault. The fixture used a loopback transport adapter and a
simulated signed-out OIDC issuer. No production session or deployment
was changed.
- The default local browser startup hit the host's embedded PostgreSQL
resource limit. The passing run used a separate disposable database on
the test PostgreSQL process.
- The full local `pnpm test:run` sweep was stopped after about 31
minutes once CI completed the full suite. It had reported 33 failures in
the unchanged runner API unit/integration files; both files pass in
isolation (1,749 + 28 tests). The local sweep did not reach the later
workspace/serialized groups. CI completed all of those groups
successfully.
- Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at
5/5 with its only thread resolved. All CI gates are green on this
commit, including all general/serialized server groups, workspace tests,
Runner checks, typecheck, build, and all eight browser shards
([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)).

## Risks

- Recovery depends on valid Cloud origin and stack metadata. Incomplete
metadata shows an unavailable message instead of a password form.
- Browsers with session storage disabled use the manual Cloud link,
since automatic retries cannot be bounded across documents.
- The external identity provider's email/social login was not completed
in this local test. Existing Cloud authentication owns that flow.
- No migration, credential format, membership rule, or production
deployment changes.

## Model Used

OpenAI GPT-6 through Codex. The exact served variant and context-window
limit are not exposed in this session. Used reasoning, repository tools,
code execution, and browser testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:45:27 -05:00

225 lines
9.2 KiB
Markdown

# Deployment Modes
Status: Canonical deployment and auth mode model
Date: 2026-02-23
### Paperclip Cloud sign-in
Cloud-managed instances use Cloud for human sign-in. The instance `/auth`
route waits for deployment metadata before rendering; it never renders the
email/password form when health identifies a Cloud-managed instance. A missing
instance session returns through Cloud's `/v1/stacks/:slug/entry-redirect`.
Cloud renews the tenant session from the existing Cloud session, or sends the
user through its sign-in flow. The original tenant path, query, and fragment
travel as `returnTo` so the user returns to the same task.
Both the Cloud origin and stack slug come from the server's health metadata
(`PAPERCLIP_CLOUD_API_ORIGIN` and `PAPERCLIP_STACK_SLUG`). Do not infer the
environment from the browser hostname or hardcode staging/production domains.
Missing configuration shows an unavailable state. An automatic recovery attempt
is limited per browser tab until a session is verified, with a five-minute
expiry and an explicit retry link if recovery fails. Network/server failures
show an error rather than treating the user as signed out.
Self-hosted authenticated instances retain their instance sign-in form, and
`local_trusted` instances retain their normal access path.
## 1. Purpose
Paperclip supports two runtime modes:
1. `local_trusted`
2. `authenticated`
`authenticated` supports two exposure policies:
1. `private`
2. `public`
This keeps one authenticated auth stack while still separating low-friction private-network defaults from internet-facing hardening requirements.
Paperclip now treats **bind** as a separate concern from auth:
- auth model: `local_trusted` vs `authenticated`, plus `private/public`
- reachability model: `server.bind = loopback | lan | tailnet | custom`
## 2. Canonical Model
| Runtime Mode | Exposure | Human auth | Primary use |
|---|---|---|---|
| `local_trusted` | n/a | No login required | Single-operator local machine workflow |
| `authenticated` | `private` | Login required | Private-network access (for example Tailscale/VPN/LAN) |
| `authenticated` | `public` | Login required | Internet-facing/cloud deployment |
## Reachability Model
| Bind | Meaning | Typical use |
|---|---|---|
| `loopback` | Listen on localhost only | default local usage, reverse-proxy deployments |
| `lan` | Listen on all interfaces (`0.0.0.0`) | LAN/VPN/private-network access |
| `tailnet` | Listen on a detected Tailscale IP | Tailscale-only access |
| `custom` | Listen on an explicit host/IP | advanced interface-specific setups |
## 3. Security Policy
## `local_trusted`
- loopback-only host binding
- no human login flow
- optimized for fastest local startup
## `authenticated + private`
- login required
- low-friction URL handling (`auto` base URL mode)
- private-host trust policy required
- Better Auth request rate limiting is off by default for private mode to keep local/LAN repair loops from locking out the operator; set `PAPERCLIP_AUTH_RATE_LIMIT_ENABLED=true` to opt in
- bind can be `loopback`, `lan`, `tailnet`, or `custom`
## `authenticated + public`
- login required
- explicit public URL required
- stricter deployment checks and failures in doctor
- Better Auth request rate limiting is on by default; set `PAPERCLIP_AUTH_RATE_LIMIT_ENABLED=false` only when an explicit front-door limiter covers the deployment
- recommended bind is `loopback` behind a reverse proxy; direct `lan/custom` is advanced
- local stdio MCP runtime slots fail closed by default; set `PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST` only when a trusted worker/runtime host is configured to supervise those processes. Remote HTTP MCP remains the preferred public-hosted path.
### Paperclip Cloud warm-pool identity
A Cloud-managed warm-pool process initially boots under a `pool-*` origin. It
receives only Cloud's public verification set in
`PAPERCLIP_CLOUD_RUNTIME_IDENTITY_JWKS`. Before Cloud activates a claimed stack,
the existing server-to-server health request carries a short-lived Ed25519 JWS
that binds the immutable `PAPERCLIP_CLOUD_STACK_ID`, pool claim, previous
origin, canonical HTTPS origin, and slug. Paperclip verifies and persists that
one-time assertion, updates its live public/API URL provider, and acknowledges
the exact origin in `/api/health` before the first user request is admitted.
The Harness signing private key is never present in Paperclip, browsers, or
other tenant stacks. A different claim or destination cannot replace the
persisted identity. On restart, the durable identity is loaded before auth,
routes, and child-runtime configuration, even when provider variables are
temporarily stale. Self-hosted deployments continue to use their configured
`PAPERCLIP_PUBLIC_URL` and do not participate in this protocol.
## 4. Onboarding UX Contract
Default onboarding remains interactive and flagless:
```sh
pnpm paperclipai onboard
```
Server prompt behavior:
1. quickstart `--yes` defaults to `server.bind=loopback` and therefore `local_trusted/private`
2. advanced server setup asks reachability first:
- `Trusted local` → `bind=loopback`, `local_trusted/private`
- `Private network` → `bind=lan`, `authenticated/private`
- `Tailnet` → `bind=tailnet`, `authenticated/private`
- `Custom` → manual mode/exposure/host entry
3. raw host entry is only required for the `Custom` path
4. explicit public URL is only required for `authenticated + public`
Examples:
```sh
pnpm paperclipai onboard --yes
npx paperclipai onboard --yes --bind lan
npx paperclipai run --bind tailnet
```
`configure --section server` follows the same interactive behavior.
## 5. Doctor UX Contract
Default doctor remains flagless:
```sh
pnpm paperclipai doctor
```
Doctor reads configured mode/exposure and applies mode-aware checks. Optional override flags are secondary.
## 6. Board/User Integration Contract
Board identity must be represented by a real DB user principal for user-based features to work consistently.
Required integration points:
- real user row in `authUsers` for Board identity
- `instance_user_roles` entry for Board admin authority
- `company_memberships` integration for user-level task assignment and access
This is required because user assignment paths validate active membership for `assigneeUserId`.
## 7. Local Trusted -> Authenticated Claim Flow
When running `authenticated` mode, if the only instance admin is `local-board`, Paperclip emits a startup warning with a one-time high-entropy claim URL.
- URL format: `/board-claim/<token>?code=<code>`
- intended use: signed-in human claims board ownership
- claim action:
- promotes current signed-in user to `instance_admin`
- demotes `local-board` admin role
- ensures active owner membership for the claiming user across existing companies
This prevents lockout when a user migrates from long-running local trusted usage to authenticated mode.
## 8. First Admin Setup For Fresh Authenticated Installs
Fresh authenticated installs start in `bootstrap_pending` until the first
`instance_admin` exists.
For `authenticated/private`, Paperclip supports a browser-first setup path:
1. open the Paperclip URL from the private network or appliance UI
2. sign in or create a Paperclip account
3. choose `Claim this instance` on the setup screen
That browser claim promotes the signed-in session user to the first instance
admin and then falls through to normal onboarding. The endpoint is available
only to real browser session actors in `authenticated/private`; unauthenticated
requests, agent keys, board API keys, and local implicit board actors are
rejected.
This is intentionally a first-claim bootstrap contract: before an instance
admin exists, the first authenticated browser session that completes the claim
wins. Operators must keep a `bootstrap_pending` private deployment on a trusted
network and complete setup before admitting untrusted users. This behavior is
not an account-recovery or public-deployment mechanism.
The CLI fallback remains supported in all authenticated setup states:
```sh
pnpm paperclipai auth bootstrap-ceo
```
That command prints a one-time first-admin invite URL. Browser claim and
bootstrap invite acceptance share the same first-admin transaction, so whichever
path wins first makes later attempts return a conflict.
For `authenticated/public`, browser first-admin claim is intentionally disabled.
Public deployments must use the high-entropy bootstrap invite path unless a
future public-hosted setup design explicitly changes this policy.
## 9. Current Code Reality (As Of 2026-02-23)
- runtime values are `local_trusted | authenticated`
- `authenticated` uses Better Auth sessions and bootstrap invite flow
- `local_trusted` ensures a real local Board user principal in `authUsers` with `instance_user_roles` admin access
- company creation ensures creator membership in `company_memberships` so user assignment/access flows remain consistent
## 10. Naming and Compatibility Policy
- canonical naming is `local_trusted` and `authenticated` with `private/public` exposure
- no long-term compatibility alias layer for discarded naming variants
## 11. Relationship to Other Docs
- implementation plan: `doc/plans/2026-02-23-deployment-auth-mode-consolidation.md`
- V1 contract: `doc/SPEC-implementation.md`
- operator workflows: `doc/DEVELOPING.md` and `doc/CLI.md`
- invite/join state map: `doc/spec/invite-flow.md`