## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
17 KiB
Docker Quickstart
Run Paperclip in Docker without installing Node or pnpm locally.
All commands below assume you are in the project root (the directory containing package.json), not inside docker/.
Building the image
docker build -t paperclip-local .
The Dockerfile installs common agent tools (git, gh, curl, wget, ripgrep, python3) and the Claude, Codex, and OpenCode CLIs.
Build arguments:
| Arg | Default | Purpose |
|---|---|---|
USER_UID |
1000 |
UID for the container node user (match your host UID to avoid permission issues on bind mounts) |
USER_GID |
1000 |
GID for the container node group |
CLI_TOOLS_CACHE_EPOCH |
empty | Refresh the CLI-install layer; CI supplies the current ISO week |
PAPERCLIP_BUILD_VERSION |
empty | Runtime version when Git metadata is unavailable |
PAPERCLIP_BUILD_COMMIT |
empty | Source commit written into the server build stamp and runtime environment |
Changing the build version or commit preserves the CLI-install cache. The tool layer refreshes when its weekly epoch, base image, installation command, or earlier build inputs change. Local builds can set a new epoch explicitly to refresh tools without clearing the entire build cache.
docker build -t paperclip-local \
--build-arg USER_UID=$(id -u) --build-arg USER_GID=$(id -g) .
Standard images and downstream composition
The Docker workflow publishes the standard production target for Linux AMD64
and ARM64. Canonical master pushes also publish
ghcr.io/paperclipai/paperclip:sha-<FULL_SHA> and a GitHub/Sigstore attestation
for its immutable multi-platform digest. Downstream services can compose their
own images from this public base without rebuilding Core.
Stamped standard images include the build-owned remote provider pack at
/opt/paperclip-runner/provider-pack and set
PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH to that directory. Downstream
compositions and the explicit cloud target inherit both. The pack is required
on the controller for remote OpenCode and ACPX runs, even when the sandbox has
preinstalled providers; setting the environment variable alone cannot repair an
image that omits the pack. Rebuild and redeploy from a standard image containing
the pack if runs fail with runner_remote_provider_artifact_incompatible.
Local Docker builds must supply a full PAPERCLIP_BUILD_COMMIT to include a
qualified pack. Unstamped builds remain usable for local adapters and skip pack
generation. The pack stays root-owned and readable by remapped runtime UIDs.
The legacy recurring public -cloud publisher is retired. Master pushes,
release tags, and manual Docker dispatches no longer build that variant.
Existing -cloud tags and digests remain in the registry for rollback; their
release-channel aliases no longer advance. This change deletes no images,
cache tags, or migrator artifacts.
The cloud Dockerfile target remains available for explicit
preview builds. Those requests still publish a
full-SHA -cloud tag when needed. They do not advance a release channel or
replace downstream private composition.
A published image alone does not prove source tests or migration compatibility. Downstream deployment tooling must verify source proof, the standard image attestation, the exact-source migrator, and its own composed image before rollout. Resolve immutable digests instead of deploying mutable tags.
One-liner (build + run)
docker build -t paperclip-local . && \
docker run --name paperclip \
-p 3100:3100 \
-e HOST=0.0.0.0 \
-e PAPERCLIP_HOME=/paperclip \
-e BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
-e PAPERCLIP_TOOL_ACTION_SIGNING_SECRET=$(openssl rand -hex 32) \
-v "$(pwd)/data/docker-paperclip:/paperclip" \
paperclip-local
Open: http://localhost:3100
Data persistence:
- Embedded PostgreSQL data
- uploaded assets
- local secrets key
- local agent workspace data
All persisted under your bind mount (./data/docker-paperclip in the example above).
Docker Compose
Quickstart (embedded SQLite)
Single container, no external database. Data persists via a bind mount.
BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
PAPERCLIP_TOOL_ACTION_SIGNING_SECRET=$(openssl rand -hex 32) \
docker compose -f docker/docker-compose.quickstart.yml up --build
Defaults:
- host port:
3100 - persistent data dir:
./data/docker-paperclip
Optional overrides:
PAPERCLIP_PORT=3200 PAPERCLIP_DATA_DIR=../data/pc \
docker compose -f docker/docker-compose.quickstart.yml up --build
Note: PAPERCLIP_DATA_DIR is resolved relative to the compose file (docker/), so ../data/pc maps to data/pc in the project root.
If you change host port or use a non-local domain, set PAPERCLIP_PUBLIC_URL to the external URL you will use in browser/auth flows.
Pass OPENAI_API_KEY and/or ANTHROPIC_API_KEY to enable local adapter runs.
Full stack (with PostgreSQL)
Paperclip server + PostgreSQL 17. The database is health-checked before the server starts.
BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
docker compose -f docker/docker-compose.yml up --build
PostgreSQL data persists in a named Docker volume (pgdata). Paperclip data persists in paperclip-data.
Untrusted PR review
Isolated container for reviewing untrusted pull requests with Codex or Claude, without exposing your host machine. See doc/UNTRUSTED-PR-REVIEW.md for the full workflow.
docker compose -f docker/docker-compose.untrusted-review.yml build
docker compose -f docker/docker-compose.untrusted-review.yml run --rm --service-ports review
Authenticated Compose (Single Public URL)
For authenticated deployments, set one canonical public URL and let Paperclip derive auth/callback defaults:
services:
paperclip:
environment:
PAPERCLIP_DEPLOYMENT_MODE: authenticated
PAPERCLIP_DEPLOYMENT_EXPOSURE: private
PAPERCLIP_PUBLIC_URL: https://desk.koker.net
PAPERCLIP_PUBLIC_URL is used as the primary source for:
- auth public base URL
- Better Auth base URL defaults
- bootstrap invite URL defaults
- hostname allowlist defaults (hostname extracted from URL)
For fresh authenticated/private Docker or appliance-style installs, the first
admin can now be claimed entirely from the browser after sign-in. Open the
Paperclip URL, sign in or create an account, then choose Claim this instance
on the setup screen. This browser claim is disabled for authenticated/public;
public deployments should run the high-entropy CLI invite fallback instead:
pnpm paperclipai auth bootstrap-ceo
Granular overrides remain available if needed (PAPERCLIP_AUTH_PUBLIC_BASE_URL, BETTER_AUTH_URL, BETTER_AUTH_TRUSTED_ORIGINS, PAPERCLIP_ALLOWED_HOSTNAMES).
Set PAPERCLIP_ALLOWED_HOSTNAMES explicitly only when you need additional hostnames beyond the public URL host (for example Tailscale/LAN aliases or multiple private hostnames).
Optional Vercel Connect credentials
Vercel Connect's backend integration is retained for controlled testing and
existing Vercel-backed connections, but its new-connection UI is currently
withheld from Apps → Browse. Setting
PAPERCLIP_VERCEL_CONNECT_ENABLED=true does not expose a customer-facing setup
entry. Native provider setup screens remain unchanged. Vercel-hosted deployments use the
workload OIDC token Vercel injects. Other hosted and self-hosted deployments
can provide PAPERCLIP_VERCEL_CONNECT_ACCESS_TOKEN as a deployment bootstrap
secret only when that token type is accepted by the live Connect API:
services:
paperclip:
environment:
PAPERCLIP_VERCEL_CONNECT_ENABLED: "true"
PAPERCLIP_VERCEL_CONNECT_ACCESS_TOKEN: ${PAPERCLIP_VERCEL_CONNECT_ACCESS_TOKEN}
Do not save that access token in a company secret or connection config. It is instance bootstrap authority for the operator-selected Vercel account. A token's long expiry and broad Vercel scope do not prove Connect compatibility; validate it with connector metadata before rollout. Workload OIDC takes precedence when both authorities are present. Turning the feature flag off hides new Vercel-backed setup; existing connections keep resolving while workload OIDC or the bootstrap token remains available. Missing or invalid authority fails closed. See the Vercel Connect operator guide.
Claude + Codex Local Adapters in Docker
The image pre-installs:
claude(Anthropic Claude Code CLI)codex(OpenAI Codex CLI)python3(used to run local subscription sign-in in a pseudo-terminal)
Start the container, then open Apps → Connections to connect a Claude or Codex subscription in your browser. You can also enter a provider API key there. You do not run a shell command or configure a CLI home.
docker run --name paperclip \
-p 3100:3100 \
-e HOST=0.0.0.0 \
-e PAPERCLIP_HOME=/paperclip \
-v "$(pwd)/data/docker-paperclip:/paperclip" \
paperclip-local
Notes:
- The image provides the CLI and Python prerequisites for local subscription sign-in.
- Adapter environment checks in Paperclip surface missing authentication or CLI prerequisites.
Podman Quadlet (systemd)
The docker/quadlet/ directory contains unit files to run Paperclip + PostgreSQL as systemd services via Podman Quadlet.
| File | Purpose |
|---|---|
docker/quadlet/paperclip.pod |
Pod definition — groups containers into a shared network namespace |
docker/quadlet/paperclip.container |
Paperclip server — joins the pod, connects to Postgres at 127.0.0.1 |
docker/quadlet/paperclip-db.container |
PostgreSQL 17 — joins the pod, health-checked |
Setup
-
Build the image (see above).
-
Copy quadlet files to your systemd directory:
# Rootless (recommended) cp docker/quadlet/*.pod docker/quadlet/*.container \ ~/.config/containers/systemd/ # Or rootful sudo cp docker/quadlet/*.pod docker/quadlet/*.container \ /etc/containers/systemd/ -
Create a secrets env file (keep out of version control):
cat > ~/.config/containers/systemd/paperclip.env <<EOL BETTER_AUTH_SECRET=$(openssl rand -hex 32) PAPERCLIP_TOOL_ACTION_SIGNING_SECRET=$(openssl rand -hex 32) POSTGRES_USER=paperclip POSTGRES_PASSWORD=paperclip POSTGRES_DB=paperclip DATABASE_URL=postgres://paperclip:paperclip@127.0.0.1:5432/paperclip # OPENAI_API_KEY=sk-... # ANTHROPIC_API_KEY=sk-... EOL -
Create the data directory and start:
mkdir -p ~/.local/share/paperclip systemctl --user daemon-reload systemctl --user start paperclip-pod
Quadlet management
journalctl --user -u paperclip -f # App logs
journalctl --user -u paperclip-db -f # DB logs
systemctl --user status paperclip-pod # Pod status
systemctl --user restart paperclip-pod # Restart all
systemctl --user stop paperclip-pod # Stop all
Quadlet notes
- First boot: Unlike Docker Compose's
condition: service_healthy, Quadlet'sAfter=only waits for the DB unit to start, not for PostgreSQL to be ready. On a cold first boot you may see one or two restart attempts injournalctl --user -u paperclipwhile PostgreSQL initialises — this is expected and resolves automatically viaRestart=on-failure. - Containers in a pod share
localhost, so Paperclip reaches Postgres at127.0.0.1:5432. - PostgreSQL data persists in the
paperclip-pgdatanamed volume. - Paperclip data persists at
~/.local/share/paperclip. - For rootful quadlet deployment, remove
%hprefixes and use absolute paths.
Onboard Smoke Test (Ubuntu + npm only)
Use this when you want to mimic a fresh machine that only has Ubuntu + npm and verify:
npx paperclipai onboard --yescompletes- the server binds to
0.0.0.0:3100so host access works - onboard/run banners and startup logs are visible in your terminal
Build + run:
./scripts/docker-onboard-smoke.sh
Open: http://localhost:3131 (default smoke host port)
Useful overrides:
HOST_PORT=3200 PAPERCLIPAI_VERSION=latest ./scripts/docker-onboard-smoke.sh
PAPERCLIP_DEPLOYMENT_MODE=authenticated PAPERCLIP_DEPLOYMENT_EXPOSURE=private ./scripts/docker-onboard-smoke.sh
SMOKE_DETACH=true SMOKE_METADATA_FILE=/tmp/paperclip-smoke.env PAPERCLIPAI_VERSION=latest ./scripts/docker-onboard-smoke.sh
Notes:
- Persistent data is mounted at
./data/docker-onboard-smokeby default. - Container runtime user id defaults to your local
id -uso the mounted data dir stays writable while avoiding root runtime. - Smoke script defaults to
authenticated/privatemode soHOST=0.0.0.0can be exposed to the host. - Smoke script defaults host port to
3131to avoid conflicts with local Paperclip on3100. - Smoke script also defaults
PAPERCLIP_PUBLIC_URLtohttp://localhost:<HOST_PORT>so bootstrap invite URLs and auth callbacks use the reachable host port instead of the container's internal3100. - In authenticated mode, the smoke script defaults
SMOKE_AUTO_BOOTSTRAP=trueand drives the real bootstrap path automatically: it signs up a real user, runspaperclipai auth bootstrap-ceoinside the container to mint a real bootstrap invite, accepts that invite over HTTP, and verifies board session access. - Run the script in the foreground to watch the onboarding flow; stop with
Ctrl+Cafter validation. - Set
SMOKE_DETACH=trueto leave the container running for automation and optionally write shell-ready metadata toSMOKE_METADATA_FILE. - Set
SMOKE_CONTAINER_NAMEto fix the container's name up front. Automation that has to collect diagnostics when the script fails needs a name it already knows, rather than one it can only read back out of a successful run. Defaults to the image name. - The container's logs are dumped to
SMOKE_LOG_FILE(default$TMPDIR/<container name>.log) before the script tears the container down, so a run that never became ready still leaves its logs behind. - The image definition is in
docker/Dockerfile.onboard-smoke.
General Notes
- The
docker-entrypoint.shadjusts the containernodeuser UID/GID at startup to match the values passed viaUSER_UID/USER_GID, avoiding permission issues on bind-mounted volumes. - Paperclip data persists via Docker volumes/bind mounts (compose) or at
~/.local/share/paperclip(quadlet).
Native Runner build cache
The image compiles the native Runner in runner-build, before copying the
application source. A pinned cargo-chef generates a dependency recipe in
runner-plan. The separate runner-deps stage compiles that recipe with the
package-owned Rust compiler. Both the dependency build and the real binary use
the release profile and locked Cargo dependencies. The recipe stage never
modifies source in the checkout.
Changes to Rust source or embedded protocol inputs rebuild the real binary but
can reuse compiled dependencies when the recipe is unchanged. Dependency
manifests, the Cargo lockfile, target metadata, or compiler changes invalidate
the relevant cache. Ordinary server or UI changes can reuse the entire native
build through the existing registry cache (mode=max). Each platform gets its
own native build; no cross-architecture binary is reused. No additional GitHub
Actions cache is created. A cold build also installs the recipe generator and
compiles dependencies, so the savings apply after those layers are available.
Cloud builds import one registry cache: the first available full-SHA cache in
the current commit's ten-entry first-parent ancestry, with the legacy cache
as a final fallback. Each build still exports its own SHA cache with
mode=max. In fresh-builder checks, importing several historical manifests
missed native layers that a single matching manifest reused. The selector
inspects metadata after Docker login, stops at the first available cache, and
permits a cold build if no cache can be read.
The application build inherits that stage and still runs the normal server build, including Cargo, binary staging, and generated-contract checks. Rust input file times are normalized in both stages so fresh checkouts do not force Cargo to rebuild unchanged source. Changes made by build scripts still reach Cargo's normal validation. The final application copy excludes Cargo's target directory as before. Cache misses only cost compilation time.
Pull requests that change the Dockerfile, Docker ignore rules, or Runner native
inputs also build the isolated runner-build target in Docker Runner check.
The check runs bash scripts/check-docker-runner-cache.sh against a disposable
copy of tracked source and the actual Docker ignore rules. It compiles a baseline
and exports a local cache, removes that builder, changes a Rust metadata constant,
and rebuilds on a fresh builder using only the exported cache. It requires a
cached dependency build, an unchanged dependency recipe, and changed metadata
from the real binary. It also verifies that a dependency declaration change
alters the recipe. The probe exports small metadata results instead of importing
a large test image into the Docker daemon. Temporary builders and cache files
are removed afterward. It catches missing embedded inputs before the post-merge
build. It uses a GitHub-hosted runner with read-only repository access and never
publishes images or registry caches. Allow up to 20 minutes for its cold build and
source rebuild.