mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Repo-only project workspaces are materialized by a server-side `git clone`, and isolated `git_worktree` runs refresh their base ref with server-side `git fetch` > - Both operations run outside the agent process with no credentials, so private GitHub repositories can never be cloned or refreshed — agent-scoped credential env bindings do not reach them > - The company secret store already has a well-known GitHub token convention (`GITHUB_TOKEN` / `GH_TOKEN` / `PAPERCLIP_GITHUB_TOKEN`, consumed by the external-object provider for API reads), but nothing server-side consults it for git > - This pull request resolves that token per run and authenticates the managed clone and every base-ref refresh with it through an ephemeral credential helper > - The benefit is that isolated workspaces work on private repositories with one company secret, while public repositories and self-hosted ambient git configuration keep working unchanged ## Linked Issues or Issue Description **Subsystem affected** Server workspace materialization (`server/src/services/heartbeat.ts`) and execution-workspace realization (`server/src/services/workspace-runtime.ts`). **Problem or motivation** A project workspace configured with only a private GitHub `repoUrl` cannot be used for isolated `git_worktree` runs: the managed `git clone` runs with a sanitized, credential-less environment, and plain git cannot consume a bare token env variable without a credential helper. There is no way to give the server a git credential — storing a `GH_TOKEN` company secret has no effect on server-side git, and a credential-less private clone hangs on a terminal prompt until the ten-minute clone timeout. Base-ref refreshes (`git fetch`) during worktree realization have the same gap. **Proposed solution** A `git-credentials` module resolves a token per run — company secret by well-known name (`GITHUB_TOKEN`, `GH_TOKEN`, `PAPERCLIP_GITHUB_TOKEN`), then `GITHUB_TOKEN`/`GH_TOKEN` in the server process environment for self-hosted deployments, then none — and builds a git invocation that authenticates via an inline credential helper. The token travels in an env variable; it never appears in argv, URLs, or on disk. Only `https://github.com` remotes are authenticated; everything else keeps ambient behavior. The provider is a single factory seam so a future brokered credential source can replace it without touching call sites. **Alternatives considered** - A GitHub OAuth "connect your account" flow: heavier product surface, needs app registration and callback custody; out of scope for a server credential and better served by a dedicated connector later. The provider seam keeps that path open. - `gh auth setup-git`: writes helper configuration to disk and requires a global token env; rejected in favor of per-invocation config with no persistent state. - Embedding the token in the clone URL: leaks into argv, error messages, and `.git/config`; rejected. ## What Changed - New `server/src/services/git-credentials.ts`: `createGitRemoteAuthProvider` (memoized per run, one secret resolution and one audit event), `buildGitAuthInvocation` (helper-reset + inline helper, `x-access-token` username, `GIT_TERMINAL_PROMPT=0`), `isGitHubHttpsRemoteUrl` host gating (rejects ssh/GHES/http/other hosts/userinfo URLs), `describeGitAuthFailure`, and the canonical `scrubGitCredentialText`. Secret resolutions pass a `system` consumer access context so they are recorded as secret access events. - `ensureManagedProjectWorkspace` (now exported) accepts an optional auth provider; the clone env spreads the token after `sanitizeRuntimeServiceBaseEnv` (which strips `PAPERCLIP_*`), always sets `GIT_TERMINAL_PROMPT=0`, distinguishes "credential rejected" from "no credential configured — add a GITHUB_TOKEN or GH_TOKEN company secret" in the error, and removes the partially created directory on clone failure so a timeout-killed clone cannot be adopted as a broken checkout by the next run. - `refreshRemoteTrackingBaseRef` (now exported) captures the remote URL it already looked up, asks the provider for an invocation, and attributes failed authenticated fetches to the credential in a scrubbed warning. The optional provider threads through `detectDefaultBranch`, `resolveAuthoritativeBaseRef`, `inspectExecutionWorkspaceBaseDrift`, `realizeExecutionWorkspace`, and `ensurePersistedExecutionWorkspaceAvailable`; heartbeat builds one provider per run for both the anchor-resolution clone path and workspace realization/restore. - `github-external-object-provider.ts` imports the shared secret-name list; `isGitHubDotCom` is exported from `github-fetch.ts`. - Docs: "Private repositories and repo-only project workspaces" section in the execution-workspaces guide, cross-linked from the secrets deploy doc. ## Verification - `cd server && npx vitest run src/__tests__/git-credentials.test.ts` — resolution chain order and precedence, env fallback, memoization, audited access context, host-gating matrix, invocation shape (token absent from argv), scrubber, failure descriptions, and a real-git `git credential fill` round trip that proves the helper executes and answers with the env-carried token (no network). - `cd server && npx vitest run src/__tests__/heartbeat-managed-clone-credentials.test.ts` — clones behave byte-identically with no provider or a null-returning provider (local repos, no network), authenticated-failure errors name the credential, non-auth failures do not mention credentials, partial clone directories are removed, pre-existing non-git directories keep the "Using it as-is" path, and the sanitizer spread order keeps the token env alive. - `cd server && npx vitest run src/__tests__/workspace-runtime.test.ts` — new `refreshRemoteTrackingBaseRef` cases: provider offered the remote URL and null keeps behavior identical; failed authenticated fetch warning names the credential; unauthenticated failure warning stays credential-free. - `pnpm --filter @paperclipai/server typecheck` is clean. - Manual (optional, networked): store a `GH_TOKEN` company secret, configure a repo-only project workspace pointing at a private GitHub repository, run an isolated-workspace issue — the managed clone succeeds and the worktree run proceeds. ## Risks - Every new parameter is optional; with no provider the git invocations are byte-identical to before. Public repos and ambient credential helpers keep working whenever no token resolves. - Precedence change when a token exists: a stored company secret now wins over ambient helpers for `https://github.com` remotes (the helper list is reset for that invocation). The rejected-credential error names the secret so an operator can fix or remove it. - `GIT_TERMINAL_PROMPT=0` on the managed clone is the one always-on change: a credential-less private clone now fails fast with a clear message instead of hanging until the ten-minute timeout (it could only ever "succeed" interactively on a TTY dev server). - The token is scoped to the git process env for one invocation; it is never written to agent env, run context, disk, or logs, and error text is scrubbed of URL userinfo. - No migrations, no image changes (git ships in the image). ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
126 lines
8.9 KiB
Markdown
126 lines
8.9 KiB
Markdown
---
|
|
title: Execution Workspaces And Runtime Services
|
|
summary: How project runtime configuration, execution workspaces, and issue runs fit together
|
|
---
|
|
|
|
This guide documents the intended runtime model for projects, execution workspaces, and issue runs in Paperclip.
|
|
|
|
Paperclip now presents this as a workspace-command model:
|
|
|
|
- `Services` are long-running commands that stay supervised.
|
|
- `Jobs` are one-shot commands that run once and exit.
|
|
- Raw runtime JSON is still available for advanced config, but it is no longer the primary mental model.
|
|
|
|
## Project runtime configuration
|
|
|
|
You can define how to run a project on the project workspace itself.
|
|
|
|
- Project workspace runtime config describes the services and jobs available for that project checkout.
|
|
- This is the default runtime configuration that child execution workspaces may inherit.
|
|
- Defining the config does not start anything by itself.
|
|
|
|
## Runtime control: manual and heartbeat-driven
|
|
|
|
Workspace commands can be controlled manually from the UI, and heartbeat runs also start services automatically.
|
|
|
|
- Project workspace services are started and stopped from the project workspace UI, and project jobs can be run on demand there.
|
|
- Execution workspace services are started and stopped from the execution workspace UI, and execution-workspace jobs can be run on demand there.
|
|
- Heartbeat runs also auto-start the workspace's runtime services at the beginning of an issue run. `ensureRuntimeServicesForRun` (`server/src/services/workspace-runtime.ts`, called from `server/src/services/heartbeat.ts`) starts each service whose desired state resolves to `running` — which is the default when no explicit per-service desired state is set. A running service that matches an existing reuse key is reused rather than restarted.
|
|
- You can opt a service out of that auto-start by setting its desired state to `stopped`/`manual` in the runtime config; those services stay UI-controlled.
|
|
- Paperclip does not automatically restart workspace services on server boot — services only come back up when the next run (or a manual start) brings them up.
|
|
|
|
## Execution workspace inheritance
|
|
|
|
Execution workspaces isolate code and runtime state from the project primary workspace.
|
|
|
|
- An isolated execution workspace has its own checkout path, branch, and local runtime instance.
|
|
- The runtime configuration may inherit from the linked project workspace by default.
|
|
- The execution workspace may override that runtime configuration with its own workspace-specific settings.
|
|
- The inherited configuration answers "which commands exist and how to run them", but any running service process is still specific to that execution workspace.
|
|
|
|
## Issues and execution workspaces
|
|
|
|
Issues are attached to execution workspace behavior, not to automatic runtime management.
|
|
|
|
- An issue may create a new execution workspace when you choose an isolated workspace mode.
|
|
- An issue may reuse an existing execution workspace when you choose reuse.
|
|
- Multiple issues may intentionally share one execution workspace so they can work against the same branch and running runtime services.
|
|
- Running an issue auto-starts the workspace's `running`-desired runtime services for the duration of the run (see "Runtime control" above); it does not stop them when the run ends unless they are ephemeral and no other run holds a lease.
|
|
|
|
## Execution workspace lifecycle
|
|
|
|
Execution workspaces are durable until a human closes them.
|
|
|
|
- The UI can archive an execution workspace.
|
|
- Closing an execution workspace stops its runtime services and cleans up its workspace artifacts when allowed.
|
|
- Shared workspaces that point at the project primary checkout are treated more conservatively during cleanup than disposable isolated workspaces.
|
|
|
|
## Resolved workspace logic during heartbeat runs
|
|
|
|
Heartbeat resolves a workspace for the run (code location and session continuity) and also brings up that workspace's runtime services.
|
|
|
|
1. Heartbeat resolves a base workspace for the run.
|
|
2. Paperclip realizes the effective execution workspace, including creating or reusing a worktree when needed.
|
|
3. Paperclip persists execution-workspace metadata such as paths, refs, and provisioning settings.
|
|
4. Heartbeat passes the resolved code workspace to the agent run.
|
|
5. Heartbeat calls `ensureRuntimeServicesForRun` to start the workspace's `running`-desired runtime services, running the lazy runtime provision command first if one is configured and has not yet run (see "Lazy runtime provisioning" below).
|
|
|
|
## Lazy runtime provisioning
|
|
|
|
Some workspaces need heavy one-time setup — seeding a database, warming caches — before their runtime services can start. That work can be deferred to the first runtime-service start instead of running eagerly during workspace preparation.
|
|
|
|
- Configure a **runtime provision command** on the project's workspace strategy (Project properties → execution workspace), or override it per execution workspace on the workspace's Configuration tab.
|
|
- When set, workspace preparation stays lean and the command runs exactly once, immediately before the first runtime-service start for that workspace. Leaving it empty keeps the legacy eager path (all setup during workspace provisioning).
|
|
- The command's outcome is recorded as a `workspace_runtime_provision` operation on the execution workspace and surfaced on the workspace detail page:
|
|
- **Deferred** — configured but not yet run (no runtime service has started yet).
|
|
- **Provisioned at <time>** — the command completed successfully.
|
|
- **Provisioning failed** — the command failed; the workspace detail links to the runtime logs for the failing operation.
|
|
- While the command runs, the runtime service shows a **Provisioning…** state before it transitions to starting/running.
|
|
|
|
## Private repositories and repo-only project workspaces
|
|
|
|
A project workspace can be **repo-only**: a `Repo URL` with no local path. The server then
|
|
materializes a managed checkout on demand (`git clone` into a managed directory) and, for
|
|
isolated `git_worktree` runs, refreshes the base ref (`git fetch`) before preparing each
|
|
worktree. Both operations run on the server, outside any agent process — so agent-scoped
|
|
credential env bindings do not apply to them.
|
|
|
|
For **private GitHub repositories**, store a token as a **company secret** named one of
|
|
`GITHUB_TOKEN`, `GH_TOKEN`, or `PAPERCLIP_GITHUB_TOKEN` (checked in that order; Settings →
|
|
Secrets). The server resolves it per run and authenticates managed clones and base-ref
|
|
fetches with it. Details and caveats:
|
|
|
|
- Scope: only `https://github.com/...` repo URLs are authenticated this way. SSH URLs, GitHub
|
|
Enterprise hosts, and other providers keep ambient behavior (system git config/credential
|
|
helpers on the server host). URLs that embed their own credentials are never overridden.
|
|
- Fallback: with no matching company secret, the server falls back to a `GITHUB_TOKEN` or
|
|
`GH_TOKEN` variable in the **server process environment** (useful for self-hosted single-tenant
|
|
deployments), then to unauthenticated access — public repos keep working with no setup.
|
|
- The token never appears in command lines, URLs, or on disk; it is passed to git through an
|
|
ephemeral credential helper. Each resolution is recorded as a secret access event.
|
|
- This is separate from the **agent push credential**: agents pushing branches/PRs still need
|
|
`GH_TOKEN`/`GITHUB_TOKEN` bound at agent or project scope (see
|
|
[deploy/secrets](../../deploy/secrets.md)) so the token reaches the agent process env. The
|
|
same company secret can back both uses via a binding.
|
|
|
|
## Cross-run persistence (no-remote-git contract)
|
|
|
|
Code state moves between runs through the local execution-workspace cwd alone — not through a git remote.
|
|
|
|
- Each run's prepare step bundles the local worktree to the run's remote dir over ssh, with no `git remote` configured.
|
|
- The adapter's restore step at the end of the run writes any new remote commits back into the local worktree directly.
|
|
- Adapters must never `git push` from runtime code, and must never assume a remote exists.
|
|
- A failed restore is a run-level error and records `workspace_finalize=failed` on the execution workspace, which gates dependent issue wakes until the next successful finalize.
|
|
|
|
The invariant is enforced by the "no-remote-git contract" case in `packages/adapter-utils/src/ssh-fixture.test.ts`, which asserts a remote-only commit reaches the local worktree with no remote configured at any point.
|
|
|
|
## Current implementation guarantees
|
|
|
|
With the current implementation:
|
|
|
|
- Project workspace command config is the fallback for execution workspace UI controls.
|
|
- Execution workspace runtime overrides are stored on the execution workspace.
|
|
- Heartbeat runs auto-start the workspace's `running`-desired runtime services (via `ensureRuntimeServicesForRun`); services set to `stopped`/`manual` stay UI-controlled.
|
|
- A configured runtime provision command runs once, lazily, before the first runtime-service start.
|
|
- Server startup does not auto-restart workspace services.
|