mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
181 lines
9.9 KiB
Markdown
181 lines
9.9 KiB
Markdown
# @paperclipai/plugin-kubernetes (alpha)
|
|
|
|
First-party Paperclip sandbox-provider plugin for Kubernetes.
|
|
|
|
**Alpha:** the default backend (`sandbox-cr`) is built on `kubernetes-sigs/agent-sandbox` v1alpha1 — expect breaking changes as that CRD evolves toward Beta. A stable fallback backend (`job`, using `batch/v1` Job) is available for clusters without agent-sandbox installed, but it does NOT support multi-command exec (paperclip-server's adapter-install pattern requires sandbox-cr).
|
|
|
|
## Prerequisites
|
|
|
|
### For `sandbox-cr` backend (default, recommended)
|
|
|
|
1. A Kubernetes cluster running k8s 1.27+
|
|
2. [`kubernetes-sigs/agent-sandbox`](https://github.com/kubernetes-sigs/agent-sandbox) controller installed in the cluster (alpha — installs the `sandboxes.agents.x-k8s.io/v1alpha1` CRD and controller)
|
|
3. Paperclip-server running with access to the cluster (in-cluster via `inCluster: true` or external via `kubeconfig`)
|
|
|
|
### For `job` backend (stable fallback)
|
|
|
|
1. A Kubernetes cluster running k8s 1.27+
|
|
2. Paperclip-server with cluster access — no additional controllers or CRDs required
|
|
|
|
## Installation
|
|
|
|
```bash
|
|
paperclipai plugin install @paperclipai/plugin-kubernetes
|
|
```
|
|
|
|
Or, for local development:
|
|
|
|
```bash
|
|
paperclipai plugin install --local /path/to/paperclip/packages/plugins/sandbox-providers/kubernetes
|
|
```
|
|
|
|
## Backends
|
|
|
|
The plugin supports two backend modes, selected via the `backend` config field:
|
|
|
|
| Backend | Default | Stability | Multi-command exec | Requires |
|
|
|---|---|---|---|---|
|
|
| `sandbox-cr` | Yes | Alpha | Yes | `kubernetes-sigs/agent-sandbox` controller |
|
|
| `job` | No | Stable | No | Nothing beyond k8s 1.27+ |
|
|
|
|
**`sandbox-cr` (default):** Creates a `Sandbox` CR (`agents.x-k8s.io/v1alpha1`) whose controller provisions a long-lived pod running `sleep infinity`. paperclip-server execs individual commands into the running pod — this is the multi-command adapter-install pattern. When you `releaseLease`, the Sandbox CR is deleted and the controller tears down the pod.
|
|
|
|
**`job` (stable fallback):** Creates a `batch/v1` Job. The container entrypoint runs once and exits — no multi-command exec possible. Use this when you cannot install agent-sandbox, or when you need strictly stable Kubernetes APIs. Note: paperclip-server's adapter-install pattern will not work in job mode.
|
|
|
|
### Migrating from `job` to `sandbox-cr`
|
|
|
|
1. Install the agent-sandbox controller: `kubectl apply -f https://github.com/kubernetes-sigs/agent-sandbox/releases/latest/download/install.yaml`
|
|
2. Update your environment config to set `backend: "sandbox-cr"` (or remove `backend` since `sandbox-cr` is the default)
|
|
3. New leases will use the Sandbox CR backend. Existing leases created with `job` mode continue to use job semantics until they are released.
|
|
|
|
## Configuration
|
|
|
|
Create a `sandbox` environment with `driver: kubernetes`. One of these auth fields is required:
|
|
|
|
- `inCluster: true` — use the in-pod ServiceAccount credentials (when paperclip-server runs inside the same cluster).
|
|
- `kubeconfig: <YAML>` — inline kubeconfig (stored as a company secret).
|
|
- `kubeconfigSecretRef: <secret-uuid>` — reference to an existing Paperclip secret.
|
|
|
|
Common optional fields:
|
|
|
|
| Field | Default | Purpose |
|
|
|---|---|---|
|
|
| `backend` | `"sandbox-cr"` | `sandbox-cr` (alpha, requires agent-sandbox controller) or `job` (stable, one-shot entrypoint). |
|
|
| `adapterType` | `"claude_local"` | One of the supported adapter types (claude_local, codex_local, gemini_local, cursor_local, opencode_local, pi_local). Determines runtime image + env keys + egress allow-list. |
|
|
| `namespacePrefix` | `"paperclip-"` | Prefix for the per-company tenant namespace. |
|
|
| `companySlug` | derived from companyId | Override the auto-derived company slug. |
|
|
| `imageRegistry` | (none) | Override the default registry for agent runtime images. |
|
|
| `imageAllowList` | `[]` | Glob patterns of allowed `target.imageOverride` values. Empty = no override permitted. |
|
|
| `imagePullSecrets` | `[]` | Names of pre-created Docker image pull secrets in the tenant namespace. |
|
|
| `egressAllowFqdns` | `[]` | Additional FQDNs (beyond adapter defaults like `api.anthropic.com`). |
|
|
| `egressAllowCidrs` | `[]` | Additional CIDRs to allow egress to. |
|
|
| `egressMode` | `"standard"` | `standard` (NetworkPolicy + CIDRs) or `cilium` (CiliumNetworkPolicy + FQDN allow-list). |
|
|
| `runtimeClassName` | (none) | e.g. `kata-fc` for Firecracker-backed microVMs. Cluster must have the RuntimeClass installed. |
|
|
| `serviceAccountAnnotations` | `{}` | Annotations applied to per-tenant ServiceAccount (e.g. IRSA `eks.amazonaws.com/role-arn`). |
|
|
| `jobTtlSecondsAfterFinished` | `900` | Seconds after a Job completes before garbage-collection. |
|
|
| `podActivityDeadlineSec` | `3600` | Hard ceiling on a single run's wall-clock time. |
|
|
|
|
Full JSON Schema in `src/manifest.ts`.
|
|
|
|
### Task-scoped egress grants
|
|
|
|
Keep provider-level egress defaults narrow, then grant only the destinations a task needs through its execution workspace settings:
|
|
|
|
```json
|
|
{
|
|
"executionWorkspaceSettings": {
|
|
"networkEgress": {
|
|
"allowFqdns": ["github.com", "pypi.org"],
|
|
"allowCidrs": []
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
The provider creates a workload-owned policy selected by the task run label, so the additional destinations do not become reachable from other concurrent agent pods. Cilium mode enforces FQDNs directly. Standard NetworkPolicy mode cannot express FQDNs, so an FQDN grant permits public IPv4 TCP 80/443 for that run while excluding private, loopback, link-local, CGNAT, and multicast ranges. Network failures that look policy-related include the grant path in stderr, and the sandbox exposes the effective policy through `PAPERCLIP_NETWORK_EGRESS_*` environment variables.
|
|
|
|
## What gets created in your cluster
|
|
|
|
For each company that runs agents (created lazily on first dispatch):
|
|
|
|
```
|
|
Namespace paperclip-{companySlug} (PSS: restricted enforce + audit)
|
|
ServiceAccount paperclip-tenant-sa
|
|
Role paperclip-tenant-role (only get pods/log)
|
|
RoleBinding paperclip-tenant-rb
|
|
ResourceQuota paperclip-quota (pods, requests/limits cpu+memory)
|
|
LimitRange paperclip-limits (container max/min/default/defaultRequest)
|
|
NetworkPolicy paperclip-deny-all (deny ingress + egress baseline)
|
|
NetworkPolicy paperclip-egress-allow (DNS + paperclip-server callback + user CIDRs)
|
|
OR CiliumNetworkPolicy paperclip-egress-fqdn if egressMode=cilium
|
|
```
|
|
|
|
For each agent run (sandbox-cr backend):
|
|
|
|
```
|
|
Sandbox CR pc-{ulid} (agents.x-k8s.io/v1alpha1; explicit delete on release)
|
|
Pod pc-{ulid}-{podSuffix} (managed by Sandbox controller; torn down on CR delete)
|
|
Secret pc-{ulid}-env (owned by Sandbox CR; cascade-deleted)
|
|
```
|
|
|
|
For each agent run (job backend):
|
|
|
|
```
|
|
Job pc-{ulid} (backoffLimit: 0, ttlSecondsAfterFinished from config)
|
|
Pod pc-{ulid}-{podSuffix} (owned by Job; cascade-deleted)
|
|
Secret pc-{ulid}-env (owned by Job; cascade-deleted)
|
|
```
|
|
|
|
## Security baseline
|
|
|
|
Every agent pod is:
|
|
|
|
- non-root (`runAsUser: 1000`, `runAsGroup: 1000`, `runAsNonRoot: true`)
|
|
- drops ALL Linux capabilities, `allowPrivilegeEscalation: false`
|
|
- `readOnlyRootFilesystem: true` with explicit `emptyDir` mounts for `/workspace`, `/home/paperclip`, `/home/paperclip/.cache`, `/tmp`
|
|
- `seccompProfile: RuntimeDefault`
|
|
- Tini as PID 1 (reaps zombies, forwards signals)
|
|
- `fsGroupChangePolicy: OnRootMismatch` (fast PVC startup; openclaw-operator lesson)
|
|
- `automountServiceAccountToken: true` (for the agent shim's paperclip-server callback)
|
|
|
|
Plus per-namespace `pod-security.kubernetes.io/enforce: restricted` and a deny-all NetworkPolicy baseline with explicit egress allow-list (DNS, paperclip-server, configured FQDNs/CIDRs).
|
|
|
|
The per-run Secret carrying the bootstrap token and adapter API keys has `ownerReferences` pointing at the owning Job, so a single `kubectl delete job …` cascades cleanly to the Pod and Secret.
|
|
|
|
## Optional Kata-FC microVM isolation
|
|
|
|
For stronger isolation, install [Kata Containers](https://github.com/kata-containers/kata-containers) with the Firecracker hypervisor, then set `runtimeClassName: kata-fc` in the plugin config. Each agent pod will run inside a Firecracker microVM. Requires nested-virt-capable nodes (bare-metal or specific cloud instance types).
|
|
|
|
## Roadmap
|
|
|
|
- **Phase A (done):** `sandbox-cr` backend — multi-command exec via agent-sandbox Sandbox CRD.
|
|
- **Phase B:** Warm pool support — pre-provisioned Sandbox CRs for sub-second cold starts. The `SandboxOrchestrator` interface reserves optional `pause?`/`resume?` extension slots.
|
|
- **Phase C:** Kata-FC + snapshots — `runtimeClassName: kata-fc` with VM snapshot for fast restore.
|
|
- **Phase D:** Contribute back to agent-sandbox upstream if their Beta model diverges from our needs. The `SandboxOrchestrator` interface (`src/sandbox-orchestrator.ts`) is the clean swap point — a new implementation can be added without touching `plugin.ts` business logic.
|
|
|
|
## Lessons learned (from openclaw-operator)
|
|
|
|
This plugin adopts patterns from `openclaw-rocks/openclaw-operator`:
|
|
|
|
- Tini PID 1 (issue #471 — zombie helper processes)
|
|
- Read-only rootFS with explicit writable mounts (issue #456 — ~/.config not writable)
|
|
- Strategic merge on reconcile (issue #446 — preserve third-party annotations)
|
|
- Multi-storage-class testing (issue #448 — `local-path-provisioner` differences)
|
|
- Image version compat matrix (issue #462 — runtime deps cannot resolve after upgrade)
|
|
|
|
## Development
|
|
|
|
```bash
|
|
cd packages/plugins/sandbox-providers/kubernetes
|
|
pnpm install --ignore-workspace
|
|
pnpm test # unit tests only (fast)
|
|
pnpm typecheck
|
|
pnpm build
|
|
```
|
|
|
|
To run the kind-cluster integration test (requires `kubectl --context kind-paperclip` and a pre-loaded alpine image; see `test/integration/end-to-end-run.test.ts`):
|
|
|
|
```bash
|
|
RUN_K8S_INTEGRATION_TESTS=1 pnpm test test/integration/end-to-end-run.test.ts
|
|
```
|