mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/plugin-task-execution
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c57c0f7498 |
fix(sandbox-providers): accept bsdtar listings in the syncOut tarball confinement check (#11289)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers (Daytona, Kubernetes) sync run results back to the host with a sandbox-authored tarball > - Before extraction, a confinement check parses the host `tar -tvf` listing and fails closed on unparseable lines > - The parser only understands the GNU tar listing dialect; macOS ships bsdtar, whose ls-style listing never matches > - Every sandbox syncOut on a macOS host therefore aborts with "refusing tarball with an unparseable entry listing", and the run fails at copy-back > - This pull request teaches the parser both dialects while keeping the fail-closed and traversal guarantees > - The benefit is that Daytona and Kubernetes sandbox runs work on macOS hosts, with no behavior change on Linux ## Linked Issues or Issue Description No existing issue. Bug description: **What happened?** On a macOS host, every Daytona sandbox run fails at syncOut. The adapter reports: `Daytona syncOut refusing tarball with an unparseable entry listing: -rw-r--r-- 0 daytona daytona 7560 Aug 11 21:43 AGENTS.md`. The Kubernetes provider has the same parser and fails the same way. **Expected behavior** The confinement check accepts a well-formed listing from the host tar, whichever dialect the host tar emits. It still rejects members that escape the extraction directory, and it still fails closed on lines it cannot parse. **Steps to reproduce** 1. Run Paperclip on macOS (system tar is bsdtar). 2. Configure an agent with the Daytona sandbox provider. 3. Trigger any run that syncs files back from the sandbox. 4. The run fails at syncOut with the unparseable-entry-listing error, because bsdtar prints `<perms> <links> <user> <group> <size> <Mon> <day> <time|year> <name>` while the parser expects the GNU `<perms> <owner>/<group> <size> <date> <time> <name>` shape. **Operating system** macOS (bsdtar 3.5.3). Linux hosts with GNU tar are unaffected. ## What Changed - Extracted the listing-line parse in both providers' `file-sync.ts` into an exported `parseTarVerboseListingLine` that accepts the GNU/busybox dialect and the bsdtar (libarchive) dialect. - The GNU shape now requires the slash-joined `<owner>/<group>` field. This keeps the two shapes mutually exclusive. Without it, a bsdtar line with numeric uid/gid satisfies the loose GNU pattern shifted by one field, which would hide a leading `../` from the traversal check. - Unparseable lines still fail closed. This includes device-node entries, whose size column is `major,minor` in both dialects. - Made the path-traversal fixture in the Daytona suite portable: GNU spells member renaming `--transform`, bsdtar spells it `-s`. - Added a Daytona test that refuses a sandbox-authored tarball carrying a symlink whose target escapes the extraction dir. - Added parser unit tests for both dialects (file, dir, symlink, hardlink, numeric owner, year-form dates, fail-closed lines) to both providers' suites. ## Verification - `pnpm test` in `packages/plugins/sandbox-providers/daytona`: 136/136 pass on a macOS host. On unpatched `master` the round-trip test fails there with the unparseable-entry-listing error. - `pnpm test` in `packages/plugins/sandbox-providers/kubernetes`: the new parser tests pass; no new failures against the `master` baseline on the same host. - `pnpm typecheck` passes in both packages. - CI runs the same suites on Linux/GNU tar and proves the GNU path is unchanged. ## Risks - Low risk. The GNU pattern is one token stricter (`<owner>/<group>` must contain `/`). GNU and busybox tar always print the slash-joined owner field, so accepted GNU listings are unchanged. - The bsdtar branch only widens acceptance on hosts that were failing 100% of syncOuts before, so no working deployment changes behavior. - The check still fails closed on anything neither pattern matches. ## Model Used Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking + tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0ccba45e4d |
feat(sandbox-providers): native providers honor postUploadCommands (Daytona executes; Kubernetes executes) (#10347)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies, so runtime handoffs have to preserve the exact behavior an agent asked for. > - The sandbox-provider layer is where uploaded files and follow-up commands become a real in-sandbox operation. > - The provider-delegable sync-in seam already exists from the prior PR; without this follow-up, native providers can still accept command-bearing uploads and silently drop the commands. > - That is a fail-open gap for native sandbox providers, because the file transfer succeeds while the intended post-upload work never runs. > - This pull request teaches Daytona and Kubernetes to execute `postUploadCommands` in order, inside the sandbox, after the files land. > - The benefit is consistent and safer sync semantics: native providers either run the commands as requested or fail fast instead of pretending the operation completed fully. ## Linked Issues or Issue Description This PR builds on the previously merged provider-delegable sync-in seam and closes the remaining gap for native providers that still dropped `postUploadCommands`. Problem: - A sync operation could include ordered `postUploadCommands`, but a native provider could finish the file upload and skip the commands entirely. - That creates fail-open behavior for command-bearing uploads, especially when the caller relies on the provider to execute the follow-up action in the sandbox. Proposed fix: - Execute `postUploadCommands` inside the sandbox after file placement. - Preserve the provided command order. - Fail fast on the first non-zero exit or timeout. - Keep command execution verbatim and confine any provided `cwd` under the workspace root. Related public PR: - Refs: #10340 ## What Changed - Daytona `performSyncIn` now executes ordered `postUploadCommands` through the existing `executeCommand` seam. - Kubernetes `performSyncIn` now executes ordered `postUploadCommands` through its streaming pod exec path. - Added workspace confinement for provided `cwd` values and defaulted missing `cwd` to the remote root. - Added tests covering the new post-upload command execution behavior in both provider packages. ## Verification - Latest validation recorded on the handoff branch: `@paperclipai/plugin-sdk` and `@paperclipai/plugin-kubernetes` typechecks passed. - Daytona Vitest: `63/63` passing in `plugin.test.ts`. - Kubernetes Vitest: `195/195` passing, including `file-sync.test.ts`. - `git log --oneline origin/master..HEAD` showed a single expected commit on the branch. ## Risks - Command execution semantics are stricter now, so malformed commands or a bad `cwd` will fail the sync instead of being ignored. - The change makes provider behavior more explicit, which can surface previously hidden failures in callers that assumed commands were optional. - Timeout behavior may differ slightly between providers, so the failure mode is intentionally fail-fast. ## Model Used OpenAI Codex, GPT-5-based coding agent with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7f766526a6 |
feat(sandbox): add task-scoped egress grants (#10155)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9edde68373 |
feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - AI agents run in sandboxed execution environments (Kubernetes pods, Daytona workspaces, etc.) and need to sync files between the host and those environments — for workspace setup, asset delivery, and output retrieval > - The existing sync path for Kubernetes uses a base64-over-exec chunk loop: each ~4 MB chunk requires its own `execInPod` round-trip, so large syncs balloon into many exec calls with corresponding overhead > - `execInPod` supports piped stdin/stdout, meaning the full transfer can be done as a single exec that streams a raw `tar` archive over the data channel — one round-trip regardless of file size, with nothing base64-encoded and nothing buffered whole in memory on either side > - PR-1 (#10013, merged) added the `onEnvironmentSyncIn`/`onEnvironmentSyncOut` opt-in hook API to the sandbox provider interface and documented the protocol; PR-2 (#10028, merged) implemented these hooks for the Daytona provider > - This pull request implements the same two lifecycle hooks in the Kubernetes sandbox provider, so workspace/asset file sync streams through one `execInPod` per operation instead of the chunk loop > - The benefit is significantly fewer exec round-trips for large syncs and flat memory use on both host and pod, with security properties preserved: atomic replace, secret-mode enforcement, path confinement, TOCTOU-safe snapshot, and member-confinement on host-assembled archives from sandbox-authored tar output ## Linked Issues or Issue Description This is the third and final PR in a sequential series: - Refs #10013 — PR-1: opt-in sync hook API + provider docs (merged) - Refs #10028 — PR-2: native file-sync lifecycle hooks for Daytona provider (merged) **Feature:** Native single-exec file-sync lifecycle hooks for the Kubernetes sandbox provider. *Motivation:* The existing Kubernetes sync path encodes files as base64 and loops over `execInPod` one chunk at a time (~4 MB per exec). For large workspaces or asset sets this is slow and resource-intensive. The Kubernetes `execInPod` API supports piped stdin/stdout, enabling a raw-`tar` streaming transfer that needs only one exec regardless of file count or size and never buffers the whole payload in memory. *Proposed solution:* Implement `onEnvironmentSyncIn` and `onEnvironmentSyncOut` in the Kubernetes provider using a streaming `execInPod` with a tar pipeline — for syncIn the host builds the archive on disk and streams its raw bytes into the pod's stdin (`head -c <exact-size> | tar -x`, no base64); for syncOut in-pod `tar` writes to the exec's stdout and the host streams those bytes straight to a file. Path confinement, atomic replace, secret-mode enforcement, TOCTOU protection, and a streamed-bytes fail-closed guard are all enforced. ## What Changed - **New `src/file-sync.ts`** in `packages/plugins/sandbox-providers/kubernetes/` — `performSyncIn` and `performSyncOut` over an injected pod-exec closure, keeping transfer logic hermetically unit-testable - **New `execInPodStreaming` in `src/pod-exec.ts`** — a streaming exec primitive that binds a caller-supplied stdin readable and a stdout writable to the exec WebSocket data channel, added alongside the existing `execInPod` (which is unchanged). This lets a transfer stream raw bytes to/from disk instead of buffering the payload as a single string - **Updated `src/plugin.ts`** — registers `onEnvironmentSyncIn`/`onEnvironmentSyncOut`; resolves the `sandbox-cr` pod exactly like `onEnvironmentExecute` and delegates; `job` backend rejects file-sync calls explicitly (out of scope) - **syncIn path:** host builds the tarball to a temp file → streams its raw bytes over exec stdin, bounded in-pod by `head -c <exact-archive-size> | tar -x` (no base64 anywhere) → extract into a `/proc/self/fd`-pinned reserved `0700` staging dir → `chmod`-before-`mv -f` atomic replace per file (directory mappings use `followSymlinks`→`-h`) - **syncOut path:** in-pod validate + realpath-snapshot each source (closes the validation→copy TOCTOU window) → single-exec `tar -c` streamed over exec stdout → host streams that stdout straight to a temp file through a byte-counting transform → member-confined extraction of the sandbox-authored archive - **Security properties:** secret files land at requested mode with no widened window; every interpolated path is shell-quoted and confined lexically plus via in-pod `realpath`; the outbound stream is bounded by a **streamed-bytes disk guard** (`MAX_SYNC_OUTPUT_BYTES`, 8 GiB default, per-call overridable) that fails the transfer closed — writing no target file — if an untrusted pod emits more bytes than allowed. Neither host nor pod buffers the whole payload, so there is no in-memory size cap on the transfer - **No changes** to `execInPod`, `wrapCommandWithEnv`, or `FastUploadInterceptor` (the `environmentExecute` path is untouched) - **No dependency or lockfile changes** - **New tests** in `test/unit/file-sync.test.ts` (atomic-replace, `0600` secret mode, symlink preserve/deref, dir-mapping, exclude, path-confinement rejection, streamed-output guard fail-closed) and `test/unit/pod-exec.test.ts` (streaming stdin/stdout, caller-sink error fail-closed), plus extended `test/unit/plugin.test.ts` ## Follow-up: Legacy Job-Lease Base64 Fallback Fix Addresses the Greptile 4/5 blocking finding ("Handle existing job leases", `server/src/services/environment-runtime.ts`). Job leases provisioned before the `nativeFileSyncUnsupported` metadata flag existed carry `backend: "job"` but no flag, so `supportsSync()` treated them as native-capable and routed their sync to the pod-exec hook — which the job backend rejects (it has no exec channel) instead of using the byte-identical base64 fallback. The fix adds a belt-and-suspenders gate on the persisted `backend === "job"` field alongside the existing `nativeFileSyncUnsupported` flag check, so pre-existing job leases continue syncing via the base64 fallback after deployment. No behaviour change for `sandbox-cr` leases. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-kubernetes test` — 19 files / 182 tests green, including the existing `upload-interceptor` and `pod-exec` suites - `tsc --noEmit` in the kubernetes package — 0 errors - The sync hooks are opt-in; existing `environmentExecute` behaviour is unaffected and tested by the unchanged existing suites ## Risks - **Opt-in only:** `onEnvironmentSyncIn`/`onEnvironmentSyncOut` are registered conditionally; providers that do not register them fall back to the existing chunk loop. No regression risk on the existing path. - **Shell-injection surface:** all path interpolation uses shell-quoting; paths are additionally confined lexically and via in-pod `realpath` before use. - **TOCTOU on syncOut:** the in-pod snapshot validates and records file metadata before the tar call, closing the window between validation and copy. - **Archive member confinement:** host-side reassembly rejects any tar member whose resolved path escapes the target directory, preventing a malicious in-pod tar from writing outside the intended destination. - **Untrusted-output volume:** an over-large outbound stream trips the streamed-bytes disk guard and fails closed (no target written and the temp sink is swept) rather than filling host disk or memory; the guard bounds disk unconditionally and bounds memory insofar as WebSocket write-backpressure holds. ## Model Used Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`) — produced by a Claude-based AI agent using agentic tool use and multi-step code generation. 200K context window, extended reasoning, code execution and verification capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
482f64e343 |
fix(plugin-kubernetes): resolve sandbox pod by exact name (controller labels pods with sandbox-name-hash, not sandbox-name) (#7982)
## Thinking Path Production e2e on the merged #5790 plugin failed on every fresh lease with "Failed to install the adapter runtime command" for a harness that was present in the runtime image. Tracing the lease showed the first exec resolved no pod: the exact-label fallback added during the #5790 review queries `agents.x-k8s.io/sandbox-name=<name>`, but the kubernetes-sigs agent-sandbox controller labels pods only with `agents.x-k8s.io/sandbox-name-hash` (see `sandboxLabel` in its `controllers/sandbox_controller.go`) and NAMES the backing pod exactly after the Sandbox CR. The selector matches nothing, `findPodForSandbox` returns null, execute returns "podName could not be resolved", and adapter-utils misreports it as a missing runtime command. ## What Changed Between the `status.podName` read and the label fallback, try an exact-name pod GET (`readNamespacedPod({namespace, name})`). This is collision-free, so the original review concern (name-prefix matching execing into a concurrent sandbox's pod) stays honored. A 404 falls through to the existing full-name label selector for controller versions that do set such a label. Non-404 errors propagate unchanged. ## Verification - New unit test pins the controller reality: pod named exactly like the sandbox, only a `sandbox-name-hash` label, no full-name label; fails before the fix, passes after. - Review-feedback round: the primary-path test now asserts the exact-name GET is never called, and a new test covers non-404 error propagation (403 rejects, no fallback). 153/153 plugin tests green, tsc clean. - Production-verified on our deployment: agent runs were broken on every fresh lease before this patch and complete end-to-end after it (gVisor sandbox pool, agent-sandbox controller v0.4.6; verified run with cost event and agent reply on a fresh tenant). ## Risks Low: one additional pod GET per first-exec on a fresh lease, only when `status.podName` is unset. Non-404 errors from the GET propagate unchanged (now test-pinned). ## Issue No existing issue; the defect is described in full under Thinking Path (introduced by the review-round fallback change in #5790, first hit in production e2e on 2026-06-11). ## Model Used Claude Fable 5 (claude-fable-5, Claude Code CLI, extended reasoning, tool use) ## Duplicate search Searched open and closed PRs for `findPodForSandbox`, `sandbox-name-hash`, and pod-resolution fixes; no duplicate found. Related parent: #5790 (introduced the fallback this PR repairs). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (no UI change) - [x] I have updated relevant documentation to reflect my changes (code comments; no doc surface affected) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
05ab45225a |
feat(plugin-kubernetes): self-hostable Kubernetes sandbox provider (stage 1/3: plugin package) (#5790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers are the seam that lets agent runs execute in isolated environments; today the only first-party remote provider is Daytona, a hosted third-party service > - Self-hosters running Paperclip on their own infrastructure (often Kubernetes already) have no first-party way to run agent sandboxes on a cluster they control > - That gap matters for teams with data-residency, sovereignty, or cost constraints who cannot or will not send workloads to a hosted sandbox service > - This pull request adds a Kubernetes sandbox-provider plugin as a standalone, workspace-excluded package: it implements every SandboxProvider hook the Daytona provider does, on infrastructure the operator owns > - The benefit is that any Paperclip deployment with a Kubernetes cluster gets multi-tenant, network-isolated, quota-bounded agent sandboxes with zero new external dependencies ## Linked Issues or Issue Description No existing issue. Following the feature template: - **Problem:** Paperclip's remote sandbox execution requires a hosted third-party provider. Self-hosters cannot run agent sandboxes on their own Kubernetes clusters with a first-party provider. - **Proposed solution:** A `@paperclipai/plugin-kubernetes` sandbox-provider plugin with two backends: long-lived sandboxes via the [kubernetes-sigs/agent-sandbox](https://github.com/kubernetes-sigs/agent-sandbox) CRD (multi-command exec, adapter-install pattern) and one-shot `batch/v1` Jobs (stable APIs only, no extra controllers). - **Alternatives considered:** Driving kubectl from a generic shell provider (no lifecycle/lease semantics), or requiring a hosted provider (exactly the constraint this removes). ## What Changed This is **stage 1 of 3** of a staged contribution (direction agreed with maintainers): the plugin package alone. Stage 2 (server integration: lease params, provider registration) and stage 3 (agent runtime images + CI) are companion PRs that will be cross-linked from a comment here. - New package `packages/plugins/sandbox-providers/kubernetes` (workspace-excluded, like the path already carved out in `pnpm-workspace.yaml`): src, unit + kind integration tests, operator prerequisite manifests, README, smoke-test guide - Implements the full SandboxProvider hook surface the Daytona provider implements: `validateConfig`, `probe`, `acquireLease`, `resumeLease`, `releaseLease`, `destroyLease`, `realizeWorkspace`, `execute` - Two backends: `sandbox-cr` (default; long-lived pod via the agent-sandbox `Sandbox` CR, supports multi-command exec) and `job` (one-shot `batch/v1` Job; nothing beyond k8s 1.27+ required) - Per-run adapter resolution: one environment serves mixed harnesses; the per-run `adapterType` hint is read through a local optional type extension, so the plugin typechecks and builds against the current plugin SDK and simply falls back to the environment's configured default adapter until stage 2 lands - Exec-env wrapping: the Kubernetes exec API carries no environment, so commands are wrapped to receive the run's env - Fast-upload interception for workspace realization, scoped per lease - Per-tenant isolation: derived namespace per company, RBAC, ResourceQuota, restricted-PSS pod security (runAsNonRoot, drop ALL, seccomp RuntimeDefault, no SA token automount) - Network egress policy in two flavors: native `NetworkPolicy` and `CiliumNetworkPolicy` (FQDN allowlists) - Image allowlist with glob matching, registry override, and per-run image override validation - Per-run Kubernetes Secrets carrying agent credentials, ownerRef'd to the Job or Sandbox CR for cascade GC ## Verification - Standalone build, exactly as the README documents: ```bash cd packages/plugins/sandbox-providers/kubernetes pnpm install --ignore-workspace pnpm test # 147 unit tests, 17 files, all green pnpm typecheck # clean against the in-repo plugin SDK on master pnpm build # dist/ emitted, manifest + worker entrypoints present ``` - A kind-cluster end-to-end integration test is included (`RUN_K8S_INTEGRATION_TESTS=1 pnpm test test/integration/end-to-end-run.test.ts`) - Beyond CI: this provider has been verified in a production multi-tenant deployment against five harnesses (opencode, pi, codex, gemini, claude code) with real billed runs ## Risks - **Zero behavior change for any existing deployment.** The package is workspace-excluded; nothing in the server imports or loads it until stage 2's integration lands. No existing code paths are touched. - The default `sandbox-cr` backend depends on an alpha CRD (`agents.x-k8s.io/v1alpha1`); the README flags this and the `job` backend uses only stable APIs as a fallback. - Risk surface is confined to deployments that explicitly install and configure the plugin. - The default runtime images (`ghcr.io/paperclipai/agent-runtime-*`) are published by the stage 3 companion PR (#7934); until that lands, deployments must point `runtimeImage` at their own images. ## Model Used Claude Opus 4.8 (1M context), extended thinking, with tool use (Claude Code). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (no UI changes) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending this push) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |