mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
codex/runner-pi-diagnostics-package-fix
31
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d172197117 |
feat(storage): add plain directory sync with conflict preflight (#14416)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs use workspace transport to restore and collect files. > - Some directories belong to the agent across tasks. > - Those directories need plain file transport without task Git state. > - A concurrent file edit must be detected before a merge changes any file. > - This pull request adds optional plain-directory sync and conflict preflight. > - Existing task workspace sync keeps its defaults. ## Linked Issues or Issue Description Refs #14325. This is the transport prerequisite for a replacement of its instruction revision design with current agent files. ## What Changed - Add an opt-in plain-directory transport mode to command and sandbox runtimes. - Add strict merge preflight for file edits, deletions, and directory changes. - Accept identical replay after an interrupted merge. Preserve competing changes. - Set the compiled OpenCode test executable to 0755, independent of the CI host’s file-creation mask. Preserve the original startup error if cleanup also fails. ## Verification - At `ced53ae532ce6966cad5a83575d83db61af98126`, all 212 targeted transport tests passed across workspace restore, remote managed runtime, SSH fixture, and execution-target sandbox suites. Adapter-utils typecheck passed. - Real isolated SSH retry fixture previously passed with `PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1`; stale deleted files remain absent while gitignored binary bytes survive. - Dependent PR #14420 passed real native local, legacy local, and native Daytona persistence E2E at `169fab46d5af21caa2269b4c1b29b69c933a6951`, which includes all production transport changes through `ced53ae53`; the subsequent two commits only fix the OpenCode test fixture. Nine tasks, three server restarts, and all cleanup checks passed. - A hosted OpenCode fixture failed twice at `ced53ae53`. Reproduced the failure locally and in Linux with `umask 0002`: the compiler created a group-writable executable, correctly rejected by the qualified launch boundary. Explicit 0755 permissions fix the test without weakening the production guard. The focused test and non-root Linux reproduction now pass under that same mask. - Before rebase, head `69e97de0475d34aac5d532e559a405eaf015fd2b` includes the deterministic fixture permission fix and preserves original bootstrap diagnostics. All production transport code is unchanged since the 212-test validation. Fresh Greptile review is 5/5 on this exact head with no unresolved findings. All 54 current-head checks passed, with two conditional skips. The full CI run completed successfully, including the previously failing OpenCode runner shard. - Merge validation on rebased head `c509d79dd190c5cb00dc65edfde209097ff21465`: all five commits are patch-identical to the reviewed branch. All 54 checks passed with two conditional skips, and fresh Greptile review is 5/5 with no findings. One retry cleared an npm archive 404 and a Cursor fixture timeout. ## Risks - New behavior is opt-in. Existing task snapshot behavior retains its defaults. - Generic strict merge preflight remains opt-in. The dependent agent-folder feature rebases changed paths before applying them to provide per-file last-sync-wins; it does not create a conflict-review queue. - This change adds no database migration, dependency, or UI. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and tool use assisted this change. Live provider validation used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b06034d762 |
Write mode-constrained inbound files directly to their target (#12320)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter utilities transfer files between the host and an agent environment > - A sandbox target already provides the security boundary for inbound files > - The generic fallback adds a temporary file and a rename that do not add protection inside that boundary > - This pull request writes a mode-constrained inbound file directly to its target and applies the mode after the write > - The benefit is a simpler transfer path while host targets keep the strict pre-write mode rule ## Linked Issues or Issue Description **What existing behavior does this improve?** The inbound file-sync fallback for a mode-constrained file stages the file under a temporary name, applies the mode, and renames the file into place. **Subsystem affected** `packages/adapter-utils` and `packages/plugins`. **Current behavior** A sandbox target uses a temporary path before it receives the file. The host then changes the mode and renames the file to the target path. **Proposed behavior** A sandbox target receives the file at its target path. The host applies the mode after the write. A host target still applies the mode before the first byte. **Reason and benefit** The sandbox boundary already protects the target. The direct write removes an unnecessary staging path and rename. **Breaking changes** None. The directory path and outbound transfer path keep their existing behavior. ## What Changed - Write a mode-constrained single-file inbound transfer directly to the sandbox target. - Apply the mode after the direct write and keep the confinement check before post-upload commands. - Scope the protocol comment by transfer direction and preserve the strict host-target rule. - Keep directory inbound transfers and outbound transfers unchanged. ## Verification - Run the targeted unit suite for the changed package. - Verify the suite covers direct target writes, post-write mode application, and confinement rejection. - Run `tsc --noEmit` for both changed packages. - Review the full GitHub Actions check set after the PR opens. ## Risks - A sandbox provider that assumes a temporary inbound path could expose a behavior mismatch. - The confinement check remains before post-upload commands, which limits escape risk. - Host targets keep the pre-write mode rule, so host permission behavior does not change. ## Model Used OpenAI GPT-5. This model assisted with Git operations, PR preparation, review coordination, and tool use. Context window size and reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
445547c989 |
feat(duplex): run the Daytona sandbox callback bridge over Node HTTP/2 (#12120)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers carry agent work through controlled execution channels > - The Daytona callback bridge uses a bespoke line-framed protocol over its duplex channel > - The bespoke protocol adds framing work and does not use the Node transport that already supports multiplexed streams > - This pull request carries raw bytes across the channel, adds a Node HTTP/2 bridge, and selects it for Daytona > - The benefit is one authenticated, multiplexed callback session with queue_v1 as the bounded fallback ## Linked Issues or Issue Description **Subsystem affected** The packages/plugins Daytona provider and the shared duplex execution path. **Problem or motivation** The Daytona callback bridge uses a bespoke line-framed protocol over the provider duplex channel. This adds protocol work and limits stream handling. **Proposed solution** Carry raw bytes through the cross-layer channel. Add an authenticated Node HTTP/2 host server and sandbox client gateway. Select http2_v1 for Daytona and retain queue_v1 as the fallback. **Alternatives considered** Keep the current duplex_v1 protocol. This keeps the bespoke framing path and does not provide one HTTP/2 session for callback streams. **Roadmap alignment** ROADMAP.md lists Daytona under cloud and sandbox agents. This change improves the shipped Daytona provider path. **Additional context** The branch adds no dependency. Node 24 provides the http2 module. The host token check and canonical path parser remain the single dispatch path. ## What Changed - Carry raw Uint8Array chunks through the adapter, plugin, worker, runtime, and Daytona layers. - Encode bytes as base64 only across the JSON-RPC hop, because JSON has no binary type. - Add the bounded host HTTP/2 server and the in-sandbox HTTP/2 client gateway. - Authenticate every stream with the per-run bridge token before route work. - Parse the path once and reuse the canonical result for route and forwarding work. - Select http2_v1 for Daytona and fall back once to queue_v1 when the client preface is absent. - Add transport, session, stream, and fallback telemetry. - Mark HTTP/2 as the preferred transport and queue_v1 as the soft-deprecated fallback. ## Verification - `npx vitest run packages/adapter-utils/src` — 990 passed and 4 skipped. - `npx vitest run server/src/__tests__/plugin-worker-manager-duplex.test.ts` — 32 passed. - `npx vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 220 passed and 6 skipped. - `npx tsc --noEmit` in `packages/adapter-utils`, `packages/shared`, `packages/plugins/sdk`, and `server` — clean. - No `package.json` or `pnpm-lock.yaml` file changed. - The live Daytona test skips when `DAYTONA_API_KEY` is absent. - The root `npx tsc --noEmit` command has a pre-existing missing `packages/adapters/droid-local` reference on this branch and on `master`. ## Risks - The transport change affects several duplex layers and could expose byte-boundary errors. - A missing HTTP/2 client preface falls back once to queue_v1 and records `preface_missing`. - The host token check and canonical path parser must remain on the shared dispatch path. - The live Daytona test needs `DAYTONA_API_KEY` and does not run in this agent sandbox. ## Model Used OpenAI GPT-5, tool-enabled coding agent with repository inspection, GitHub CLI, and shell execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c5050396c7 |
fix(daytona-duplex): chunk host-to-sandbox writes and make a transport close legible (#11986)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agent work through adapters and sandbox providers > - The Daytona duplex path sends host input through a provider pseudo-terminal WebSocket > - Large messages exceed the provider limit, and a transport close can look like a process exit > - This pull request chunks UTF-8 input and carries transport-close state through the duplex path > - The benefit is reliable large input and accurate loss reporting ## Linked Issues or Issue Description **What happened?** The Daytona duplex path sent a full input payload as one WebSocket message. A payload above the provider limit closed the channel. The wait path also mapped a non-numeric exit result to a process exit without exit data. **Expected behavior** The provider must receive large input as ordered UTF-8 chunks. A transport close without exit data must record `transport_closed`, while a numeric exit must record `provider_exit`. **Steps to reproduce** 1. Start a Daytona duplex session. 2. Send an input payload larger than 65536 bytes. 3. Observe that one message closes the provider channel. 4. End a session without a numeric exit code. 5. Observe that the loss reason reports a process exit. **Paperclip version or commit** Commit `1761e79ec9097c65d94f90a8ba20416f8ab718a6`. **Deployment mode** Built from source with the Daytona sandbox provider. ## What Changed - Add a shared UTF-8 byte chunker with a 32768-byte cap. - Route both Daytona pseudo-terminal write paths through the chunker. - Preserve multi-byte UTF-8 sequences across read-side chunks. - Carry an explicit `transportClosed` state through the worker and host wait paths. - Record `transport_closed` for a reason-less transport close and `provider_exit` for a numeric exit. - Keep orderly completion suppression for both exit paths. ## Verification - The Daytona plugin suite passes 194 tests. - The adapter-utils broker, codec, and telemetry suites pass 73 tests. - The plugin SDK duplex and worker RPC host suites pass 37 tests. - The server plugin worker manager duplex suite passes 78 tests. - The execution target sandbox and ACPX execute suites pass 257 tests. - TypeScript checks pass for adapter-utils, plugin SDK, server, and the standalone Daytona plugin. ## Risks The chunk size adds a loop for large input payloads. The 32768-byte cap stays below the provider limit. The optional loss field preserves compatibility for other providers. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faab2620ad |
feat(sandbox): add a duplex transport for Daytona behind a default-off kill switch (#11750)
## Thinking Path > - Paperclip is an open source app that manages AI agents for work > - Paperclip runs agents in local and remote sandbox environments > - A sandbox needs a bounded channel for commands and asynchronous input > - Daytona needs a real pseudo-terminal transport for this channel > - The sandbox gateway also needs a mode that handles channel loss safely > - This pull request adds the Daytona transport and gateway mode behind a default-off kill switch > - The benefit is a tested foundation for later transport selection ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): sandbox providers, plugin SDK, server settings, and shared types. **Problem or motivation** The merged sandbox protocol has no runtime transport for Daytona. The generated sandbox gateway also has no duplex mode. A later transport-selection change needs both parts and a safe per-run gate. **Proposed solution** Add a Daytona `duplexCommandStream` transport over a raw pseudo-terminal. Add a generated gateway mode named `duplex_v1`. Add the `enableSandboxDuplexBridge` setting with a default value of `false`. Keep transport selection disabled until a later pull request. **Alternatives considered** Keep the protocol unused until the transport-selection change. This would delay provider tests and leave the gateway path without direct coverage. **Roadmap alignment** This change supports the completed Roadmap item for cloud and sandbox agents. It extends the merged sandbox channel foundation in pull request #11738. **Additional context** The Daytona provider remains an untrusted boundary. Deployments must use least-privilege provider credentials and provider-side quota controls. Operators must name an owner for duplex telemetry retention before rollout. ## What Changed - Add the Daytona `duplexCommandStream` capability over a raw pseudo-terminal. - Add a launch wrapper that disables echo and newline translation for NDJSON frames. - Close channels on lease release, destroy, resume of a stopped worker, and worker shutdown. - Declare the capability in the Daytona manifest and set `PLUGIN_VERSION` to `0.1.5`. - Add the worker-to-host notification sink at `ctx.duplexChannel.data` and `ctx.duplexChannel.exit`. - Add the generated sandbox gateway mode `PAPERCLIP_API_BRIDGE_MODE=duplex_v1`. - Add channel-loss results of `409 outcome_indeterminate` and `503 bridge_unavailable`. - Add the per-run setting `enableSandboxDuplexBridge`, with a default value of `false`. - Add unit tests, generated-source codec tests, lifecycle tests, and a credential-gated live Daytona test. ## Verification - Daytona suite: 185 tests pass. - Adapter utilities: 754 tests pass and 4 tests skip. - Plugin SDK: 62 tests pass. - Shared package: 28 tests pass. - Server duplex tests pass. - Shared, plugin SDK, server, and Daytona TypeScript checks pass. - The live Daytona test passes 3 cases when `DAYTONA_API_KEY` is set. - The live Daytona test skips 3 cases without `DAYTONA_API_KEY`. - CI must run the full workspace typecheck, test, and build gates after PR creation. ## Risks - The Daytona control plane and pseudo-terminal remain untrusted boundaries. - The duplex gateway changes behavior only when the mode and per-run setting enable it. - A lost channel fails requests without replay, so callers must handle indeterminate outcomes. - The transport-selection change must require both `duplexCommandStream === true` and `enableSandboxDuplexBridge === true`. - The provider credential and quota limits need operator control before rollout. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8161244284 |
feat(sandbox): add opt-in duplex command-stream foundation (capability, protocol, bounded host route, frame codec) (#11738)
## Thinking Path > - Paperclip provides a control plane for companies that run AI agents. > - Sandboxed agents need a safe execution path for persistent command streams. > - The existing callback transport does not provide a bounded, generic duplex route. > - The host must control capability access, route identity, protocol limits, and close behavior. > - This pull request adds an opt-in duplex command-stream foundation across the sandbox layers. > - The feature stays inert because no current provider declares the capability. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox command execution needs a persistent host-to-sandbox stream. The current callback bridge uses a file transport and does not provide this generic route. **Proposed solution** Add a fail-closed provider capability, generic worker protocol messages, a host-owned bounded route, cross-layer service mediation, and a versioned newline-delimited frame codec. **Alternatives considered** Keep the file transport and add feature-specific commands. This does not provide one reusable duplex contract or host-owned route bounds. **Roadmap alignment** This work supports the completed Cloud / Sandbox agents roadmap area and the safe autonomy goal in the product definition. **Additional context** The change passed a two-stage security review. The final code review verdict was approve after fixes for active-stream bounds and service-layer capability mediation. ## What Changed - Add the opt-in `duplexCommandStream` provider capability with fail-closed narrowing. - Add duplex open, write, stop, and close requests and data and exit notifications to the plugin worker protocol. - Add a host-owned route with bounds for chunk size, cumulative bytes, lifetime, protocol errors, pending requests, and pre-bind buffering. - Add close acknowledgement handling with worker retirement when the close remains unconfirmed. - Wire `openDuplexChannel` through the execution target, runtime service, and plugin worker. - Add a versioned frame codec with shared wire-compatibility vectors and split UTF-8 handling. ## Verification - `server/src/__tests__/plugin-worker-manager-duplex.test.ts` passes 18 tests. - `server/src/__tests__/environment-execution-target-duplex.test.ts` passes 11 tests. - `packages/adapter-utils/src/duplex-frame-codec.test.ts` passes 38 tests. - `server/src/__tests__/sandbox-capability-contract.test.ts` passes 15 tests. - Setup-token pseudo-terminal regression tests pass 47 tests. - Server TypeScript check passes. - Continuous integration will run the full required test, typecheck, build, and policy checks. ## Risks - Providers that opt into the capability must implement the complete worker protocol. - Route limit defaults can close a stream when a workload exceeds the configured bounds. - The capability remains disabled for current providers, so current production behavior does not change. ## Model Used OpenAI GPT-5 (`gpt-5`), with tool use and code execution. The model reviewed and prepared this pull request from the supplied implementation and verification record. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
233be4b36c |
feat: parallelize sandbox file-sync behind a provider opt-in capability (#11736)
## Thinking Path > - Paperclip runs AI agents through local and remote execution adapters. > - Sandbox providers move workspace and asset files before and after agent runs. > - Serial file transfers delay startup and teardown when several operations do not depend on each other. > - Providers need an opt-in contract so existing providers keep their serial behavior. > - This pull request adds a bounded scheduler and routes inbound and outbound sync operations through it. > - The benefit is shorter sandbox setup and teardown with stable errors, clear telemetry, and a safe opt-in path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): packages/shared, packages/adapter-utils, packages/plugins, and server. **Problem or motivation** Sandbox sync processes the workspace, assets, and referenced projects in series. This adds avoidable wait time to agent startup and teardown. **Proposed solution** Add a fail-closed provider capability named concurrentSyncOperations. Use a bounded scheduler with a limit of four operations. Preserve operation order for error reporting. Keep non-opted-in providers on the serial path. **Alternatives considered** Increase the serial transfer speed or add provider-specific schedulers. Those options do not provide one shared contract or stable behavior across providers. **Roadmap alignment** ROADMAP.md lists cloud and sandbox agents as a product area. This change improves sandbox execution without changing the control-plane contract. **Additional context** The Daytona provider opts in. Board trials on this commit showed overlap for inbound sync and outbound restore, with no referenced-project staging failures. ## What Changed - Add the concurrentSyncOperations sandbox capability and fail-closed parsing. - Add a bounded settle-all scheduler with stable input-order errors. - Parallelize inbound workspace, asset, and referenced-project sync operations when the provider opts in. - Parallelize outbound workspace and asset restore operations when the provider opts in. - Surface referenced-project failure text in run logs and server telemetry. - Add Daytona sync spans and the capability declaration. - Preserve in-flight upload scratch tarballs during workspace wipe. - Add unit and regression tests for the scheduler, coordinators, provider behavior, telemetry, and wipe race. ## Verification - Run the adapter-utils and server type checks. - Run the targeted adapter-utils, server, and Daytona test suites. - Run the full automated sweep. - Review six cold Daytona trials, with three serial and three parallel runs. - Confirm that parallel trials show inbound overlap and outbound restore overlap. - Confirm that providers without the capability keep serial behavior. ## Risks - Providers must opt in only when their file operations can run safely at the same time. - A provider that declares the capability incorrectly can expose transfer races. - The scheduler keeps a limit of four to bound resource use. - Providers without the capability keep the prior serial behavior. ## Model Used OpenAI GPT-5 in the Codex runtime. The model used tool calls, code inspection, and GitHub workflow support. The model did not author the implementation commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d5208d30c1 |
feat(adapter-utils): add host-side pack span to managed-runtime tarball build (#11072)
## Thinking Path > - Paperclip runs AI agents through adapter execution services. > - The adapter runtime records spans for each stage of agent startup. > - The managed runtime packs workspace tarballs before it uploads them. > - These pack operations had no host span, so `stage.sync` omitted pack time. > - This pull request adds one `pack` span around both tarball builds. > - The span nests under the active `stage.sync` step and improves trace detail. ## Linked Issues or Issue Description **What existing behavior does this improve?** The managed runtime workspace sync builds a git-history tarball and a workspace-overlay tarball before upload. **Subsystem affected** The change affects `packages/adapter-utils`, which provides adapter execution and managed runtime support. **Current behavior** The host builds both tarballs without an OpenTelemetry span. The `stage.sync` trace therefore omits the host pack duration. **Proposed behavior** The host wraps both tarball builds in one `pack` span. The executor parents this span under the active startup step. **Reason and benefit** The trace shows the time that the host spends packing workspace data. Operators can use the existing runtime span tree to find sync delays. **Breaking changes** None. The default span runner remains a no-op runner, and the existing control flow remains unchanged. ## What Changed - Add an optional `runtimeSpan` runner to the managed runtime preparation path. - Create one host `pack` span around the two workspace tarball builds. - Parent the `pack` span under the active startup step. - Add unit coverage for span emission and span nesting. ## Verification - Run `tsc --noEmit` for `@paperclipai/adapter-utils`. - Run the `@paperclipai/adapter-utils` Vitest suite. - Confirm the suite reports 448 passed tests and 4 skipped tests. - Confirm the trace test records `pack` under `stage.sync`. ## Risks The change adds optional tracing only. The default no-op runner preserves behavior when tracing is not configured. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The model reviewed the handoff and opened this pull request. The implementation author supplied the commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b7e0814a0 |
feat(acp): stream Daytona sandbox agent output and remove the host output poll (#11049)
## Thinking Path > - Paperclip is the open source app that manages AI agents for work > - Sandbox providers let agents run in remote and isolated environments > - Daytona session commands need a path that sends agent output to the host without host polling > - Host polling adds delay and repeats provider output work > - This pull request adds typed execute.log notifications and a log sink for incremental output > - This pull request adds an optional ACP session stream with final-result replay protection > - The benefit is lower output delay while the default flags keep current behavior unchanged ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. The change spans the plugin SDK, Daytona provider, adapter utilities, and server execution services. **Problem or motivation** The Daytona ACP bridge polls a host output file while an agent command runs. This adds delay and can repeat work. The host also needs a safe route for provider output chunks. **Proposed solution** Add a typed `execute.log` notification with host-issued invocation correlation. Add an ordered log sink to the environment execute path. Add an optional ACP session-log path that parses newline-delimited JSON frames and removes the host output poll for that path. **Alternatives considered** Keep the output-file poll as the only path. This keeps the current behavior but does not provide timely output. The new path stays behind flags, so the existing path remains the default fallback. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`, including Daytona support. ## What Changed - Add the typed `execute.log` worker-to-host notification and company-scoped host route. - Add ordered `stdout` and `stderr` chunk delivery before the final execute result. - Add the Daytona session log sink and the optional ACP streamed session path. - Add monotonic frame handling so live and final output reach the host once. - Keep `useLogStream` and `streamAgentSessionOutput` off by default. - Add unit and integration coverage for the notification, execution target, runtime, and Daytona paths. ## Verification - Run adapter-utils tests: 445 tests pass locally. - Run server environment tests: 73 tests pass locally. - Run Daytona plugin tests: 131 tests pass locally. - Run TypeScript checks for shared, adapter-utils, and server. - Review the pull request checks after GitHub completes them. - All required GitHub checks pass on the current head. ## Risks The new paths change output delivery only when a feature flag enables them. The final execute result remains available for parsing and fallback. The main risk is a provider stream or frame-order error; the final-result parser limits that risk. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The runtime did not supply a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes No operator documentation change applies because both new flags remain disabled by default. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a69a3d6b0 |
refactor(observability): stop writing detailed per-step timing to the run log (#10776)
## Thinking Path > - Paperclip tracks agent work for the team. > - The startup timing path already emits detailed spans. > - The run log repeats the same detailed timing. > - That makes the event larger than it needs to be. > - This pull request removes the redundant per-step timing fields from the run-log payload. > - The benefit is a smaller log with the same detail still available in traces. ## Linked Issues or Issue Description No public GitHub issue or open pull request matches this change. I described the enhancement below. **What existing behavior does this improve?** `run.startup.step` event output. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `run.startup.step` writes `roundTrips`, `providerExecMs`, `providerGetMs`, `createRuntimeMs`, and `ensureSessionMs`. The same detail already exists in the spans. **Proposed behavior** `run.startup.step` keeps only `step`, `durationMs`, and `outcome`. The heartbeat lifecycle timestamps stay unchanged. **Reason and benefit** This change removes redundant data from the run log. It keeps the useful detail in trace spans. It also makes the payload smaller and easier to read. The revert path stays clear because the removed data has one producer chain. **Breaking changes** Yes. Consumers that read the removed fields must switch to span data or the remaining event fields. The heartbeat lifecycle timestamps do not change. **Additional context** I searched GitHub for related open PRs and issues. I found no match. ## What Changed - Removed the redundant per-step timing fields from the run-log payload. - Removed the now-dead producer chain that fed those fields. - Kept the span-level timing data and the heartbeat lifecycle timestamps. ## Verification - `adapter-utils` typecheck clean - `server` typecheck clean - `startup-timing` suite: 30 passed - `adapter-utils` `execute` suite: 99 passed - `server` `environment-execution-target` suite: 21 passed ## Risks - Consumers that still read the removed fields will need a code change. - This is a one-way-door data removal for the run log. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a267e0328 |
feat(sandbox): stage referenced projects into the run sandbox (#10469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs can reference more than one project > - Each referenced project must land in its own sandbox tree so one run does not mix files across projects > - The anchor workspace must keep its own git history and overlay rules > - Fail closed on sync and confinement errors, and keep the other projects alive > - This pull request stages referenced projects into isolated project directories under the run sandbox root > - The benefit is safer multi-project runs with clear failure isolation ## Linked Issues or Issue Description No public issue exists. Related PR: #10448. ## What Changed - Thread additional referenced-project sources through the run prepare path and the runtime layers. - Stage each referenced project into its own `project-<projectId>` directory under the runtime root. - Keep the anchor workspace history and overlay semantics unchanged. - Fail closed on confinement or sync errors for one project, and keep the other projects running. - Keep the path inert by default behind the multi-project workspace-sync kill-switch. - No documentation update was needed for this code-only runtime change. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/sandbox-file-sync.test.ts src/command-managed-runtime.test.ts src/sandbox-managed-runtime.test.ts src/remote-managed-runtime.test.ts` - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` - The branch points at `871378e4064a95adc8e4647e442ec1359768de61` on `origin/feat/stage-referenced-projects-into-sandbox`. - GitHub CI is green on PR #10469. ## Risks - A referenced project can skip if confinement or sync fails. - A new runtime tree layout can affect tools that assume one project root. - The kill-switch keeps the path inert until operators enable it. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes or confirmed no docs update was needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7083c275c8 |
refactor(sandbox): retire the dead noProfile flag from the exec path (#10461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The sandbox exec path starts agent commands and passes runtime options to the server and plugin layers > - This path kept a noProfile flag after the exec wrappers stopped sourcing a login profile > - The flag no longer changed behavior, so it left dead API surface in the protocol and runtime helpers > - This pull request removes that dead flag from the plugin protocol, the server drivers, and the managed-runtime helpers > - It also updates the tests and points the agent runtime README at the sandbox requirements file > - The benefit is a smaller and clearer exec-path contract with no behavior change ## Linked Issues or Issue Description - No public GitHub issue exists. ### What happened? The sandbox exec path kept a `noProfile` field after the exec wrappers stopped sourcing a login profile. ### Expected behavior The plugin protocol, server drivers, and managed-runtime helpers should not expose or forward a dead field. ### Steps to reproduce 1. Run a managed-runtime command through the sandbox exec path. 2. Inspect the protocol payload and runtime helper inputs. 3. Observe that `noProfile` is present even though it no longer changes behavior. ### Paperclip version or commit `60c7da86fc7a6c1dbf37bbcd86e25ecaaff01607` ### Deployment mode Built from source (pnpm dev / pnpm build) ### Additional context This pull request removes the dead field, updates the affected tests, and updates the README note for the sandbox profile path. ## What Changed - Removed noProfile from the plugin protocol and the server exec-path call sites. - Updated the managed-runtime helpers to use the narrower exec-path contract. - Updated the affected tests and added the README pointer to SANDBOX-REQUIREMENTS.md. ## Verification - `git grep -n "noProfile" -- packages/ server/` returns zero matches. - `tsc --noEmit` passed for `@paperclipai/adapter-utils`, `@paperclipai/plugin-sdk`, and `@paperclipai/server`. - `command-managed-runtime.test.ts` passed: 22/22. - `environment-runtime.test.ts` passed: 24/24. ## Risks - Low risk. The flag was already a no-op. - A hidden external caller may still send the removed field. ## Model Used - OpenAI GPT-5, tool-enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7797995038 |
perf(plugin-daytona): opt-in no-profile fast path for default-PATH execs (#10352)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - The Daytona adapter turns tasks into shell commands and manages execution overhead > - Many short-lived exec calls still pay for login-shell profile sourcing even when the binary already resolves on the sandbox default PATH > - That extra startup work adds latency on the hot path for repeated command execution > - This pull request adds an opt-in fast path that skips profile sourcing only when the caller explicitly requests it and the command does not need shell initialization > - The benefit is lower per-call latency for eligible commands without changing the conservative default behavior for commands that need the profile ## Linked Issues or Issue Description This change does not reference a public GitHub issue. It follows the same Daytona startup-speed work as merged PR #10335 and narrows the execution path for eligible commands while keeping the default login-shell behavior intact. ## What Changed - Added an optional `noProfile` flag to `PluginEnvironmentExecuteParams`. - Refactored Daytona login-shell script assembly so the profile and nvm sourcing block is omitted only on the explicit fast path. - Preserved environment prefixing, `cd`, shell quoting, `NONINTERACTIVE_GIT_ENV`, stdin handling, and `durationMs` behavior on both paths. - Added regression tests for the fast path omission, the preserved execution parameters, and the default profile-sourcing path. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run src/plugin.test.ts` - `pnpm --filter @paperclipai/plugin-sdk tsc --noEmit` - Reverted the guard locally to confirm the two behavior tests fail again, then restored the change. ## Risks - If a caller opts into `noProfile` for a command that depends on shell initialization, the command can fail to resolve its binary. - The API comment and opt-in design keep that risk narrow; the default path remains unchanged. ## Model Used OpenAI GPT-5 (Codex tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
030dd9d15c |
feat(adapter-utils): provider-delegable syncIn seam with ordered post-upload commands (#10340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter/runtime layer has to move files into sandboxes safely and efficiently > - The current sync-in path needs a provider-delegable seam so providers can use their native upload transport when available > - The upload contract also needs ordered post-upload commands so extracted content can be finalized fail-fast after transfer > - The fallback path still has to preserve current behavior when the provider does not expose native sync verbs > - This pull request adds the contract and runtime seam for provider-delegable sync-in, plus the single-stream collapse flag > - The benefit is fewer round trips, a cleaner provider-owned upload path, and a compatible fallback for existing runners ## Linked Issues or Issue Description No public GitHub issue exists for this change. The underlying feature request is described below in the repository's feature-request format. ### Problem or motivation Paperclip needs a sync-in path that lets each provider choose the best available upload transport instead of forcing the harness to orchestrate uploads the same way every time. The runtime also needs a way to describe ordered post-upload commands so providers can finalize extracted content fail-fast after transfer. ### Proposed solution Extend the sync contract with ordered post-upload commands, forward that contract through the plugin and environment runtime layers, and make client syncIn always available. When a provider advertises native sync verbs, the client should delegate to that transport; otherwise it should fall back to the existing tarball/write/extract behavior and then run the post-upload commands in order. ### Alternatives considered Keeping upload orchestration entirely host-side would avoid a contract change, but it would block provider-specific transport optimizations and keep the harness responsible for a path the provider can do more efficiently. A separate post-upload API would add another surface without improving the existing sync flow. ### Roadmap alignment This work aligns with the broader runtime and adapter roadmap because it improves provider integration without changing the external product model. It is an additive contract change that preserves backward compatibility for providers that do not expose native sync verbs. ### Additional context The fallback path still needs to preserve existing observable behavior, including command ordering, cwd confinement, and fail-fast execution. The single-stream progress flag is part of the same transport improvement so smaller writes can collapse to a single round trip when the runner supports it. ## What Changed - Added ordered `postUploadCommands` support to the sync operation contract and SDK mirror. - Plumbed the sync-in contract through the plugin and environment runtime layers. - Implemented a runtime client `syncIn` path that delegates to native provider transport when available, otherwise uses the generic tarball/write/extract fallback. - Preserved fail-fast execution of ordered post-upload commands in the fallback path. - Flipped the sandbox runner's single-stream stdin progress flag to collapse small `writeFile` operations to a single round trip. - Added and updated tests for contract forwarding, fallback behavior, cwd rejection, fail-fast behavior, and single-stream collapse. ## Verification - `pnpm --filter @paperclipai/plugin-sdk exec vitest run protocol.postupload.test.ts` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run environment-sync-negotiation.test.ts` - `pnpm --filter @paperclipai/adapter-utils exec vitest run command-managed-runtime.test.ts` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target.test.ts` - Local typecheck and targeted suite runs reported in the handoff passed before PR creation. ## Risks - The new fallback path could diverge from the previous inline upload behavior if the tarball/extract contract changes. - Provider-native sync handling may expose provider-specific edge cases if a runner advertises sync verbs but does not fully honor the contract. - The single-stream flag changes transport behavior for small uploads, so regressions would likely show up as round-trip or upload failures. ## Model Used OpenAI Codex (GPT-5, tool-using coding agent). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
197718bc00 |
feat(acpx): per-step round-trip + provider-latency attribution for sandbox startup (#10222)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-backed runs need more precise startup observability so operators can see where time is spent before an adapter is invoked > - Aggregate startup timing hides which boundary is actually slow, especially in remote execution where the bottleneck can move between host round-trips, provider-boundary latency, and handshake phases > - This pull request keeps the existing startup timing channel additive while attributing the latency to the specific startup steps that caused it > - The benefit is better diagnosis of sandbox startup regressions without changing control flow or introducing a schema migration ## Linked Issues or Issue Description Refs: #10204 This PR extends the existing sandbox run-startup timing observability with per-step round-trip and provider-latency attribution for the Daytona startup path. It keeps the event payload additive and free-form, and it leaves the control flow, database schema, and external adapter interfaces unchanged. ## What Changed - Added per-step round-trip counting for the host-to-sandbox execute seam - Added provider-boundary duration accumulation for the Daytona execute and re-fetch steps - Split the ACP handshake timing into `createRuntimeMs` and `ensureSessionMs` while preserving the warm-handle skip - Kept the startup timing payload additive and did not add a schema migration ## Verification - `adapter-utils` acpx-engine and startup-timing suites: pass - Daytona plugin suite: pass with mocked SDK and injected-clock duration assertions - `server` environment-execution-target suite: pass - `tsc --noEmit` for adapter-utils and server: pass - Git validation: fetched `origin/feat/sandbox-start-step-timing-attribution`, confirmed it matches the authorized submit SHA `53c573266e618af05c57a0720aaa9d9e0452de61`, and confirmed the branch contains only the expected commit on top of `origin/master` - Searched GitHub for duplicate or related PRs/issues; found one closely related merged PR and no open duplicate on this branch - Checked `ROADMAP.md`; the broad sandbox-agent roadmap section does not call out this specific startup-timing attribution work as a duplicate ## Risks - Low risk: the change is additive and only enriches existing timing data - Downstream consumers that assume aggregate-only startup timing may need to tolerate the additional per-step fields - The finer spawn/initialize/session split remains a follow-up in the external ACP client because that hook is not available here yet ## Model Used OpenAI GPT-5, tool-using coding agent ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e00818574 |
fix(runtime): support in-place workspace realization (#10230)
## Thinking Path > - Paperclip is the control plane people use to coordinate AI agents and their execution environments. > - Environment realization decides where an agent runs and which filesystem and toolchain are authoritative. > - Copy-based realization is unsafe for container-anchored tasks because absolute paths such as `/app` can point outside the synchronized tree and task-specific binaries may be absent. > - That mismatch can let an agent successfully verify work in a phantom writable path while sync-back silently discards the result. > - Existing task environments already provide the authoritative filesystem and toolchain, so they should be executed in place rather than copied. > - Copy mode still needs explicit confinement rules so aliases target the synchronized workspace and unsynchronized writable paths fail visibly. > - This pull request adds typed realization metadata, propagates the authoritative root through orchestration, and teaches Codex to honor it. > - The benefit is that container-anchored tasks operate on verifier-visible state with the intended tools, while copy mode remains safe and backward compatible. ## Linked Issues or Issue Description No public GitHub issue exists for this defect. GitHub duplicate searches for in-place execution, workspace realization, and authoritative workspace roots found no related pull request to link. ### What happened? Environment-backed agent runs were always realized through a copied workspace. Tasks anchored to absolute container paths could therefore write outside the synchronized tree, and task-provided toolchains were unavailable in the copy. A run could report success even though sync-back discarded its output. ### Expected behavior Existing task environments should run against their real authoritative root and toolchain. Copy-mode runs should map declared absolute aliases into the synchronized tree and reject writable paths that cannot be restored. ### Steps to reproduce 1. Run a Codex task environment whose required files live under `/app` or `/workspace` and whose required binary exists only in the task container. 2. Observe that copy realization changes the effective filesystem/toolchain or permits writes outside the synchronized root. 3. Complete and verify the task inside the agent sandbox. 4. Observe that the verifier cannot see out-of-tree artifacts or that task-specific commands were unavailable. ### Reproduction context - Paperclip commit: `3a16b91217483d2c233926de5b7f7bc3a1077924` - Deployment: built from source in a task-container execution environment - Adapter: Codex local - Database: not database-related - Access context: agent execution ## What Changed - Added typed `copy | in_place` workspace-realization metadata, authoritative roots, confined aliases, and outbound restore paths to shared execution-target contracts. - Selected in-place realization for existing task environments and skipped archive prepare/restore when the authoritative environment is used directly. - Propagated the authoritative root into adapter context so Codex uses it for cwd and `PAPERCLIP_WORKSPACE_*` semantics, including ACP execution. - Bound copy-mode aliases such as `/app` to the synchronized workspace and rejected writable out-of-tree paths without explicit restore mappings. - Added focused regression coverage while preserving existing copy-mode archive restore behavior. ## Verification - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/execute.remote.test.ts server/src/__tests__/environment-run-orchestrator.test.ts` — 48 passed, 4 skipped. - `pnpm -r typecheck` — passed. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY -u AWS_SESSION_TOKEN pnpm test:run` — passed across all general and serialized Vitest shards. - `pnpm build` — passed. - Codex `k=1` acceptance run completed July 24, 2026 at 23:54:30 UTC with 4 completed, 0 exceptions, and mean reward 1.0: `build-cython-ext`, `openssl-selfsigned-cert`, `prove-plus-comm`, and `sqlite-db-truncate` each received terminal grade 1.0 against real task-environment paths and toolchains. ## Risks - In-place mode deliberately exposes the authoritative task root to the adapter; incorrect environment metadata could point execution at the wrong root. Typed metadata and focused orchestration tests cover selection and propagation. - Copy-mode writable-path validation is stricter and may reject previously accepted unsafe configurations. The rejection is intentional and produces a visible error instead of silently losing output. - The acceptance run is focused on four Codex task-environment workloads, not a broad cross-adapter benchmark. Existing copy-mode archive tests and the full repository suite remain green. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact model ID and context-window size were not exposed to this runtime. Capabilities used: extended reasoning, repository editing, shell execution, test/build execution, Git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b247bf7150 |
feat(runtime): opt-in sandbox file-sync lifecycle hooks (API + provider docs) (#10013)
## Thinking Path > - Paperclip is an open-source AI-agent management platform; agents run tasks inside sandboxed environments (Daytona, Kubernetes, E2B, etc.) > - The control-plane ↔ sandbox file-transfer path flows through the `environmentExecute` seam in `protocol.ts` — the only verb available to plugins — which forces a base64-over-exec chunked loop for every file move: workspace files, assets, Codex home sync > - This transport is correct and safe, but it bypasses provider-native bulk/streaming APIs (Daytona `uploadFiles`, K8s `FastUploadInterceptor` / volume mounts), leaving significant throughput on the table for large workspaces > - The right fix is an opt-in seam extension: providers with faster native transfer declare two optional verbs; providers that do not opt in stay on the existing fallback with zero code or behavior change required > - This PR adds the first layer of that extension — two optional verbs (`environmentSyncIn` / `environmentSyncOut`) in the plugin SDK, the runtime plumbing to prefer the native path for the two clean destroy-then-replace cases, and a doc for the contract > - The core correctness invariant is byte-identical fallback: if no provider opts in, execution is exactly what ships today; `assertSyncOperationsConfined` enforces host-side path confinement for providers that do opt in > - No provider advertises the verbs yet → zero production behavior change; future PRs wire up Daytona and K8s providers against this contract ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Description follows the `feature_request` issue template: **Subsystem affected:** packages/plugins — plugin system; packages/adapter-utils — adapter runtime; server/ — EnvironmentRuntimeService **Problem or motivation:** Sandbox file transfers currently always use a base64-over-exec chunked loop regardless of what the underlying provider supports. For workspaces larger than a few MB this becomes the dominant wall-clock cost of every sandbox run, and it bypasses bulk/stream APIs that providers like Daytona already expose natively. **Proposed solution:** Add two optional, opt-in plugin hooks — `onEnvironmentSyncIn` / `onEnvironmentSyncOut` — to the plugin SDK. When a provider defines both hooks and both are advertised via the existing `supportedMethods` negotiation, the runtime prefers the native path for the two clean destroy-then-replace transfer cases; all other cases fall back to the existing byte-identical base64 transport. **Alternatives considered:** An unconditional verb would require every provider to implement or stub the verb. The opt-in / `METHOD_NOT_IMPLEMENTED` pattern (already used by `environmentExecute`) preserves backward compatibility with zero provider changes required. **Roadmap alignment:** Consistent with the ✅ "Cloud / Sandbox agents" and ✅ "Plugin system" milestones; extends the plugin seam rather than adding control-plane-level logic. **Additional context:** Searched open pull requests and issues for duplicate sandbox file-sync / native-transfer work; none found. ## What Changed - **`packages/plugins/sdk`** - `protocol.ts`: two new optional `HostToWorkerMethods` — `environmentSyncIn` / `environmentSyncOut` — plus generic `SyncOperation`, `SyncFileMapping`, and `SyncOutcome` types - `define-plugin.ts`: optional `onEnvironmentSyncIn` / `onEnvironmentSyncOut` fields on `PluginDefinition`; worker advertises each verb only when its hook is defined (else `METHOD_NOT_IMPLEMENTED`, mirroring `environmentExecute`) - `worker-rpc-host.ts`: route new verbs to plugin hooks - `index.ts`: re-export new public types - **`packages/adapter-utils`** - `command-managed-runtime.ts`: expose optional `syncIn` / `syncOut` on `CommandManagedRuntimeRunner` (available only when both verbs are advertised); add `assertSyncOperationsConfined` host-side path-confinement guard - `sandbox-managed-runtime.ts`: `SandboxManagedRuntimeClient` gains optional `syncIn` / `syncOut`; orchestrator prefers native path for default-provision asset inbound and workspace-download-into-fresh-dir outbound; all other paths keep the existing base64 fallback - `sandbox-file-sync.test.ts` (new): 234-line characterization suite — native-opt-in branch, fallback branch, `assertSyncOperationsConfined` escape-path rejection, `followSymlinks` → tar `-h` - `command-managed-runtime.test.ts`: negotiation + native-sync + confinement tests - **`server/src/services/environment-runtime.ts`**: `EnvironmentRuntimeService` delegates to `syncIn` / `syncOut`, gated on advertised support - **`server/src/services/environment-execution-target.ts`**: minor typing fix alongside the new verbs - **`doc/plugins/SANDBOX_FILE_SYNC_HOOKS.md`** (new): documents the full contract — opt-in / no-op guarantee, operation ordering, provider-may-tar, atomicity, `followSymlinks`, secret modes (0600, no window), path confinement, `operationId` opacity, resource bounds, shell-quoting ## Verification ```bash # SDK suite pnpm --filter packages/plugins/sdk test # Adapter-utils suite (includes new sandbox-file-sync characterization tests) pnpm --filter packages/adapter-utils test # Expected: 255 pass / 4 skip # Type-check across affected packages pnpm --filter packages/plugins/sdk typecheck pnpm --filter packages/adapter-utils typecheck # Server changed-file spot check: cd server && npx tsc --noEmit --skipLibCheck 2>&1 | grep -E "environment-(runtime|execution-target)" | head -20 ``` Key behavioral invariant to spot-check: with no provider opting in (the current state), run any sandbox task and confirm file-transfer behavior is byte-for-byte identical to what the pre-PR code produces. The characterization tests assert this at the unit level. ## Risks - **Zero production risk today**: no provider advertises `environmentSyncIn` / `environmentSyncOut`, so the new code paths are unreachable in production; all real traffic stays on the existing base64 fallback - **Path confinement**: `assertSyncOperationsConfined` rejects any `targetPath` that escapes the declared root — this is the primary security boundary for future providers. The test suite covers escape-path rejection - **Atomicity**: the contract delegates atomicity to providers; the doc explicitly calls out that directory-level ops are not guaranteed atomic - **Secret transport**: credential assets (e.g., Codex `auth.json`, directory mappings) continue to use the existing tar path — they do not go through the new verbs in any current provider > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: Anthropic Model: `claude-sonnet-4-6` (Claude Sonnet 4.6) Context window: 200 K tokens Capabilities: extended tool use, multi-file code generation, agentic reasoning via the Paperclip agent framework ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bac7307ec4 |
fix(adapter-utils): improve sandbox restore failure diagnostics (#8903)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters share managed-runtime helpers from `@paperclipai/adapter-utils` so local and sandboxed runs can prepare, execute, and restore workspaces consistently. > - The sandbox managed runtime syncs workspaces in both directions by creating tar archives on the local host and inside the remote sandbox. > - A prior fix avoided archiving `.` during upload because tar self-entries can force chmod/utime on a directory the command user does not own. > - The restore/download path still archived the remote workspace as `.`, and failed managed commands only surfaced stderr. > - This pull request applies entry-based tar creation to the sandbox restore path and includes stdout in managed command failure diagnostics. > - The benefit is that restore failures become visible to operators and sandbox workspace restore avoids the same directory-metadata failure class already fixed for upload. ## Linked Issues or Issue Description No existing public issue found; describing in-PR (bug). - **What happens:** sandbox workspace restore can fail while creating `workspace-download.tar` from the remote workspace when the tar command includes a `.` self-entry and the command user cannot update metadata on the workspace directory. If the failing command writes its diagnostic to stdout, the managed runtime error can collapse to a generic failed shell command without the useful tar message. - **Expected behavior:** restore should archive the workspace entries without a `.` self-entry, and failed managed runtime commands should include useful stdout/stderr diagnostics. - **Where:** `packages/adapter-utils/src/sandbox-managed-runtime.ts` restore/download path and `packages/adapter-utils/src/command-managed-runtime.ts` command error formatting. - **Related public context:** #7836 fixed the upload side of the same tar self-entry failure class. ## What Changed - Added stdout-aware failed-command formatting in the command managed runtime, keeping diagnostics bounded to the tail of stdout/stderr. - Added remote workspace tarball creation that names top-level entries explicitly instead of archiving `.` during sandbox restore. - Preserved empty-workspace restore support by creating a valid empty tarball when the remote workspace has no entries. - Added regression coverage for stdout diagnostics, restore tar members, and empty workspace restore tarballs. ## Verification - `pnpm vitest run packages/adapter-utils/src/command-managed-runtime.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 2 files, 17 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `git diff --check origin/master..HEAD` — passed. - Local public-hygiene/PII scan of the committed diff checked for internal ticket refs, local/private URLs, common token/key patterns, private-key blocks, and email-like values — passed. - GitHub duplicate/related search found no public issue or PR for `workspace-download.tar Permission denied` or `sandbox restore tar permission denied`; #7836 is linked as related prior work. - PR CI on commit `3ad4c8d` — all GitHub Actions lanes, security scans, and aggregate `verify` passed. - Greptile Review on commit `3ad4c8d` — Confidence Score 5/5; the prior P2 thread is resolved with no open P2s, recommendations, or follow-ups. ## Risks Low. This is limited to shared adapter runtime error formatting and sandbox restore archive construction. Archive contents should remain equivalent apart from the removed `.` self-entry, and the new diagnostics are bounded to avoid dumping unbounded command output. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent with shell and GitHub CLI access. Context window size was not reported by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (n/a; internal adapter-runtime behavior only) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a27e5ad002 |
Add ephemeral sandbox runtime status plumbing (#8593)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandboxed agent runs can spend meaningful time preparing a remote workspace before the agent transcript shows useful output. > - Operators need short, current progress text for those setup phases, but that text should not become durable run history. > - The existing live-run websocket path already carries run updates to the UI, so the backend can reuse that channel instead of adding polling. > - This pull request adds an ephemeral runtime-progress contract, a process-local status store, and heartbeat integration for sandbox-managed runs. > - The benefit is a clearer active-run experience without database migrations or persistent progress rows. ## Linked Issues or Issue Description Refs #248 No exact public GitHub issue was found for this status-message plumbing. The underlying problem is that active sandboxed runs currently have setup phases, such as workspace sync and restore, where the operator cannot see concise current progress through the live run state. This PR addresses that gap for the backend/runtime layer while keeping progress messages ephemeral. GitHub search performed for related or duplicate work: `sandbox runtime status`, `sandbox restore index`, and `runtime progress`. No direct duplicate PR was found. ## What Changed - Added shared runtime-progress types and the `heartbeat.run.progress` live event type. - Added a process-local heartbeat run runtime-status store with TTL, bounded/redacted messages, and terminal cleanup. - Threaded runtime progress callbacks through heartbeat execution and active/live run serialization. - Emitted sandbox-managed runtime phase updates for sync, adapter startup, restore/export, and finalization paths. - Added backend and adapter-utils tests for ephemeral status behavior, terminal cleanup, live serialization, and sandbox progress callbacks. ## Verification - `pnpm install --frozen-lockfile` - Local PII scan before push: high-confidence secret patterns, internal issue links, local user paths, and private URL patterns checked across all three split diffs; no real secrets or internal links found. The only secret-like text is an intentional fake test fixture (`sk-test-secret`). - `git diff --check origin/master..feat/sandbox-runtime-status` - `pnpm exec vitest run server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/agent-live-run-routes.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 4 files, 23 tests passed. - `pnpm run typecheck` passed on both top stacks that include this branch: `feat/sandbox-status-ui` and `fix/sandbox-restore-index-sync`. - `pnpm run build` passed on both top stacks that include this branch; Vite reported existing CSS `::highlight` and chunk-size warnings. - `pnpm run test:run` was attempted on `fix/sandbox-restore-index-sync`; it failed in two unrelated broad-suite tests. One depends on this host's Git default branch behavior, and one depends on local Claude model-discovery environment. The changed focused suites above pass. ## Risks - Runtime progress is process-local by design, so status disappears after TTL, terminal cleanup, or server restart. - Clients that do not consume `heartbeat.run.progress` simply keep existing behavior. - Message redaction is intentionally generic; overly specific phase details should stay out of runtime-progress payloads. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent, with shell/tool execution in a local worktree. Exact context-window metadata is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip CTO <cto@paperclip.local> Co-authored-by: Paperclip CTO <noreply@paperclip.ing> |
||
|
|
2c98c8e1e5 |
Fix sandbox git publishing and large workspace uploads (#8422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox and SSH runtimes need to preserve agent work across isolated execution environments > - Git-backed workspaces were being copied mostly as filesystem archives, which breaks when `.git` points outside the mounted workspace and makes sandbox agents unable to publish their own branches > - Large ignored dependency trees could also be swept into the sandbox overlay, causing multi-GB transfers and max-string failures in some sandbox clients > - This pull request makes sandbox runtime setup use a git-backed HEAD sync plus a small dirty/untracked overlay, and bounds sandbox file transfers so large archives do not need one huge string > - The benefit is that sandbox agents can commit and push from a usable git checkout without uploading dependency trees such as `node_modules` ## Linked Issues or Issue Description Refs #8395 No public duplicate issue or PR was found after searches for `sandbox git workspace`, `git push sandbox`, and `node_modules sandbox upload`. Bug report: - What happened: sandbox-backed agent workspaces could receive a `.git` file that pointed at host-only git state, leaving the sandbox unable to run normal git workflows. The sandbox overlay upload could also include ignored dependency directories, creating very large transfers. - Expected behavior: sandbox and remote runtimes should prepare a usable git-backed workspace, copy only the necessary workspace overlay, and restore git history plus file changes without depending on a host-only `.git` path. - Steps to reproduce: 1. Run an agent in a sandbox-backed workspace whose local git checkout is a worktree. 2. Ask the agent to complete a GitHub workflow that requires commit/push access. 3. Observe that git operations can fail inside the sandbox, and ignored dependency trees can be uploaded as part of the workspace overlay. - Paperclip version or commit: reproduced against `master` before this PR, base `7aa212296eb1`. - Deployment mode: local dev / sandbox-backed runtime. - Installation method: built from source. - Agent adapters involved: local adapters using shared adapter-utils runtime preparation. - Database mode: not database-related. - Access context: agent runtime. - Local verification environment: Node.js v25.6.1, pnpm 9.15.4, macOS arm64. - Privacy checklist: all pasted output was reviewed for secrets, private hostnames, local usernames, and internal instance links. ## What Changed - Added a GitHub workflow push preflight so agent runs can detect missing push credentials when a workflow explicitly needs GitHub publishing. - Added shared git workspace sync helpers for shallow HEAD import/export and dirty/untracked overlay tracking. - Updated sandbox managed runtime setup to use git history plus a selected overlay instead of uploading the full local workspace for git-backed workspaces. - Bounded sandbox archive upload/download paths so large payloads stream or chunk instead of materializing one oversized string. - Excluded `.git` and ignored dependency trees from sandbox upload, download, and restore baselines while preserving local ignored directories during sync-back. - Added focused tests for git workspace sync, sandbox overlay selection, transfer chunking, and heartbeat push-preflight behavior. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts` - `pnpm --filter @paperclipai/adapter-utils typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm build` - Public-hygiene scan of the PR diff and commit messages for internal issue ids, local paths, private hostnames, and obvious token patterns. ## Risks - Medium risk: this changes sandbox runtime synchronization semantics for git-backed workspaces, especially around dirty tracked files, untracked files, deleted paths, and ignored files. - The main mitigation is focused test coverage for upload contents, restore exclusions, and git round-trip behavior. - The SSH runtime keeps the current bundle-based implementation from `master`; this PR only aligns shared excludes and sandbox behavior with that model. ## Model Used OpenAI Codex, GPT-5-based coding agent, tool-enabled shell/git/GitHub workflow, with code execution and repository inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
07e98d2b2c |
feat(adapter-utils): add observable sandbox sync progress (#8395)
## Thinking Path > - Paperclip is the control plane for running AI-agent companies, so long-running remote work needs to stay observable to human operators. > - Cloud / sandbox agents are an active roadmap area, and their workspace sync path is part of the runtime substrate every remote coding run depends on. > - In the sandbox and SSH execution-target flows, Paperclip logged that sync had started, then often went silent for the full transfer window. > - That made large remote syncs feel stalled and also hid a real performance problem in the command-managed sandbox upload path. > - The first part of this pull request threads a throttled progress-reporting surface through the adapter execution-target stack so sync and restore work can emit meaningful updates. > - The second part fixes the command-managed sandbox transport itself: it removes the old serial 32KB append bottleneck, but also falls back away from the single-stream path when a provider-backed sandbox runner cannot surface mid-flight stdin progress. > - The result is that sandbox and SSH transfers are both faster and more observable, including the live Daytona-style sandbox case that previously only emitted `0%` and `100%`. ## Linked Issues or Issue Description No public GitHub issue exists for this bug, so it is described inline below following the bug report template. ### What happened - Remote sandbox and SSH workspace syncs could spend a long time transferring data while only logging a start line (`Syncing workspace and runtime assets to sandbox environment`) and, at best, a terminal line. - In the command-managed sandbox path, the original upload implementation also paid a large performance cost by appending base64 data in many small sequential remote writes (thousands of serial 32KB round-trips on a large workspace). - After the initial transport rewrite, live provider-backed sandbox runs still only emitted `0%` and `100%` because the single-stream stdin RPC buffered progress until completion. ### Expected behavior - Long-running sandbox and SSH syncs should periodically report how much of the transfer is complete (a percentage and/or MB transferred) so an operator can tell the run is healthy and making progress rather than stuck. - The main sandbox upload path should not be artificially slow. - A transfer that fails partway should leave an explicit failure marker in the log rather than a dangling intermediate percentage. ### Steps to reproduce 1. Run an agent against a sandbox (command-managed) or SSH (remote-managed) execution target with a non-trivial workspace. 2. Watch the run log during the workspace/runtime asset sync phase. 3. Observe that the log shows the sync start line and then stays silent for the full transfer (live provider-backed sandbox runs only show `0%` then `100%`). ### Paperclip version or commit - Branch `PAPA-825-provide-status-updates-when-syncing-sandboxes` off `master`. ### Deployment mode - Self-hosted / local instance using sandbox (command-managed) and SSH (remote-managed) execution targets, including provider-backed sandbox runners. ## What Changed - Added shared throttled runtime progress reporting and threaded `onProgress` through the adapter execution-target surface and adapter `execute.ts` entrypoints. - Added sync and restore progress reporting for the command-managed sandbox path and the SSH/remote-managed path, including git import/export progress where totals are known. - Reworked command-managed sandbox transfer behavior so uploads use the faster single-stream path when appropriate, but fall back to chunked progress-emitting writes when the runner cannot expose mid-stream stdin progress. - Marked provider-backed environment sandbox runners as not supporting single-stream stdin progress so live sandbox runs emit meaningful intermediate updates instead of only `0%` and `100%`. - Emit an explicit terminal failure marker (`failed at NN% (x/y MB)`) when an SSH/tar transfer rejects, so a failed sync no longer leaves a dangling intermediate percentage in the log. - Run the SSH sync/restore size estimate (local directory walk / remote `du` probe) concurrently with the transfer instead of awaiting it before opening the pipe, so progress instrumentation no longer adds startup latency proportional to workspace file count. - Added and extended focused regression coverage for runtime progress throttling and the new failure marker, command-managed sandbox transfers, sandbox orchestration, SSH transfer progress, and environment execution-target wiring. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/runtime-progress.test.ts packages/adapter-utils/src/ssh-fixture.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` - `pnpm exec vitest run packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/environment-execution-target.test.ts` - `npx tsc --noEmit` for `packages/adapter-utils` ## Risks - The provider-backed sandbox fallback now prefers chunked command-managed writes when progress hooks are active, so small-to-medium uploads may trade some raw throughput for observable intermediate progress on runtimes that cannot surface true mid-stream stdin progress. - Progress percentages on tar-based transfers still depend on estimates in some cases, so operators may briefly see MB-only lines before the estimate resolves, then near-final clamping before the terminal `100%` line. - This PR changes shared execution-target behavior used by multiple adapters, so regressions would most likely appear in remote runtime setup/teardown flows rather than in a single adapter. ## Model Used - Initial implementation: OpenAI GPT-5.4 via Codex local agent (`codex_local`), high reasoning mode. - Observability follow-ups (failure marker, concurrent size estimate, added tests): Claude Opus 4.8 via Claude Code (`claude_local`). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
486fb88a15 |
Add Cloudflare sandbox provider plugin (#5687)
> _Stacked on top of #5685 → #5686. Diff against master includes commits from earlier PRs in the stack — review focuses on the two new commits (`Extend sandbox callback bridge for Worker-hosted plugins` + `Add Cloudflare sandbox provider plugin`)._ ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Each agent runs in a sandbox environment, and operators choose which provider backs that sandbox — today E2B and Daytona are bundled with the platform > - Cloudflare Workers + Durable Objects + the Sandbox SDK offer a credible new option: globally distributed, cheap idle, and operator-deployable as a single Worker > - To plug it in, Paperclip needs (a) a provider plugin that speaks the `PaperclipPluginManifestV1` lifecycle and (b) a small operator-deployed Worker — the **bridge** — that adapts Paperclip's runtime RPCs to the Cloudflare Sandbox SDK > - The plugin extends the existing sandbox-callback-bridge with a `bridge.transport: "worker"` discriminator so the platform routes runtime RPCs through the Worker bridge instead of the in-process runner > - This pull request adds the plugin, the bridge Worker template, and the supporting adapter-utils + server hooks the new transport needs > - The benefit is that operators can run sandboxes on Cloudflare's edge with no new platform code beyond installing the plugin and deploying the Worker ## What Changed **Shared support (`Extend sandbox callback bridge for Worker-hosted plugins`):** - `packages/adapter-utils/src/sandbox-callback-bridge.{ts,test.ts}`: expose `expectedHostHeader` so plugin-side bridge clients can verify the canonical request envelope before forwarding. - `packages/adapter-utils/src/command-managed-runtime.{ts,test.ts}`: relax the always-fresh runner construction so callers can re-use a runner across exec calls (Worker-hosted bridges hold the runner inside a Durable Object). - `server/src/services/environment-runtime.ts` + `environment-runtime.test.ts`: route Worker-hosted bridges through the same env-shaping path as E2B and pin the `requestEnv` contract. - `server/src/services/plugin-environment-driver.ts`: thread an optional `issueId` through the runtime descriptor so bridges can scope leases to the originating issue (used by Cloudflare to map a sandbox to the issue/workflow for billing and audit). - `packages/plugins/sdk/src/protocol.ts`: add `issueId?` to `PluginEnvironmentDriverBaseParams` and the new `bridge.transport: "worker"` discriminator that the new plugin declares. - `server/__tests__/heartbeat-plugin-environment.test.ts`: pin the heartbeat path against the new runtime descriptor. **The Cloudflare plugin itself (`Add Cloudflare sandbox provider plugin`):** - `packages/plugins/sandbox-providers/cloudflare/`: plugin entry, manifest, plugin runtime (lifecycle + bridge client), config parsing, and Vitest coverage. Manifest declares `bridge.transport: "worker"` so the platform routes runtime RPCs through the bridge client. - `bridge-template/`: a Worker template the operator deploys with `wrangler`. Owns Durable Object-backed sessions (`sessions.ts`), exec/stream routes (`exec.ts`, `routes.ts`), and an HMAC auth layer (`auth.ts`) that pins the `Host` header surface. Includes the SDK-contract-correct exec implementation, lease recovery, and chunked stdout/stderr streaming. - Tests cover lease/session handoff (`bridge-template/src/exec.test.ts`, `routes.test.ts`), bridge client request shaping (`src/bridge-client.test.ts`), and end-to-end plugin behavior (`src/plugin.test.ts`) including streamed exec output. 27 tests in total. - `README.md` walks the operator through deploying the bridge Worker, registering the plugin, and configuring the runtime. ## Verification - `pnpm typecheck` - `pnpm exec vitest run --no-coverage packages/adapter-utils/src/sandbox-callback-bridge.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-plugin-environment.test.ts` - `(cd packages/plugins/sandbox-providers/cloudflare && pnpm test)` — 27 passing For an operator-side smoke test: 1. Deploy the bridge: `cd packages/plugins/sandbox-providers/cloudflare/bridge-template && wrangler deploy` 2. Register the plugin in your Paperclip instance, point its bridge URL at the deployed Worker, set the HMAC shared secret. 3. Create a sandbox environment whose provider is `cloudflare`, then run a Codex or Claude job against it. ## Risks - Adds a new `bridge.transport: "worker"` code path, but the existing E2B / Daytona transports go through the same shaped helpers and have explicit test coverage that pins their behavior unchanged. - The Worker bridge stores session state in a Durable Object; operator instances must be aware of the corresponding Cloudflare costs (DO requests, storage). Documented in the README. - The `issueId` plumbing is optional throughout — existing plugins that don't supply it continue to work. ## Model Used - Provider: Anthropic - Model: Claude Opus 4.7 (1M context) - Capabilities used: extended reasoning, tool use (Read/Edit/Bash/Grep) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots — N/A, no UI change - [x] I have updated relevant documentation to reflect my changes (plugin README, bridge-template README) - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
778e775c35 |
Add secrets provider vaults and remote import (#5429)
## Thinking Path > - Paperclip orchestrates AI-agent companies and needs secrets handling to work across local development, hosted operators, and governed agent execution. > - The affected subsystem is the company-scoped secrets control plane: database schema, server services/routes, CLI workflows, and the Secrets settings UI. > - The gap was that secrets were local-only and operators could not manage provider vaults or import existing remote references without exposing plaintext. > - This branch adds provider vault configuration plus an AWS Secrets Manager remote-import path while preserving company boundaries, binding context, and audit trails. > - I kept the PR to a single branch PR, removed unrelated lockfile/package drift, rebased the full branch onto the current `public-gh/master`, and addressed fresh Greptile findings. > - The benefit is a reviewable implementation of provider-backed secrets with focused tests covering provider selection, import conflicts, deleted secret reuse, rotation guards, and AWS signing behavior. ## What Changed - Added provider vault support for company secrets, including provider config storage, default vault handling, health checks, binding usage, access events, and remote import preview/commit. - Added an AWS Secrets Manager provider using SigV4 request signing, bounded request timeouts, namespace guardrails, cached runtime credential resolution, and external-reference linking without plaintext reads. - Added Secrets UI surfaces for vault management and remote import, plus CLI/API documentation for setup and operations. - Stabilized routine webhook secret binding paths and SSH environment-driver fixture bindings discovered during verification. - Addressed Greptile and CI findings: no lockfile/package drift, monotonic migration metadata, disabled-vault default races, soft-deleted secret hiding/recreate behavior, remove behavior with disabled vaults, soft-deleted external-reference re-import, non-active rotation guards, managed-secret soft deletion through PATCH, and per-call AWS SDK credential client churn. - Rebased this branch onto `public-gh/master` at `0e1a5828` and force-pushed with lease to keep this as the single PR for the branch. ## Verification - `git fetch public-gh master` - `git rebase public-gh/master` - `git diff --name-only public-gh/master...HEAD | grep '^pnpm-lock\.yaml$' || true` confirmed `pnpm-lock.yaml` is not in the PR diff. - Confirmed migration ordering: master ends at `0081_optimal_dormammu`; this PR adds `0082_dry_vision` and `0083_company_secret_provider_configs`. - Inspected migrations for repeat safety: new tables/indexes use `IF NOT EXISTS`; foreign keys are guarded by `DO $$ ... IF NOT EXISTS`; column additions use `ADD COLUMN IF NOT EXISTS`. - `pnpm -r typecheck` passed before the Greptile follow-up commits. - `pnpm test:run` ran the full stable Vitest path before the Greptile follow-up commits; it completed with 3 timing-related failures under parallel load: `codex-local-execute.test.ts`, `cursor-local-execute.test.ts`, and `environment-service.test.ts`. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/codex-local-execute.test.ts src/__tests__/cursor-local-execute.test.ts src/__tests__/environment-service.test.ts` passed on targeted rerun (`24/24`). - `pnpm build` passed before the Greptile follow-up commits. Vite reported existing chunk-size/dynamic-import warnings. - After Greptile follow-up commits: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` passed (`26/26`). - After Greptile follow-up commits: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/aws-secrets-manager-provider.test.ts src/__tests__/secrets-service.test.ts` passed (`39/39`). - After Greptile follow-up commits: `pnpm --filter @paperclipai/server typecheck` passed. - Captured Storybook screenshots from `ui/storybook-static` for visual review. - Latest PR checks on `5ca3a5cf`: `policy`, serialized server suites 1/4-4/4, `Canary Dry Run`, `e2e`, `security/snyk`, and `Greptile Review` pass; aggregate `verify` is still registering the completed child checks. - Greptile review loop continued through the latest requested pass; all Greptile review threads are resolved and the latest `Greptile Review` check on `5ca3a5cf` passed with 0 comments added. ## Screenshots Before: the provider-vault and remote-import surfaces did not exist on `master`; these are after-state screenshots from the Storybook fixtures.    ## Risks - Migration risk: this adds new secret provider tables and extends existing secret rows. The migrations were checked for monotonic ordering and idempotent guards, but reviewers should still inspect upgrade behavior carefully. - Provider risk: AWS support uses direct SigV4 requests. Automated tests cover signing, request timeouts, vault-config selection, namespace guardrails, pending-version archival, sanitized provider errors, and service-level cleanup paths. A real-vault AWS smoke test remains deployment validation for an operator with AWS credentials rather than an unverified merge blocker in this local branch. - UI risk: the Secrets page and import dialog are large new surfaces; screenshots are included above for reviewer inspection. - Verification risk: the full local stable test command hit parallel-load timing failures, although the exact failed files passed when rerun directly. - Operational risk: remote import intentionally avoids plaintext reads; operators must understand that imported external references resolve at runtime and may fail if AWS permissions change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent with local shell/tool use in the Paperclip worktree. Exact context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
fe3904f434 |
Stabilize runtime probes and Codex env tests (#5445)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - Adapters expose a Test action that probes the configured runtime —
install, resolvability, hello — to give operators a fast yes/no on
whether an environment is healthy
> - The Codex test path was running its hello probe directly without
going through the managed-runtime preparation that production runs use,
so a healthy production setup could still report a probe failure
> - The plugin worker manager wasn't surfacing terminated workers
cleanly, leaving the runtime probe waiting on a dead worker until the
request timed out
> - This pull request routes the Codex test probe through
`prepareAdapterExecutionTargetRuntime` (so it sees the same managed
Codex home production sees), exposes `commandCwd` on
`createCommandManagedRuntimeClient` so callers can target a per-probe
directory without leaking the workspace `remoteCwd`, and propagates
plugin-worker termination as a usable error instead of a hang
> - The benefit is the Codex Test action mirrors production behavior
end-to-end, and probes against a terminated plugin worker fail fast
instead of timing out
## What Changed
- `packages/adapter-utils/src/command-managed-runtime.ts`: rename the
`remoteCwd` knob to `commandCwd` so callers can target a per-probe
directory without inheriting the workspace cwd; matching test coverage
in `command-managed-runtime.test.ts`
- `packages/adapter-utils/src/sandbox-callback-bridge.{ts,test.ts}`:
small fixes to keep callback bridge stop semantics deterministic
- `packages/adapters/codex-local/src/server/test.ts`: thread the Codex
hello probe through `prepareAdapterExecutionTargetRuntime` +
`prepareManagedCodexHome` so the probe sees the same managed home
production sees; new `test.remote.test.ts` covers the remote probe path
- `packages/adapters/cursor-local/src/server/execute.ts`: small
probe-side cleanup that aligns with the new commandCwd contract
- `server/src/services/plugin-worker-manager.ts`: surface plugin-worker
termination as a structured error so callers fail fast; new
`plugin-worker-terminated.cjs` fixture and
`plugin-worker-manager.test.ts` cases pin the behavior
## Verification
- `pnpm vitest run --no-coverage --project @paperclipai/adapter-utils
--project @paperclipai/adapter-codex-local --project
@paperclipai/adapter-cursor-local --project @paperclipai/server` —
1749/1750 passing (1 unrelated skip)
- `pnpm typecheck` clean
## Risks
Low–medium. The `remoteCwd → commandCwd` rename is a parameter renaming
on an internal helper used only by adapter test/execute paths in this
repo. The plugin-worker-terminated path was previously a hang; failing
fast may surface latent timeouts as explicit termination errors in
callers that already expected them.
## Model Used
Claude Opus 4.7 (1M context)
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable — new tests cover
commandCwd, plugin-worker termination, and Codex remote test path
- [x] If this change affects the UI, I have included before/after
screenshots — N/A (no UI)
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
---
> **Stacked PR.** Sits on top of #5444 which adds the per-run runtime
API surface this PR builds on. Cumulative diff against `master` includes
that PR's content; the files touched by *this* PR's commit are listed
under "What Changed" above. Will rebase onto `master` and force-push
once #5444 merges.
|
||
|
|
9578dc3da7 |
Wire per-adapter sandbox install commands through test and execute paths (#5280)
> **Stacked PR.** Sits on top of the e2b sandbox chain — #5278 (stdin staging) and #5279 (honest-resolvability + login-profiles). The cumulative diff against `master` includes both of those PRs' content; the files touched by *this* PR's commit are the new `maybeRunSandboxInstallCommand` helper in `packages/adapter-utils/src/execution-target.ts` and the per-adapter `index.ts`/`server/test.ts`/`server/execute.ts` wiring under `packages/adapters/{claude,codex,cursor,gemini,opencode,pi}-local/`. The honest resolvability check from #5279 is what gives this PR's install command a meaningful "did it actually land on PATH" follow-up. ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Sandbox execution targets are ephemeral — each fresh lease starts from a template image that may or may not have the agent CLIs preinstalled > - When a CLI isn't preinstalled, the resolvability probe fails at `command -v` and the hello probe never runs > - There's no shared mechanism for "before you probe or provision, install the CLI on this sandbox" > - This pull request adds a `SANDBOX_INSTALL_COMMAND` constant per adapter and a `maybeRunSandboxInstallCommand` helper that runs it via the existing sandbox login shell, captures structured output, and never throws (so the resolvability + hello probe still run after); each adapter's `test()` and `execute()` share the constant so the two callsites can't drift > - The benefit is a fresh sandbox lease without a preinstalled CLI now installs it once via `sh -lc` before the resolvability probe and before managed-runtime provisioning, with a uniform `<adapter>_install_command_run` check on the test report ## What Changed - `packages/adapter-utils/src/execution-target.ts`: add `AdapterSandboxInstallCommandCheck` and `maybeRunSandboxInstallCommand` (runs the install via existing sandbox shell, captures exit/stdout/stderr, returns a structured info/warn check, never throws) - Add `SANDBOX_INSTALL_COMMAND` to each adapter's `index.ts` so `test()` and `execute()` share a single source of truth - Wire each of the 6 affected adapter `testEnvironment()`s to call `maybeRunSandboxInstallCommand` before `ensureAdapterExecutionTargetCommandResolvable` - Pass `installCommand: SANDBOX_INSTALL_COMMAND` through `prepareAdapterExecutionTargetRuntime` in each adapter's `execute()` - Per-adapter install commands use npm globals where possible so binaries land on a PATH segment the template already exports: - claude → `npm install -g @anthropic-ai/claude-code` - codex → `npm install -g @openai/codex` - cursor → `curl https://cursor.com/install -fsS | bash` - gemini → `npm install -g @google/gemini-cli` - opencode → `npm install -g opencode-ai` - pi → `npm install -g @mariozechner/pi-coding-agent` SSH and local targets ignore `installCommand` (SSH runtime takes no such param; local short-circuits before runtime prep), so this is a no-op for non-sandbox environments. ## Verification - `pnpm typecheck` clean - `pnpm vitest run --no-coverage --project @paperclipai/adapter-utils` and per-adapter projects pass - Manual sandbox matrix (claude, codex, cursor, gemini, opencode, pi) — each goes `install_command_run → resolvable → hello_probe_passed` (Codex and Pi land on `hello_probe_auth_required`, which is the configured-credentials problem, not an install issue) - SSH no-regression: SSH Claude still passes; the helper short-circuits on non-sandbox targets ## Risks Medium — adds a network/CPU cost (npm install / curl) on every fresh sandbox lease. Cost is bounded (one-time per lease, typically tens of seconds for npm globals), and the helper never throws so a failing install still lets the report run resolvability and hello probes. If a sandbox image already has the CLI, the install is an idempotent reinstall. ## Model Used Claude Opus 4.7 (1M context) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots — N/A (no UI) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
076067865f |
Migrate SSH environment callback to bridge (#5116)
> **Stacked PR (part 3 of 7).** Depends on: - PR #5114 - PR #5115 > Diff against `master` includes commits from earlier PRs in the stack — the new commit in this PR is the topmost one. ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Agents executing on a remote SSH-backed environment need a way to call back into > the Paperclip control plane (run events, log streaming, signals) > - When the SSH host can't reach the Paperclip host (NAT, firewalls, or simply not > on the same network), the run silently fails or hangs — a recurring class of > failure during SSH testing > - In sandboxed environments we already solved this with a callback bridge that > tunnels back through the existing connection; SSH was the odd one out > - This PR migrates SSH execution to use the same callback bridge, so every > adapter's remote run uses one consistent reverse-channel. Per-adapter SSH glue > is deleted in favour of a shared `CommandManagedRuntimeRunner` built from the > SSH spec > - The benefit is fewer SSH-specific failure modes, a smaller code surface, and > one place to evolve the callback contract going forward ## What Changed - Added `createSshCommandManagedRuntimeRunner` in `packages/adapter-utils/src/ssh.ts` that adapts an SSH spec into a generic command-managed-runtime runner (with cwd, env, and timeout handling) - Removed `paperclipApiUrl` from `SshRemoteExecutionSpec`; the bridge URL now flows through the shared runner - Reworked `execution-target.ts` to use the SSH runner alongside sandbox runners via a unified `CommandManagedRuntimeRunner` interface - Simplified `remote-managed-runtime.ts` and `sandbox-managed-runtime.ts` to consume the shared runner abstraction - Deleted per-adapter SSH callback wiring from claude-local, codex-local, cursor-local, gemini-local, opencode-local, pi-local execute.ts files - Removed `environment-runtime-driver-contract.test.ts` (the contract is now enforced by `environment-execution-target.test.ts`) - Added/updated `execute.remote.test.ts` cases for each adapter to cover the SSH runner path ## Verification - `pnpm --filter @paperclipai/adapter-utils test` - `pnpm test -- execute.remote` (covers all six local adapters' SSH paths) - Manual QA: ran a claude-local agent against an SSH-backed environment, confirmed the agent successfully called back to `/api/agent-callback/*` endpoints during the run ## Risks - Refactor touches all six local adapters. If any adapter had subtle SSH-specific behaviour that wasn't captured in tests, it could regress. Mitigation: each adapter's `execute.remote.test.ts` was extended. - `paperclipApiUrl` removal from `SshRemoteExecutionSpec` is a breaking type change for any internal consumer. Verified no external plugins consume this type. - The new `CommandManagedRuntimeRunner` shape is a public surface in `@paperclipai/adapter-utils`; downstream plugins implementing custom runners may need updates, but no such plugins exist in this repo. ## Model Used - OpenAI GPT-5.4 (reasoning effort: high) via Codex CLI - Provider: OpenAI - Used to author the code changes in this PR ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots — N/A - [ ] I have updated relevant documentation to reflect my changes — N/A - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a7b45938b7 |
Let sandbox providers declare shell defaults (#5114)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - Agents execute in sandboxed remote environments served by pluggable
sandbox
> providers (E2B today, more later)
> - Today every sandbox command runs under `sh -lc` regardless of what
the
> provider's container actually ships
> - That misses bash-only shell init on E2B (which ships bash) and
prevents
> future providers from declaring a different default — there's no way
for a
> provider to say "I have bash, use it"
> - This PR adds a `shellCommand` field to sandbox execution targets so
providers
> can declare their preferred shell ("bash" for E2B), threads it through
the
> sandbox-managed-runtime client, callback bridge, and execution-target
shell
> helper, and validates the value at the lease-metadata boundary
> - The benefit is that sandbox commands run under the right shell on
the right
> provider, and adding new sandbox providers only needs to declare a
shell
> preference
## What Changed
- Added `packages/adapter-utils/src/sandbox-shell.ts` exporting
`preferredShellForSandbox(shellCommand)` (returns `"bash"` if input is
`"bash"`,
else `"sh"`)
- Added `shellCommand?: "bash" | "sh" | null` to
`AdapterSandboxExecutionTarget`
and `CommandManagedRuntimeSpec`; threaded it through
`runAdapterExecutionTargetShellCommand`,
`prepareAdapterExecutionTargetRuntime`,
and `startAdapterExecutionTargetPaperclipBridge`
- `createCommandManagedRuntimeClient`, `prepareCommandManagedRuntime`,
and
`createCommandManagedSandboxCallbackBridgeQueueClient` now take an
optional
`shellCommand` and use `preferredShellForSandbox` to pick the shell
- `startSandboxCallbackBridgeServer` accepts a `shellCommand` for its
server
startup, readiness probe, and stop hook
- E2B sandbox plugin declares `shellCommand: "bash"` in `leaseMetadata`
- `resolveEnvironmentExecutionTarget` reads `shellCommand` from lease
metadata
(validating against `"bash" | "sh" | null`)
- `environment-runtime.ts` adds `"shellCommand"` to
`INTERNAL_PLUGIN_SANDBOX_CONFIG_KEYS`
so the field round-trips through internal plugin config without leaking
to
external plugin metadata
- Updated tests in `command-managed-runtime.test.ts`,
`execution-target-sandbox.test.ts`, `sandbox-callback-bridge.test.ts`,
`environment-execution-target.test.ts`
## Verification
- `pnpm --filter @paperclipai/adapter-utils test`
- `pnpm --filter @paperclipai/server test --
environment-execution-target`
- `pnpm --filter @paperclipai/sandbox-providers-e2b test`
- Manual QA: boot a Paperclip instance, create an E2B-backed
environment, run a
claude_local agent against it, and confirm the run completes (verifies
bash
shell semantics flow through the callback bridge end-to-end)
## Risks
- E2B sandbox commands now run under `bash -lc` instead of `sh -lc`.
Bash is a
strict superset for the commands we issue (no busybox-only flags in our
shell
scripts), so risk is low. The shellCommand field is opt-in via lease
metadata —
providers that don't declare it stay on `sh`.
- New optional field on `CommandManagedRuntimeSpec` and
`AdapterSandboxExecutionTarget`.
Consumers ignoring the field retain previous behaviour (sh).
- Lease metadata now carries an additional field. Existing leases
without
`shellCommand` resolve to `null` and fall back to sh — backwards
compatible.
## Model Used
- OpenAI GPT-5.4 (reasoning effort: high) via Codex CLI
- Provider: OpenAI
- Used to author the code changes in this PR
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] If this change affects the UI, I have included before/after
screenshots — N/A (no UI changes)
- [ ] I have updated relevant documentation to reflect my changes — N/A
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
a4ac6ff133 |
Add sandbox callback bridge for remote environment API access (#4801)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Agents can run inside sandboxed environments like E2B, which are isolated from the host network > - Sandboxed agents need to call back to the Paperclip API to report progress, post comments, and update issue status > - But sandbox environments cannot reach the Paperclip server directly because they run in isolated network namespaces > - This PR adds a callback bridge that proxies API requests from the sandbox to the Paperclip server, running as a local HTTP server on the host that forwards authenticated requests > - The bridge is started automatically when an adapter launches a sandbox execution, and torn down when the run completes > - The benefit is sandboxed agents can interact with the Paperclip API without requiring network-level access to the host, enabling E2B and similar providers to work end-to-end ## What Changed - Added `sandbox-callback-bridge.ts` in `packages/adapter-utils/` — a lightweight HTTP bridge server that accepts requests from sandbox environments and proxies them to the Paperclip API with authentication - Added request validation and security policy: the bridge only forwards requests to the configured API URL, validates content types, enforces size limits, and rejects non-API paths - Wired the bridge into all remote adapter execute paths (claude, codex, cursor, gemini, pi) — the bridge starts before the agent process and the bridge URL is passed via environment variables - Updated `environment-execution-target.ts` to prefer the explicit API URL from environment lease metadata for sandbox callback routing - Fixed Claude sandbox runtime setup to work with the bridge configuration - Added comprehensive test coverage for bridge request handling, policy enforcement, and sandbox execution integration - Fixed browser bundling — the bridge module is excluded from the frontend bundle via the adapter-utils index export ## Verification - `pnpm test` — all existing and new tests pass, including bridge unit tests and sandbox execution integration tests - `pnpm typecheck` — clean - Manual: configure an E2B environment, run an agent task, verify the agent can post comments and update issue status through the bridge ## Risks - Medium. This is a new network-facing component (HTTP server on localhost). The security policy restricts forwarding to the configured API URL only and validates all requests, but any proxy introduces attack surface. The bridge binds to localhost only and is scoped to the lifetime of a single agent run. ## Model Used Codex GPT 5.4 high via Paperclip. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
70679a3321 |
Add sandbox environment support (#4415)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The environment/runtime layer decides where agent work executes and how the control plane reaches those runtimes. > - Today Paperclip can run locally and over SSH, but sandboxed execution needs a first-class environment model instead of one-off adapter behavior. > - We also want sandbox providers to be pluggable so the core does not hardcode every provider implementation. > - This branch adds the Sandbox environment path, the provider contract, and a deterministic fake provider plugin. > - That required synchronized changes across shared contracts, plugin SDK surfaces, server runtime orchestration, and the UI environment/workspace flows. > - The result is that sandbox execution becomes a core control-plane capability while keeping provider implementations extensible and testable. ## What Changed - Added sandbox runtime support to the environment execution path, including runtime URL discovery, sandbox execution targeting, orchestration, and heartbeat integration. - Added plugin-provider support for sandbox environments so providers can be supplied via plugins instead of hardcoded server logic. - Added the fake sandbox provider plugin with deterministic behavior suitable for local and automated testing. - Updated shared types, validators, plugin protocol definitions, and SDK helpers to carry sandbox provider and workspace-runtime contracts across package boundaries. - Updated server routes and services so companies can create sandbox environments, select them for work, and execute work through the sandbox runtime path. - Updated the UI environment and workspace surfaces to expose sandbox environment configuration and selection. - Added test coverage for sandbox runtime behavior, provider seams, environment route guards, orchestration, and the fake provider plugin. ## Verification - Ran locally before the final fixture-only scrub: - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - Ran locally after the final scrub amend: - `pnpm vitest run server/src/__tests__/runtime-api.test.ts` - Reviewer spot checks: - create a sandbox environment backed by the fake provider plugin - run work through that environment - confirm sandbox provider execution does not inherit host secrets implicitly ## Risks - This touches shared contracts, plugin SDK plumbing, server runtime orchestration, and UI environment/workspace flows, so regressions would likely show up as cross-layer mismatches rather than isolated type errors. - Runtime URL discovery and sandbox callback selection are sensitive to host/bind configuration; if that logic is wrong, sandbox-backed callbacks may fail even when execution succeeds. - The fake provider plugin is intentionally deterministic and test-oriented; future providers may expose capability gaps that this branch does not yet cover. ## Model Used - OpenAI Codex coding agent on a GPT-5-class backend in the Paperclip/Codex harness. Exact backend model ID is not exposed in-session. Tool-assisted workflow with shell execution, file editing, git history inspection, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |