mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 20:50:08 +02:00
codex/plugin-task-execution
54
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a2461c4c2b |
Reframe capability as plugin-provided Runner execution
Use the Runner execution name throughout the guide, capability comments, diagnostics, and test descriptions. Keep the environment-driver RPC contract and project preparation flow intact. Validation: SDK build, server typecheck, and 22 focused host/worker tests passed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
62467a9a5a |
Merge master and preserve task admission dispatch
Keep the new idle-work tracking import alongside the plugin task dispatcher. The remaining changes merge cleanly from master. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d35a584f83 |
Support project-aware environment task admission
Submit accepts multiple project IDs and the host validates their company scope before dispatch. Keep PRP compatibility for the Runner connection while documenting provider-owned resource preparation and task lifecycle. Verified SDK build, server typecheck, and 21 focused task and worker contract tests including multi-project scope rejection. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3367b75ccc |
fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A host can stop an idle instance to reduce unused compute. > - Zero active runs do not prove that requests, saved work, or cleanup are complete. > - A client can disconnect while its request still writes, and failed cleanup can remain only in memory or on disk. > - This pull request holds new admission and checks accepted work, cleanup, and persisted work under an owned drain. > - The host gets an empty report only while those checks remain valid. ## Linked Issues or Issue Description Refs #13413. That companion PR uses the same owned-hold vocabulary for runtime services. This PR covers HTTP requests, scheduler work, accounting, cleanup, and the instance work inventory. It does not include the preview gateway changes. **What existing behavior does this improve?** The instance task-drain API reports process counters. It does not prove that a host can safely stop the instance for idle sleep. **Current behavior** A quiet run set can coexist with an unfinished request, cleanup after a completed run, future work, or a failed accounting write. **Proposed behavior** Provide a bounded owned idle hold. Block new ingress, track accepted handler promises, inspect durable and local work, and return `none` only when the same hold stays quiet through the checks. Keep normal deployment drains compatible. **Reason and benefit** Hosts can identify eligible idle instances without treating a disconnect or failed cleanup write as completed work. ## What Changed - Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked release to task drain. Existing holds cannot be replaced by another API request. - Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new WebSocket upgrades. Track async handlers in nested Express routers and error middleware until they settle, even after the response or client disconnect. - Count accepted live-event WebSocket authentication through settlement, even after disconnects. Count detached built-in agent, managed-home and runtime-service startup reconciliation after readiness. - Keep health and control mutations tracked. Count control-request authentication separately from the read-only report, including concurrent user/company/membership writes. - Pause new scheduler admissions during idle holds. Count work already in flight, including database backups, and reject scans whose work generation changes. Periodic backups block sleep without a host wake schedule. - Inspect accounting and orphan-cleanup spool directories without skipping temporary or malformed entries. Retain orphan tokens through queue splices, flush failures, and buffer overflow. Keep failed usage capture counted until its database failure fence is written. - Check persisted work across companies in a bounded read-only transaction. Include deferred agent-file cleanup and saved watchdogs whose watched issues are complete. Enabled plugins and unsupported retained work remain blockers. - Require the exact idle owner and completed startup tracking before returning an empty report. Keep reports free of tenant details. - Document the hosting protocol, retry behavior, conservative blockers, and the remaining external provider-stop race. ## Verification - Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper extractions are preserved. GitHub confirms no merge conflicts. - Focused heartbeat renderer/run-log, drain, control-auth, admission, route and PostgreSQL inventory checks: 234 tests passed in 10 suites after the rebase. - The POST task-drain contract includes the expected `409` conflict response. The 84 OpenAPI and instance-settings route tests and server typecheck passed after that final documentation fix. - Final accepted-upgrade/startup regression run: 104 tests passed in five suites, including success and failure after readiness or disconnect. Final server typecheck also passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm install --frozen-lockfile` and the PR diff secret scan passed. No dependency or lockfile changes are added by this PR. - The prior verification covered HTTP disconnects and early responses, async error handlers, concurrent authentication writes, signed bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and failed accounting fences. Those tests passed. - The last full local `pnpm test:run` stopped in the general-server phase with 16,658 passing tests and five skill/connector fixture failures caused by an ancestor workspace skill directory. That full local run preceded this rebase and has not been repeated for the import conflict. Full current-head CI passed: 53 successful checks and two conditional skips, with no failures. - All eight review findings are fixed and their threads resolved, including accepted upgrade authentication, detached startup writes and the POST conflict contract. Current-head Apex review is 5/5 with no new findings; all eight review threads remain resolved. - No live provider stop or production change was performed. ## Risks - The hosting controller must use the owned protocol and hold external admission through its final validation and provider stop. A legacy drain cannot authorize idle sleep. During an idle hold, new requests receive 503 with `Retry-After: 1`; the host must handle queueing or retry before enabling this path. - This is a single-process protocol. An unexpected restart after the last validation can race an external provider stop. The host must serialize deploy/wake/sleep operations and bind the validation to the instance it stops. Multiple replicas need shared fencing. - Some retained state conservatively prevents sleep, including every enabled plugin. This PR does not promise that every inactive instance becomes eligible. - The HTTP adapter uses Express 5 router layers. Real Express tests cover nested routes, errors and disconnects. New routes must register before tracking is installed. Detached work must have durable state or explicit work tracking. - Unrecoverable in-memory cleanup debt keeps the instance awake until reconciliation. This change does not make such debt survive an unplanned process crash. No schema migration or provider configuration change is included. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact deployment variant/model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local fixture limitation and clean full CI result are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1f063ac651 |
Define remote Runner tasks by PRP compatibility
Require the submitting client PRP protocolMin/protocolMax range and validate its bounds. Describe the capability as direct remote Paperclip Runner execution throughout the SDK and documentation. Verified the SDK build, server typecheck, and all 19 focused task and worker contract tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
465140596f |
Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials. Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c84d7b35ec |
Add typed task operations to environment plugins
Provide a negotiated task RPC for environment providers that own Runner admission. Dispatch from persisted company-scoped leases and validate typed requests and receipts without coupling the host to a provider API. Cover the worker contract, host-derived identity, capability negotiation, and failed or mismatched receipts. Document retry and credential handling. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ffa32373bc |
fix(runtime): honor provider acquisition timeout defaults (#14097)
Declare provider acquisition budgets so slow Daytona creation does not hit the host’s 30-second fallback. Bound creation and setup to one deadline and preserve scoped cleanup ownership after timeout. Legacy drivers retain their original call shape. Validated by 238 provider/manifest tests, focused database and heartbeat regressions, root and standalone provider typecheck/build, and full CI. Local broad tests also expose recorded Mac baseline limitations. Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1dbffb4c04 |
fix(daytona): clean up sandbox allocation after create failure (#13979)
## Thinking Path > - Paperclip manages AI agents and their execution environments. > - The Daytona provider creates sandboxes before it returns a lease. > - Daytona can allocate a sandbox, then time out while waiting for it to start. > - The SDK throws without returning the allocated sandbox handle. > - Paperclip previously lost the resource identity and could leave a running sandbox behind. > - This pull request keeps an exact creation identity and deletes a matching failed allocation. ## Linked Issues or Issue Description Related: #13882. Searches for Daytona create cleanup and orphan fixes found no duplicate for this path. **What happened?** In [Grok qualification campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743), Daytona sandbox creation exceeded 300 seconds. No native runner or provider session started. The SDK threw, but an ownership-filtered provider lookup found the allocated sandbox still running. The disposable resource was then deleted separately. The original attempt and its failure remain retained. **Expected behavior** A failed create must retain enough identity to clean up an allocated sandbox. Cleanup must never delete another attempt's resource. If the provider cannot confirm cleanup during sandbox lease acquisition, the host must retain a durable cleanup record and retry the exact owned allocation after restart. **Steps to reproduce** 1. Have Daytona allocate a sandbox for an environment acquisition. 2. Make the SDK throw while waiting for startup, before it returns the sandbox handle. 3. Before this fix, acquisition fails without deleting the allocated sandbox. 4. The regression tests reproduce this failure and verify exact-owner deletion. **Paperclip version or commit** Observed at `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`. The same creation path exists on master; this fix starts from `efce9356b`. ## What Changed - Assign each cold create a unique provider name and creation-attempt label. - On create failure, look up that exact name and verify every ownership label before deletion. - Wait for provider deletion with bounded lookup and deletion calls. - Preserve the original create error after successful cleanup. Report both errors and the provider name when cleanup cannot be confirmed. - Treat a missing lookup after an uncertain create as unconfirmed cleanup: a delayed provider request might still create the sandbox. - Transfer only validated ownership fields through the acquire-lease RPC error; provider exceptions and arbitrary error data are not serialized. - Persist a pending-cleanup lease before the host retries deletion. Reuse the existing cleanup sweep and durable spool, retaining the original environment scope after its row is deleted. - Fence retries by company, environment, run, provider account, unique creation name, and every ownership label. Missing provisional-name lookups remain unresolved. - Journal an ownership-verified provider ID before host-driven deletion. Its absence then confirms cleanup after a lost deletion reply or final database update; failed journal writes block deletion. - Keep ownership labels separate from the mutable input Daytona modifies during create. - Report confirmed host-side deletion accurately while still rejecting the failed acquisition. - Add protocol, malformed-evidence, cross-scope denial, controller-restart, and deleted-environment regressions alongside the original bounded cleanup tests. ## Verification - SDK suite: 92 passed. Daytona plugin suite: 277 passed; six opt-in live tests skipped. Environment-runtime suite: 103 passed, then all four focused observation/recovery variants passed after adding a foreign-observation case. SDK and provider TypeScript checks and `git diff --check` pass. - Regressions failed before the respective fixes: missing durable ownership, misleading successful-cleanup message, and SDK mutation of the ownership labels. - Controlled real-Daytona proof at `d2ac1c95b32ca64daf039c0427137657f351c0e0`: create a real sandbox; inject an SDK failure and a failed immediate lookup; persist the validated ownership envelope; have a child journal the observed provider ID before deleting; deliberately drop its successful deletion receipt; reconcile absence from another fresh process and confirm it with a separate provider read. Passed, with no manual cleanup and zero model calls. The live proof uses an owner-only envelope file; actual host database/spool recovery is tested separately in the environment-runtime suite. - The preceding controlled attempt failed because the real SDK mutated the labels object. That failure is retained. The harness deleted its exact owned sandbox, and a fresh lookup confirmed absence. The fix copies the SDK input labels separately from the ownership snapshot. - [Repository CI](https://github.com/paperclipai/paperclip/actions/runs/36095925399) passes on this exact head, with a clean 5/5 review. The unchanged Cursor remote-command test passed its one bounded rerun after a 10-second timeout; the same test had already passed on the combined tree. The original failure is retained. Earlier `199f0852` CI and immediate-deletion proof are retained, without being promoted to the new source. - Combined-source verification uses temporary PR #13990 against the Grok feature branch. No Docker or Rust build runs on the developer laptop. ## Risks Creation failures now add up to ten seconds for lookup and fifteen seconds for deletion confirmation. The SDK still owns the initial creation timeout. If a request materializes only after the failed lookup, its durable pending record remains eligible for later cleanup; a never-observed creation name is never reported deleted from a 404. A previously observed provider ID can be reconciled as deleted. The host and SDK must be deployed together. Durable handoff applies to plugin-backed sandbox lease acquisition used by heartbeat and runner login. Probe and custom-image interactive setup retain immediate cleanup and explicit failure reporting, but do not use this lease-recovery path. A worker crash before delivering the failure envelope still cannot be recovered through this mechanism. A cleanup error never returns a lease. Successful creation behavior is unchanged except for the provider name and ownership label. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1483bb8bcf |
fix: pass plugin workers to issue tree resume (#13757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue tree controls can resume tasks and wake their assigned agents. > - A sandbox-backed run needs the shared plugin worker manager to acquire its lease. > - The issue tree route constructed a heartbeat service without that manager. > - A ready sandbox provider therefore appeared offline and the resumed run failed setup. > - This pull request passes the existing manager to the route and distinguishes missing wiring from a stopped worker. ## Linked Issues or Issue Description Related implementation: #10262 identified this wiring defect in several dispatch paths. This PR applies the issue-tree resume fix to current master and adds coverage in the current route and runtime suites. The other dispatch paths remain in the scope of that earlier PR. **What happened?** Resuming a paused issue subtree can fail to acquire a sandbox lease even while its provider plugin is ready and its worker is running. The error says the worker is not running because this route's heartbeat service never received the process's worker manager. **Expected behavior** Resumed work uses the same worker manager as normal issue dispatch. Missing wiring and an actually stopped worker produce distinct diagnostic errors. **Steps to reproduce** 1. Configure an agent with a plugin-backed sandbox environment and start the provider worker. 2. Pause a subtree assigned to that agent. 3. Release the hold with `metadata.wakeAgents: true`. 4. The old route constructs an unwired heartbeat service and the resumed run fails during setup. **Paperclip version or commit** Reproduced by source inspection and regression coverage against `9d19f98b50`. **Deployment mode** Server with a plugin-backed sandbox environment. ## What Changed - Pass the process's plugin worker manager from `createApp` through issue tree controls to heartbeat dispatch. - Check for a missing manager before reporting the sandbox worker as stopped. - Exercise the resume request with a shared manager and cover both runtime failure cases. Model a stopped worker explicitly in the existing infrastructure-retry fixture. Existing board and company access checks remain covered. ## Verification - Targeted Vitest route/runtime/recovery suites: 17 passed; 384 native database tests skipped because embedded Postgres cannot start on this host. - Direct server `tsc --noEmit` passed. - `pnpm -r typecheck`, server typecheck wrapper, and `pnpm build` were attempted. Their runner dependency requires `cargo`, which is absent on this host. CI must pass these checks before merge. - Full `pnpm test:run` was attempted. The general server stage reported 8,146 passed, 4,747 skipped, and 14 failed tests across 35 failed suites. Failures were embedded Postgres startup errors (with related teardown errors) and 10 cache-directory rename failures on this macOS host. The local run stopped there; Linux CI must pass the complete suite before merge. - All CI gates are green, including the native database recovery suite, workspace typecheck/build, runner checks, and browser tests. Greptile reviewed the current commit at 5/5 with no unresolved threads. - No live provider operations were performed by these tests. ## Risks Low risk: dependency forwarding and error classification only. The route retains board authorization, company boundaries, hold semantics, and existing cancellation/replay guards. The change does not replace a sandbox or discard a retained lease. No schema or API shape changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tooling, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] This fixes existing behavior and does not add planned core features - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the bug issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no private ticket or instance data - [x] Targeted local tests pass; full-suite and toolchain limits are recorded above - [x] I have added or updated tests where applicable - [x] No documentation changes are needed for dependency forwarding - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e926b13017 |
fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path > - Paperclip must stop remote execution after losing its controller. > - A bridge can remain blocked while the sandbox still incurs costs or performs actions. > - Waiting forever for that bridge prevents provider termination. > - A temporary provider outage can also exhaust cleanup attempts permanently. > - This pull request bounds bridge drain and persists cleanup retries with backoff. > - Cleanup ends only after provider confirmation, without a user accepting uncertain side effects. ## Linked Issues or Issue Description Builds on merged #13285. Related #13254 added exact provider termination receipts; merged #13272 adds explicit user retry. Merged #13352 stops active sandbox startup before waiting for setup. This PR preserves that immediate cancellation path and extends bounded teardown to ordinary release and destroy. Cleanup continues automatically after repeated provider failures. Refs #12953 for provider failures blocking execution. **What happened?** Daytona release waits for in-flight bridge activity before stop/delete. A dead bridge can prevent that wait from finishing. The host also stops cleanup after five failed attempts. **Expected behavior** Provider termination proceeds after a bounded bridge drain. Cleanup retries survive service restarts and provider outages. **Steps to reproduce** Start a sandbox command whose bridge promise never resolves, then release its lease. Separately, persist a pending-cleanup lease with five failed attempts and recover the provider. **Deployment mode** Hosted Paperclip with a Daytona provider; rebased onto master at `728f7185f` on September 14. ## What Changed - Bound bridge drain and provider lifecycle calls. Prefer stop for reusable sandboxes, with delete fallback. - Persist cleanup attempt identity, renewable in-flight deadline, and cooldown. Fence completion writes against superseded attempts. - Preserve scoped explicit Retry and its activity log. Explicit Retry can skip cooldown, but cannot take over a live cleanup attempt. - Continue cleanup after five failures with slower retries and an operator warning. - Exclude leases in cooldown before paging so they do not starve due work. - Add hung-bridge, restart, provider-recovery, and concurrent-cleanup regressions. ## Verification - Rebased onto master at `728f7185f`. The outstanding diff contains only cleanup changes; the merged controller-ownership prerequisite is excluded. - Daytona plugin suite: 160 passed, including immediate startup cancellation, graceful release, hung activity, and teardown regressions. - `pnpm exec vitest run server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31 passed. Two added integration cases verify explicit Retry during cooldown and while another cleanup owns the lease. They also verify run scoping and the activity log. - Targeted cleanup and cancellation cases in `environment-runtime.test.ts`: 20 passed. - Earlier live disposable Daytona test: provider stop ended background work, resume preserved files without restarting the old process, a new command succeeded, and the sandbox was deleted. This verifies provider behavior; it was not repeated for this rebase. - Latest-head CI and automated review are pending. Broad local tests, typecheck, and build were not rerun for this focused rebase; CI supplies those checks. ## Risks - Timing out bridge drain permits provider termination; it never supplies a stop receipt. - A crashed cleanup attempt remains protected for 15 minutes, then becomes eligible again. Repeated failures retry every 30 minutes after escalation. - The existing counter saturates at the escalation threshold; the new attempt identity and deadline prevent overlapping claims. - No schema, UI, telemetry, lockfile, or workflow change. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
|
||
|
|
0c1e7504c0 |
fix(runner): persist warm Daytona workspaces across turns (#12904)
## Thinking Path > - Daytona preserves a stopped sandbox filesystem, but deleting or replacing a sandbox removes its only remote copy. > - Warm reuse therefore improves latency but cannot be Paperclip's durability boundary. > - The host execution workspace must remain authoritative after every successful turn, while same-run recovery must avoid overwriting unexported remote work. > - Result proposal, workspace export/merge, and terminal completion need a durable, replayable ordering so a crash never starts a duplicate provider turn. > - A paid browser acceptance suite must exercise both legacy Codex and Runner Codex for three real turns on one continuously warm Daytona sandbox. ## Linked Issues or Issue Description Refs #12901. Runner Codex did not previously export successful Daytona workspace changes back to the authoritative host workspace. That made warm reuse depend on Daytona's remote filesystem and left deleted/replacement sandboxes without a reliable reconstruction path. The existing paid fixture also lacked a focused three-turn continuity case for both Codex adapters. ## What Changed - Persist versioned, atomic native workspace-sync descriptors and durable seeds in `PAPERCLIP_HOME`, without credentials or a database migration. - Classify fresh, warm, replacement, and same-run-recovery workspace preparation explicitly; ambiguous lease/root/digest evidence fails closed. - Finalize native workspace export/merge after semantic result proposal and before run completion, with idempotent replay that never submits a second provider turn. - Surface legacy Codex workspace restoration failures instead of masking them, while preserving an earlier provider error when both fail. - Keep healthy reusable Daytona leases warm for legacy and native adapters, stamp finalized workspace generations, and retain existing cleanup behavior for per-turn or unhealthy leases. - Preserve Runner Codex's provider process/session across warm turns, including bounded post-terminal tail draining and exact authority rotation. - Add the exact paid `daytona-warm-continuity` matrix: - `legacy-codex × daytona × warm-three-turn` - `runner-codex × daytona × warm-three-turn` - Drive all three turns through the browser, verify ordered file continuity and stable lease/workspace/runtime identities, capture per-turn timings, and delete the sandbox immediately after assertions. - Document `pnpm test:e2e:runner -- --suite daytona-warm-continuity`; no package script was added. ## Verification - `pnpm typecheck` — passed, including migration safety (no migration added) - Focused server/runner Vitest coverage — 144 passed - `pnpm test:e2e:runner:unit` — 114 passed - `pnpm test:e2e:runner:typecheck` — passed - `pnpm --filter @paperclipai/paperclip-runner test:codex` — 66 passed, 1 helper ignored - `native-session-executor.test.ts` — 139 passed, including safe fail-closed cleanup after remote runner identity capture failure - Paid local browser acceptance, exact post-rebase Linux/amd64 runner binary: - Runner Codex — passed in 1.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - Legacy Codex — passed in 2.7m; 3 runs; lease outcomes `created, resumed, resumed`; 10/10 matchers; cleanup passed - [Protected paid GitHub Actions campaign](https://github.com/paperclipai/paperclip/actions/runs/34026735033) against `7da42a91b95fa7fb2df126668ef7e37afb3b2b9d` — passed 2/2: - Runner Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Legacy Codex — 3 runs; lease outcomes `created, resumed, resumed`; evidence and cleanup passed - Merge/enforcement, S3 history, and Pages publication jobs passed - Paid result artifacts were scanned for both provider credentials; neither secret was present. - Current PR checks — 31 passed, 1 expected Storybook skip; Greptile 5/5; Superagent security scan passed - `git diff --check origin/master...HEAD` — passed - Confirmed no `package.json`, lockfile, migration, or SQL changes. ## Risks - Workspace synchronization now sits on the terminal-success path, so a remote export failure deliberately prevents false success. Retryable state retains its lease/seed; loss of the only unexported remote copy fails closed. - Warm provider reuse has strict identity and quiescence checks. Mismatched or ambiguous evidence blocks reuse rather than risking concurrent provider work. - The paid suite incurs Daytona and Codex cost only in the existing protected scheduled/manual workflow and explicitly destroys its sandbox after each cell. ## Model Used OpenAI Codex with GPT-5 agentic reasoning, repository inspection, real browser E2E execution, Rust/TypeScript test execution, and GitHub Actions diagnostics. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing issue or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dedicated suite invocation without adding a package script - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green on the current revision - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups on the current revision - [x] I will address all reviewer comments before requesting merge |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0a422fda52 |
feat(runner): add remote execution substrate (#12638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The current native path runs on the control-plane host. > - Remote environments need an authenticated execution-target contract. > - The contract must not change direct adapters or enable new runtimes by default. > - This pull request adds the remote execution substrate and Daytona ingress. > - The benefit is a bounded base for later remote runner transport work. ## Linked Issues or Issue Description Refs #12616. Refs #12352. **Subsystem affected** Cross-cutting. This change touches runner transport, server orchestration, plugin contracts, and shared settings. **Problem or motivation** Native execution cannot resolve an authenticated runner ingress through a remote environment. The server also lacks one provider-neutral contract for remote execution targets. **Proposed solution** Add a default-off runner preview ingress capability. Add transport-neutral runner connectivity. Add remote execution target and lifecycle handling. Add a Daytona ingress implementation with redacted credentials. **Alternatives considered** A provider-specific server path would duplicate orchestration and authorization. A public endpoint without an environment contract would weaken the trust boundary. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing runs and governed tool access. ## What Changed - Added execution-target traits for local, SSH, and sandbox environments. - Added plugin RPC contracts for runner ingress endpoints. - Added authenticated Daytona preview ingress. - Added transport-neutral PRP outbound connections. - Added remote runner artifact verification and fail-closed provider selection. - Added bounded native session resume, cancellation, and lifecycle recovery. - Preserved Codex-only selection for fresh experimental runner starts. - Preserved all direct adapter execution and finalization paths. - Removed stale Pi provider-pack requirements that security review rejected. - Kept the rollout controls off by default. - Did not change pnpm-lock.yaml, Cargo, database migrations, or GitHub workflows. ## Verification - GitHub Actions will run the repository test, typecheck, build, security, and policy gates. - Focused tests cover ingress validation, redaction, execution targets, remote lifecycle, cancellation, resume, and legacy adapter selection. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check origin/master...HEAD` passes. - The diff contains 52 files. ## Risks - Remote execution crosses a trust boundary. - The implementation validates target capabilities, artifact digests, provider-pack pins, and connection metadata. - The feature remains default-off. - Fresh native selection remains Codex-only. - Existing direct adapters remain on the legacy path. - This PR does not yet make remote Codex runnable. The next PR adds the Rust WSS and TLS transport. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
445547c989 |
feat(duplex): run the Daytona sandbox callback bridge over Node HTTP/2 (#12120)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers carry agent work through controlled execution channels > - The Daytona callback bridge uses a bespoke line-framed protocol over its duplex channel > - The bespoke protocol adds framing work and does not use the Node transport that already supports multiplexed streams > - This pull request carries raw bytes across the channel, adds a Node HTTP/2 bridge, and selects it for Daytona > - The benefit is one authenticated, multiplexed callback session with queue_v1 as the bounded fallback ## Linked Issues or Issue Description **Subsystem affected** The packages/plugins Daytona provider and the shared duplex execution path. **Problem or motivation** The Daytona callback bridge uses a bespoke line-framed protocol over the provider duplex channel. This adds protocol work and limits stream handling. **Proposed solution** Carry raw bytes through the cross-layer channel. Add an authenticated Node HTTP/2 host server and sandbox client gateway. Select http2_v1 for Daytona and retain queue_v1 as the fallback. **Alternatives considered** Keep the current duplex_v1 protocol. This keeps the bespoke framing path and does not provide one HTTP/2 session for callback streams. **Roadmap alignment** ROADMAP.md lists Daytona under cloud and sandbox agents. This change improves the shipped Daytona provider path. **Additional context** The branch adds no dependency. Node 24 provides the http2 module. The host token check and canonical path parser remain the single dispatch path. ## What Changed - Carry raw Uint8Array chunks through the adapter, plugin, worker, runtime, and Daytona layers. - Encode bytes as base64 only across the JSON-RPC hop, because JSON has no binary type. - Add the bounded host HTTP/2 server and the in-sandbox HTTP/2 client gateway. - Authenticate every stream with the per-run bridge token before route work. - Parse the path once and reuse the canonical result for route and forwarding work. - Select http2_v1 for Daytona and fall back once to queue_v1 when the client preface is absent. - Add transport, session, stream, and fallback telemetry. - Mark HTTP/2 as the preferred transport and queue_v1 as the soft-deprecated fallback. ## Verification - `npx vitest run packages/adapter-utils/src` — 990 passed and 4 skipped. - `npx vitest run server/src/__tests__/plugin-worker-manager-duplex.test.ts` — 32 passed. - `npx vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 220 passed and 6 skipped. - `npx tsc --noEmit` in `packages/adapter-utils`, `packages/shared`, `packages/plugins/sdk`, and `server` — clean. - No `package.json` or `pnpm-lock.yaml` file changed. - The live Daytona test skips when `DAYTONA_API_KEY` is absent. - The root `npx tsc --noEmit` command has a pre-existing missing `packages/adapters/droid-local` reference on this branch and on `master`. ## Risks - The transport change affects several duplex layers and could expose byte-boundary errors. - A missing HTTP/2 client preface falls back once to queue_v1 and records `preface_missing`. - The host token check and canonical path parser must remain on the shared dispatch path. - The live Daytona test needs `DAYTONA_API_KEY` and does not run in this agent sandbox. ## Model Used OpenAI GPT-5, tool-enabled coding agent with repository inspection, GitHub CLI, and shell execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b6854e61c7 |
refactor(adapter-utils): rename EffectiveSandboxCapabilities to EffectiveExecutionCapabilities (#12119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter utilities package defines shared types for agent execution targets > - The type name EffectiveSandboxCapabilities describes only one transport > - All execution target drivers return the same resolved capability snapshot > - This pull request gives the snapshot a general name and keeps the old type as a deprecated alias > - The benefit is clearer public vocabulary with source compatibility for current consumers ## Linked Issues or Issue Description **What existing behavior does this improve?** The exported capability snapshot type uses the name `EffectiveSandboxCapabilities`, although local, SSH, sandbox, and plugin drivers return it. **Subsystem affected** The change affects `packages/adapter-utils` and its server consumers. **Current behavior** The public type name points to the sandbox transport. The private parser also uses the sandbox-only name. **Proposed behavior** Use `EffectiveExecutionCapabilities` for the public type and `parseEffectiveExecutionCapabilities` for the private parser. Keep a deprecated alias for the old public type. **Reason and benefit** The new name matches the established execution-target vocabulary. The alias keeps existing type imports working during the migration. **Breaking changes** None. The runtime field, capability flags, parsed shape, and package versions do not change. **Additional context** GitHub search found no duplicate or related open issue or pull request. ## What Changed - Rename the exported interface to `EffectiveExecutionCapabilities`. - Keep `EffectiveSandboxCapabilities` as a deprecated type alias. - Rename the private parser and update its call site and references. - Add a type-level test for the deprecated alias. ## Verification - `npx tsc --noEmit -p packages/adapter-utils` - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` - `npx vitest run server/src/__tests__/environment-execution-target-capabilities.test.ts server/src/__tests__/environment-execution-target-duplex.test.ts` - The local checks passed with 133 adapter-utils tests and 31 server tests. - Reviewers can confirm that the runtime field and capability flags stay unchanged. ## Risks Low risk. The alias protects existing type imports. The change does not alter runtime behavior or serialized data. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a14e51d592 |
refactor(environment): classify environment capabilities from static driver definitions (#12045)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment runtime drivers provide workspace, lease, and custom image behavior > - Runtime code used driver identity checks and several capability-specific members > - These checks spread capability rules across the runtime and made new drivers harder to verify > - This pull request adds one general capability classifier and one static driver support table > - The benefit is one fail-closed capability model that keeps current behavior and supports future drivers ## Linked Issues or Issue Description **What existing behavior does this improve?** Environment runtime capability checks for workspace realization, custom images, lease capabilities, and duplex authorization. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** The runtime selects several capability paths from driver identity and separate capability members. Custom image gates also trust provider declarations without checking every matching live worker method. **Proposed behavior** The runtime uses one general capability classifier and one static support table. Custom image gates require both the provider declaration and every matching live worker method. The public capability names remain unchanged. **Reason and benefit** The change keeps capability rules in one place. It removes identity conditions from runtime consumers and makes unsupported drivers fail closed. **Breaking changes** None. The public names sandboxCapabilities, sandboxProviders, and EffectiveSandboxCapabilities remain available. ## What Changed - Add classifyEnvironmentCapabilities and static support definitions for all four driver families. - Add resolveCapabilities to every environment runtime driver. - Move driver traits into environment-driver-traits.ts and migrate runtime consumers. - Require provider declarations and matching live worker methods for all custom image gates. - Migrate duplex authorization to the general resolver and remove the dead sandbox-only member. - Delete the unused resolveEffectiveSandboxCapabilities wrapper and update its test. ## Verification - pnpm --filter @paperclipai/server typecheck - pnpm exec vitest run server/src/__tests__/environment-capability-contract.test.ts server/src/__tests__/environment-runtime.test.ts — 92 tests pass - pnpm exec vitest run server/src/__tests__/environment-driver-traits.test.ts server/src/__tests__/general-capability-classifier.test.ts — 12 tests pass - pnpm exec vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-execution-target-capabilities.test.ts server/src/__tests__/environment-execution-target-duplex.test.ts server/src/__tests__/environment-execution-target-duplex-kill-switch.test.ts — 66 tests pass ## Risks The main risk is a capability gate that denies a valid driver or permits an invalid driver. The static support matrix, live worker method checks, and regression tests reduce this risk. No database, public API, or published type name changes. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The deployment does not provide a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (for example, docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c62bb4b16b |
feat: environment delete with agent reassignment and consented sandbox destroy (#12053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments define where agent runs execute: local, SSH, or provider sandboxes > - Operators can create and edit environments, but the UI has no way to delete one > - The server already exposes `DELETE /environments/:id` and a delete-blast-radius preflight, but no UI consumes them, and a delete blocked by reusable sandbox leases gives the operator no path forward > - This pull request adds the delete flow to the environment configuration page: a preflight-driven modal that reassigns dependent agents, names the workspaces that hold blocking sandbox leases, and can destroy those sandboxes with explicit consent > - The benefit is that operators can retire stale environments from the UI without database surgery, and dependent agents move to a chosen replacement instead of silently falling back ## Linked Issues or Issue Description Refs #8554 Refs #11124 **Subsystem affected** Environments (server routes, environment runtime service, and the environment settings UI). **Problem or motivation** The environment configuration page has no delete control. The server delete endpoint exists, but nothing in the UI calls it. When reusable sandbox leases block a delete, the 409 error names no owner, so the operator cannot find the blocking workspace. Agents that use the environment as their default lose it silently through the FK `on delete set null`. **Proposed solution** Add a delete button with a confirmation modal on the environment edit page. The modal reads the delete-blast-radius preflight. It offers a dropdown to reassign dependent agents to another environment before the delete. It lists each workspace that holds a blocking reusable sandbox lease, with a link. When those leases are the only blocker, the confirm button destroys the sandboxes inline (`?destroyReusableSandboxLeases=true`) and then deletes. A failed teardown falls back to `pending_cleanup` for the sweep, so no sandbox is orphaned. ## What Changed - `ui/src/pages/CompanyEnvironments.tsx`: delete button on the edit page header, confirmation modal with agent reassignment select, lease-holder list, impact notes, and a consent-labeled destroy-and-delete action - `ui/src/api/environments.ts`: `deleteBlastRadius` and `remove` client methods; `remove` takes an optional `destroyReusableSandboxLeases` flag - `server/src/routes/environments.ts`: `DELETE /environments/:id` accepts `?destroyReusableSandboxLeases=true`; it destroys the environment's reusable sandbox leases first, but only when those leases are the sole delete blocker, then re-checks the blast radius before it deletes - `server/src/services/environment-runtime.ts`: new `destroyReusableSandboxLeasesForEnvironment` — destroys every reusable sandbox lease an environment still owns while the environment config (provider credentials) is still available - `server/src/services/environments.ts`: the delete blast radius now returns `reusableSandboxLeaseHolders` (lease id, workspace, issue) so clients can name what blocks a delete - `packages/shared/src/types/environment.ts`: `EnvironmentDeleteReusableLeaseHolder` type on the blast radius - Tests: route gating for the consent flag (destroy runs, mixed-blocker rejection, surviving-lease rejection), runtime destroy scoped to an environment, blast-radius holder join, and UI tests for the reassignment flow, holder links, and the consent button ## Verification - `npx vitest run server/src/__tests__/environment-routes.test.ts server/src/__tests__/environment-service.test.ts server/src/__tests__/environment-runtime.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - Manual: open Settings → Environments → edit an environment. The trash icon opens the modal. With agents on the environment, pick a reassignment target and confirm; agents move and the environment deletes. With reusable sandbox leases, the modal names the holding workspaces and the confirm button reads "Destroy N sandboxes and delete". ## Risks - The consented path destroys provider sandboxes. It runs only when reusable leases are the sole blocker, so a delete that would still be rejected never destroys anything. A failed teardown routes to `pending_cleanup` and the delete stays blocked until the sweep resolves it. - Agent reassignment issues one PATCH per agent from the client. A mid-sequence failure leaves some agents reassigned; the reassignments are valid on their own and the UI refreshes to the actual state. - Hard blockers (managed local, instance default, pending cleanup) keep the existing 409 behavior and disable the confirm button. ## Model Used - Claude (Anthropic) — Fable 5, model id `claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c5050396c7 |
fix(daytona-duplex): chunk host-to-sandbox writes and make a transport close legible (#11986)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agent work through adapters and sandbox providers > - The Daytona duplex path sends host input through a provider pseudo-terminal WebSocket > - Large messages exceed the provider limit, and a transport close can look like a process exit > - This pull request chunks UTF-8 input and carries transport-close state through the duplex path > - The benefit is reliable large input and accurate loss reporting ## Linked Issues or Issue Description **What happened?** The Daytona duplex path sent a full input payload as one WebSocket message. A payload above the provider limit closed the channel. The wait path also mapped a non-numeric exit result to a process exit without exit data. **Expected behavior** The provider must receive large input as ordered UTF-8 chunks. A transport close without exit data must record `transport_closed`, while a numeric exit must record `provider_exit`. **Steps to reproduce** 1. Start a Daytona duplex session. 2. Send an input payload larger than 65536 bytes. 3. Observe that one message closes the provider channel. 4. End a session without a numeric exit code. 5. Observe that the loss reason reports a process exit. **Paperclip version or commit** Commit `1761e79ec9097c65d94f90a8ba20416f8ab718a6`. **Deployment mode** Built from source with the Daytona sandbox provider. ## What Changed - Add a shared UTF-8 byte chunker with a 32768-byte cap. - Route both Daytona pseudo-terminal write paths through the chunker. - Preserve multi-byte UTF-8 sequences across read-side chunks. - Carry an explicit `transportClosed` state through the worker and host wait paths. - Record `transport_closed` for a reason-less transport close and `provider_exit` for a numeric exit. - Keep orderly completion suppression for both exit paths. ## Verification - The Daytona plugin suite passes 194 tests. - The adapter-utils broker, codec, and telemetry suites pass 73 tests. - The plugin SDK duplex and worker RPC host suites pass 37 tests. - The server plugin worker manager duplex suite passes 78 tests. - The execution target sandbox and ACPX execute suites pass 257 tests. - TypeScript checks pass for adapter-utils, plugin SDK, server, and the standalone Daytona plugin. ## Risks The chunk size adds a loop for large input payloads. The 32768-byte cap stays below the provider limit. The optional loss field preserves compatibility for other providers. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faab2620ad |
feat(sandbox): add a duplex transport for Daytona behind a default-off kill switch (#11750)
## Thinking Path > - Paperclip is an open source app that manages AI agents for work > - Paperclip runs agents in local and remote sandbox environments > - A sandbox needs a bounded channel for commands and asynchronous input > - Daytona needs a real pseudo-terminal transport for this channel > - The sandbox gateway also needs a mode that handles channel loss safely > - This pull request adds the Daytona transport and gateway mode behind a default-off kill switch > - The benefit is a tested foundation for later transport selection ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): sandbox providers, plugin SDK, server settings, and shared types. **Problem or motivation** The merged sandbox protocol has no runtime transport for Daytona. The generated sandbox gateway also has no duplex mode. A later transport-selection change needs both parts and a safe per-run gate. **Proposed solution** Add a Daytona `duplexCommandStream` transport over a raw pseudo-terminal. Add a generated gateway mode named `duplex_v1`. Add the `enableSandboxDuplexBridge` setting with a default value of `false`. Keep transport selection disabled until a later pull request. **Alternatives considered** Keep the protocol unused until the transport-selection change. This would delay provider tests and leave the gateway path without direct coverage. **Roadmap alignment** This change supports the completed Roadmap item for cloud and sandbox agents. It extends the merged sandbox channel foundation in pull request #11738. **Additional context** The Daytona provider remains an untrusted boundary. Deployments must use least-privilege provider credentials and provider-side quota controls. Operators must name an owner for duplex telemetry retention before rollout. ## What Changed - Add the Daytona `duplexCommandStream` capability over a raw pseudo-terminal. - Add a launch wrapper that disables echo and newline translation for NDJSON frames. - Close channels on lease release, destroy, resume of a stopped worker, and worker shutdown. - Declare the capability in the Daytona manifest and set `PLUGIN_VERSION` to `0.1.5`. - Add the worker-to-host notification sink at `ctx.duplexChannel.data` and `ctx.duplexChannel.exit`. - Add the generated sandbox gateway mode `PAPERCLIP_API_BRIDGE_MODE=duplex_v1`. - Add channel-loss results of `409 outcome_indeterminate` and `503 bridge_unavailable`. - Add the per-run setting `enableSandboxDuplexBridge`, with a default value of `false`. - Add unit tests, generated-source codec tests, lifecycle tests, and a credential-gated live Daytona test. ## Verification - Daytona suite: 185 tests pass. - Adapter utilities: 754 tests pass and 4 tests skip. - Plugin SDK: 62 tests pass. - Shared package: 28 tests pass. - Server duplex tests pass. - Shared, plugin SDK, server, and Daytona TypeScript checks pass. - The live Daytona test passes 3 cases when `DAYTONA_API_KEY` is set. - The live Daytona test skips 3 cases without `DAYTONA_API_KEY`. - CI must run the full workspace typecheck, test, and build gates after PR creation. ## Risks - The Daytona control plane and pseudo-terminal remain untrusted boundaries. - The duplex gateway changes behavior only when the mode and per-run setting enable it. - A lost channel fails requests without replay, so callers must handle indeterminate outcomes. - The transport-selection change must require both `duplexCommandStream === true` and `enableSandboxDuplexBridge === true`. - The provider credential and quota limits need operator control before rollout. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8161244284 |
feat(sandbox): add opt-in duplex command-stream foundation (capability, protocol, bounded host route, frame codec) (#11738)
## Thinking Path > - Paperclip provides a control plane for companies that run AI agents. > - Sandboxed agents need a safe execution path for persistent command streams. > - The existing callback transport does not provide a bounded, generic duplex route. > - The host must control capability access, route identity, protocol limits, and close behavior. > - This pull request adds an opt-in duplex command-stream foundation across the sandbox layers. > - The feature stays inert because no current provider declares the capability. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox command execution needs a persistent host-to-sandbox stream. The current callback bridge uses a file transport and does not provide this generic route. **Proposed solution** Add a fail-closed provider capability, generic worker protocol messages, a host-owned bounded route, cross-layer service mediation, and a versioned newline-delimited frame codec. **Alternatives considered** Keep the file transport and add feature-specific commands. This does not provide one reusable duplex contract or host-owned route bounds. **Roadmap alignment** This work supports the completed Cloud / Sandbox agents roadmap area and the safe autonomy goal in the product definition. **Additional context** The change passed a two-stage security review. The final code review verdict was approve after fixes for active-stream bounds and service-layer capability mediation. ## What Changed - Add the opt-in `duplexCommandStream` provider capability with fail-closed narrowing. - Add duplex open, write, stop, and close requests and data and exit notifications to the plugin worker protocol. - Add a host-owned route with bounds for chunk size, cumulative bytes, lifetime, protocol errors, pending requests, and pre-bind buffering. - Add close acknowledgement handling with worker retirement when the close remains unconfirmed. - Wire `openDuplexChannel` through the execution target, runtime service, and plugin worker. - Add a versioned frame codec with shared wire-compatibility vectors and split UTF-8 handling. ## Verification - `server/src/__tests__/plugin-worker-manager-duplex.test.ts` passes 18 tests. - `server/src/__tests__/environment-execution-target-duplex.test.ts` passes 11 tests. - `packages/adapter-utils/src/duplex-frame-codec.test.ts` passes 38 tests. - `server/src/__tests__/sandbox-capability-contract.test.ts` passes 15 tests. - Setup-token pseudo-terminal regression tests pass 47 tests. - Server TypeScript check passes. - Continuous integration will run the full required test, typecheck, build, and policy checks. ## Risks - Providers that opt into the capability must implement the complete worker protocol. - Route limit defaults can close a stream when a workload exceeds the configured bounds. - The capability remains disabled for current providers, so current production behavior does not change. ## Model Used OpenAI GPT-5 (`gpt-5`), with tool use and code execution. The model reviewed and prepared this pull request from the supplied implementation and verification record. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
233be4b36c |
feat: parallelize sandbox file-sync behind a provider opt-in capability (#11736)
## Thinking Path > - Paperclip runs AI agents through local and remote execution adapters. > - Sandbox providers move workspace and asset files before and after agent runs. > - Serial file transfers delay startup and teardown when several operations do not depend on each other. > - Providers need an opt-in contract so existing providers keep their serial behavior. > - This pull request adds a bounded scheduler and routes inbound and outbound sync operations through it. > - The benefit is shorter sandbox setup and teardown with stable errors, clear telemetry, and a safe opt-in path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): packages/shared, packages/adapter-utils, packages/plugins, and server. **Problem or motivation** Sandbox sync processes the workspace, assets, and referenced projects in series. This adds avoidable wait time to agent startup and teardown. **Proposed solution** Add a fail-closed provider capability named concurrentSyncOperations. Use a bounded scheduler with a limit of four operations. Preserve operation order for error reporting. Keep non-opted-in providers on the serial path. **Alternatives considered** Increase the serial transfer speed or add provider-specific schedulers. Those options do not provide one shared contract or stable behavior across providers. **Roadmap alignment** ROADMAP.md lists cloud and sandbox agents as a product area. This change improves sandbox execution without changing the control-plane contract. **Additional context** The Daytona provider opts in. Board trials on this commit showed overlap for inbound sync and outbound restore, with no referenced-project staging failures. ## What Changed - Add the concurrentSyncOperations sandbox capability and fail-closed parsing. - Add a bounded settle-all scheduler with stable input-order errors. - Parallelize inbound workspace, asset, and referenced-project sync operations when the provider opts in. - Parallelize outbound workspace and asset restore operations when the provider opts in. - Surface referenced-project failure text in run logs and server telemetry. - Add Daytona sync spans and the capability declaration. - Preserve in-flight upload scratch tarballs during workspace wipe. - Add unit and regression tests for the scheduler, coordinators, provider behavior, telemetry, and wipe race. ## Verification - Run the adapter-utils and server type checks. - Run the targeted adapter-utils, server, and Daytona test suites. - Run the full automated sweep. - Review six cold Daytona trials, with three serial and three parallel runs. - Confirm that parallel trials show inbound overlap and outbound restore overlap. - Confirm that providers without the capability keep serial behavior. ## Risks - Providers must opt in only when their file operations can run safely at the same time. - A provider that declares the capability incorrectly can expose transfer races. - The scheduler keeps a limit of four to bound resource use. - Providers without the capability keep the prior serial behavior. ## Model Used OpenAI GPT-5 in the Codex runtime. The model used tool calls, code inspection, and GitHub workflow support. The model did not author the implementation commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c1c46f1e4e |
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports subscription login through a sandbox > - The new-agent page must show login before the user creates an agent > - Test results must not expose raw sandbox diagnostics or secret values > - This pull request adds the login UI to both Test lanes and closes the diagnostic boundary > - The branch also adds durable cleanup recovery for failed sandbox teardown > - Reusable sandboxes must retain both their recorded teardown configuration and a valid lifecycle path until destruction succeeds > - The benefit is a usable login flow with fixed public checks, redacted server logs, and recoverable sandbox cleanup ## Linked Issues or Issue Description Related public work: [#9488](https://github.com/paperclipai/paperclip/pull/9488) adds first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and remote runs. Related public issue: [#2681](https://github.com/paperclipai/paperclip/issues/2681) requests Claude Code subscription support. This pull request adds the login transport and new-agent UI flow that those changes do not provide. **Subsystem affected:** Claude local adapter, server login probes, sandbox provider setup, cleanup recovery, and the new-agent UI. **Problem or motivation:** The Test lanes did not show the sandbox login panel in all supported cases. Test results also exposed raw probe diagnostics, and JSON escapes could end secret redaction early. **Proposed solution:** Surface the login capability through the bundled provider manifest. Prepare the same probe runtime in the ACP lane. Send diagnostics only to redacted server logs. Keep Test checks on fixed public messages. Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. Consume JSON escapes during redaction. Preserve failed sandbox cleanup state across retries and restarts, and prevent deletion from severing the lifecycle context of a live reusable sandbox. **Alternatives considered:** Keep raw diagnostics in Test checks or trust login URL text from the sandbox. Both choices increase information exposure. Keep separate probe behavior in the ACP lane. That choice would leave the two Test lanes inconsistent. ## What Changed - Surface the sandbox login panel on both Test lanes. - Reconcile the bundled Daytona plugin manifest so `supportsSetupTokenLogin` reaches the UI capability gate. - Prepare the ACP Test lane with the same probe runtime as the CLI Test lane. - Add the `claude_acp_login_probe_unavailable` warning when the ACP probe cannot run. - Send raw sandbox diagnostics only to redacted server logs. - Keep Test checks on fixed public messages in the ACP, managed-config, and CLI paths. - Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. - Redact JSON and escaped-JSON secret values, including escaped quotes and backslashes. - Preserve orphan cleanup records across provider failures, restarts, and unavailable plugins. - Atomically block environment deletion while a live reusable sandbox lease still depends on it. - Verify pending cleanup destroys plugin sandboxes with the provider configuration recorded on the lease, even after the current environment configuration changes. ## Verification - Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`. - Focused environment route/service/runtime coverage passes: 196 tests across 3 files. - `pnpm -r typecheck` passes. - `pnpm build` passes. - The full Vitest run completed with 4,754 passing and 28 failing tests. All 23 source-test failures reproduce unchanged on parent head `58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate executions from stale `server/dist` output. The failures are unrelated macOS path/listener and scheduler-fixture failures, so there is no new bad commit for bisect to localize. - All required CI checks pass for the current head, including build, typecheck/release registry, all server and workspace shards, serialized server suites, canary, and e2e. - A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241` reports 5/5, “safe to merge,” with no blocking failure remaining. ## Risks - A probe or redaction change could hide useful server diagnostics. - An allowlist change could reject a valid Claude login URL. - Cleanup recovery changes could affect provider teardown ordering. - An environment with a live reusable sandbox can no longer be deleted until the owning issue or execution workspace completes teardown. - The implementation keeps public Test messages fixed and sends detail to redacted server logs. ## Model Used OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code execution enabled; extended reasoning enabled. The implementation author used AI-assisted development. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the result - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation or confirmed no separate documentation change is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3061ce6901 |
feat(sandbox): stream session output by capability, drop three operator flags (#11557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandboxed agents use provider capabilities to select safe execution paths > - Session output still depends on three operator flags that duplicate capability data > - Duplicate flags can drift from the verified sandbox capability snapshot > - This pull request makes the capability snapshot the only streaming decision and removes the obsolete flags > - The benefit is default streaming with a poll fallback when a capability or stream fails ## Linked Issues or Issue Description **What existing behavior does this improve?** ACP sandbox session-output streaming and sandbox execution configuration. **Subsystem affected** Cross-cutting (multiple of the above): server/, packages/shared/, packages/adapter-utils/, and packages/plugins/. **Current behavior** Session-output streaming requires operator flags in the server and Daytona plugin configuration. Saved configurations can retain a removed key. **Proposed behavior** The verified capability snapshot selects streaming. The Daytona plugin uses persistent sessions by default, keeps bypass commands one-shot, and falls back from the log stream to polling. Removed configuration keys become inert. **Reason and benefit** One capability source prevents configuration drift. The fallback keeps output available when capability resolution or log streaming fails. **Breaking changes** The three operator flags no longer control session-output streaming. Existing saved keys load but have no effect. ## What Changed - Remove `useSessions` and `useLogStream` from the Daytona plugin configuration and manifest. - Remove `streamAgentSessionOutput` from server configuration, shared types, and execution-target plumbing. - Select streaming from `persistentProcessSessions` and `independentControlCommands`. - Keep poll fallback on capability resolution failure and stream failure. - Strip removed keys from strict fake-sandbox and catchall plugin configuration. - Update the sandbox capability documentation and focused tests. ## Verification - `tsc --noEmit` passed in `packages/shared`, `packages/adapter-utils`, `server`, and the Daytona plugin. - Daytona `plugin.test.ts` passed 139 tests. - Server capability, configuration, route, and runtime suites passed 160 tests. - `packages/adapter-utils` `execution-target-sandbox.test.ts` passed 44 tests. - The capability matrix covers stream, poll, and resolution-failure paths. - Removed-key tests cover strict fake-sandbox and catchall plugin schemas. ## Risks - A capability snapshot that lacks either required session capability uses polling. - A log stream failure uses polling and can increase request count. - Existing removed configuration keys no longer change behavior. - The isolated-worktree Daytona Vitest run has a pre-existing missing `packages/adapters/droid-local` reference. CI and standard checkouts use the committed configuration. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact runtime context window is managed by the Codex platform. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e71ce9a9d3 |
feat: sandbox provider capability contract with fail-closed effective resolution (#11463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs work through adapters and sandbox providers > - Providers need a clear contract so the server can use only verified capabilities > - A declared capability must not grant a method that the live worker did not verify > - This pull request adds manifest declarations and fail-closed effective capability resolution > - The benefit is safe provider reuse across execution targets and run lifecycles ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox providers expose different runtime methods. The server needs one safe capability contract that accounts for provider declarations, worker verification, and narrowing configuration. **Proposed solution** Add strict manifest validation for five sandbox capabilities. Resolve effective capabilities as the subset of verified, declared, and narrowed values. Store the result as a frozen execution-target snapshot. **Alternatives considered** Trusting the manifest alone could grant methods that the worker does not support. Trusting only a fixed built-in list would reject valid third-party providers. The intersection rule keeps the verified runtime ceiling and supports both provider types. **Roadmap alignment** This change supports the ACP run lifecycle track and the sandbox provider contract work in the current roadmap. **Additional context** The legacy `supportsReusableLeases` field remains supported. The nested capability validator rejects unknown keys. Missing or unavailable verification resolves all capabilities to `false`. ## What Changed - Add strict `sandboxCapabilities` manifest validation with legacy reusable-lease compatibility. - Carry declarations through the ready-driver projection. - Add fail-closed effective resolution from verified, declared, and narrowed capabilities. - Add narrowing for provider configuration, Kubernetes Job leases, and Daytona sessions. - Add a frozen read-only capability snapshot to execution targets. - Add focused tests and keep existing characterization baselines covered. - Add and update sandbox provider capability documentation. ## Verification - `npx vitest run packages/shared/src/validators/plugin.test.ts` - `npx vitest run server/src/__tests__/plugin-environment-driver-sandbox-capabilities.test.ts` - `npx vitest run server/src/__tests__/sandbox-capability-contract.test.ts` - `npx vitest run server/src/__tests__/environment-execution-target-capabilities.test.ts` - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts` - Package typechecks for shared, server, and adapter-utils pass. - Stage-2 security review suites pass with 28 tests. ## Risks The resolver fails closed when verification is absent or unavailable. Providers that rely on undeclared capabilities may see narrower behavior until they expose verified worker methods. The change does not alter the existing native-sync guard. ## Model Used OpenAI Codex, GPT-5, exact runtime model ID `gpt-5`, tool use and code execution. The implementation author used this model to assist with the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e52b8a343f |
fix: ACP run lifecycle corrections — failure settlement, workspace sync-back, lease cleanup (#11454)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run ACP sessions and manage runtime, workspace, and lease resources. > - Several failure paths left runtime bridges, staged workspaces, or environment leases active after an error. > - These leaks reduce run reliability and can leave later runs without clean resources. > - This pull request closes the failure paths, applies one teardown policy, and adds regression tests. > - The benefit is consistent failure settlement and safer reuse of agent workspaces and leases. ## Linked Issues or Issue Description **What happened?** ACP runs could leave runtime bridges, staged workspaces, or environment leases active after failures. Claude and Gemini ACP runs did not restore the sandbox workspace on teardown. Lease release stopped when one lease returned an error. **Expected behavior** Each ACP failure must return an error result and settle its resources. Teardown must run each step, release leases independently, and restore the host workspace when the sandbox ends. Pending cleanup leases must receive bounded retry attempts. **Steps to reproduce** 1. Run an ACP session that fails after runtime creation or during turn preparation. 2. Run an ACP session that fails during a warm hit or staged runtime handoff. 3. Run lease cleanup with more than one lease when the first release returns an error. 4. Inspect the result phase, teardown calls, workspace state, and lease metadata. 5. Run the regression suites listed in the Verification section. ## What Changed - Settle every ACP failure after runtime creation with an error result and one sandbox.startup span closure. - Close the ACP runtime and remove warm entries after every pre-turn failure. - Run all teardown steps, record teardown errors, release staging leases in finally, and prevent duplicate teardown. - Dispose staged runtimes after seam failures and remove borrowed staged entries with identity guards. - Add fail-open workspace sync-back teardown for Claude and Gemini ACP adapters. - Isolate lease release errors and add bounded retry sweeps for stranded pending_cleanup leases. - Atomically claim pending_cleanup retries and clamp attempt readers to keep the five-attempt bound. - Default absent provider reusableLeases values to false and align the fake provider with its runtime declaration. - Add regression tests for engine, adapter, server, and shared environment behavior. ## Verification - [x] `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 124 tests passed. - [x] `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/gemini-local/src/server/acp.test.ts` — 61 tests passed. - [x] `npx vitest run server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts server/src/__tests__/reusable-leases-default.test.ts server/src/__tests__/environment-routes.test.ts packages/shared/src/environment-support.test.ts` — passed. - [x] All listed suites ran from the repository root. - [x] GitHub CI completed successfully for `cfc349c9f232711433897915112a1c52c0e462ca`. - [x] Greptile completed with a 5/5 confidence score and no blocking finding. ## Risks The engine changes affect failure settlement and teardown order across ACP runs. The server changes add retry state to existing lease metadata without a schema migration. The adapter changes restore workspaces after sandbox execution. Regression tests cover the changed paths. GitHub CI and Greptile passed for the current head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This change fixes runtime reliability and does not duplicate a roadmap feature. ## Model Used OpenAI GPT-5 Codex. The model used tool-based repository inspection, GitHub operations, and code review support. The runtime does not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b7e0814a0 |
feat(acp): stream Daytona sandbox agent output and remove the host output poll (#11049)
## Thinking Path > - Paperclip is the open source app that manages AI agents for work > - Sandbox providers let agents run in remote and isolated environments > - Daytona session commands need a path that sends agent output to the host without host polling > - Host polling adds delay and repeats provider output work > - This pull request adds typed execute.log notifications and a log sink for incremental output > - This pull request adds an optional ACP session stream with final-result replay protection > - The benefit is lower output delay while the default flags keep current behavior unchanged ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. The change spans the plugin SDK, Daytona provider, adapter utilities, and server execution services. **Problem or motivation** The Daytona ACP bridge polls a host output file while an agent command runs. This adds delay and can repeat work. The host also needs a safe route for provider output chunks. **Proposed solution** Add a typed `execute.log` notification with host-issued invocation correlation. Add an ordered log sink to the environment execute path. Add an optional ACP session-log path that parses newline-delimited JSON frames and removes the host output poll for that path. **Alternatives considered** Keep the output-file poll as the only path. This keeps the current behavior but does not provide timely output. The new path stays behind flags, so the existing path remains the default fallback. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`, including Daytona support. ## What Changed - Add the typed `execute.log` worker-to-host notification and company-scoped host route. - Add ordered `stdout` and `stderr` chunk delivery before the final execute result. - Add the Daytona session log sink and the optional ACP streamed session path. - Add monotonic frame handling so live and final output reach the host once. - Keep `useLogStream` and `streamAgentSessionOutput` off by default. - Add unit and integration coverage for the notification, execution target, runtime, and Daytona paths. ## Verification - Run adapter-utils tests: 445 tests pass locally. - Run server environment tests: 73 tests pass locally. - Run Daytona plugin tests: 131 tests pass locally. - Run TypeScript checks for shared, adapter-utils, and server. - Review the pull request checks after GitHub completes them. - All required GitHub checks pass on the current head. ## Risks The new paths change output delivery only when a feature flag enables them. The final execute result remains available for parsing and fallback. The main risk is a provider stream or frame-order error; the final-result parser limits that risk. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The runtime did not supply a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes No operator documentation change applies because both new flags remain disabled by default. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cfed36ea6b |
feat(plugin-daytona): persistent session model with plain command dispatch (#10941)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - One core subsystem runs agent work inside sandboxes > - The Daytona provider uses that path to run user commands > - The current one-shot model does not keep a shell alive across commands > - This pull request adds an opt-in persistent session model for Daytona > - The benefit is faster command dispatch with the same sandbox boundaries ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting. This change touches `packages/adapters`, provider tests, span names, and sandbox command behavior. ### Problem or motivation The Daytona provider needs a persistent shell for repeated command dispatch. The old advisory wrapper path does not reach that goal. It also adds cost and removes the session speed gain. ### Proposed solution Add a `useSessions` driver flag. Keep it off by default. Open one Daytona session per lease when the flag is on. Send each user command into that session. Read stdout and stderr from the session logs endpoint. Run each command in a subshell so `exit` does not stop the shell. Remove the advisory `bwrap` wrapper path and its lease metadata. Add session setup and teardown spans. Keep a hard delete on teardown. ### Alternatives considered Keep the advisory `bwrap` wrapper. That path does not give a real persistent session. It also keeps extra command overhead. Keep a one-shot fallback for user commands. That would weaken the session model and hide a missing session case. ### Roadmap alignment This work fits the `Cloud / Sandbox agents` milestone in `ROADMAP.md`. It also supports the control plane goal of safe remote sandbox execution. ### Additional context The handoff verification reported `tsc --noEmit` clean and 119 Daytona unit tests passing. The handoff also reported a clean host span allowlist test and five expected commits on the branch. The security review gate remains required before merge. ## What Changed - Added an opt-in persistent session model for the Daytona sandbox provider. - Routed user commands through `executeSessionCommand` when sessions are enabled. - Removed the advisory `bwrap` command wrapper path and the lease metadata it used. - Added session lifecycle spans and span allowlist coverage. - Documented the leak bound in `DIRECTORY-CONSTRAINT-FINDINGS.md`. ## Verification - `tsc --noEmit` clean for the Daytona plugin, per handoff verification. - Daytona unit suite passes, with 119 tests, per handoff verification. - Host span allowlist test passes, per handoff verification. ## Risks - Persistent sessions can leak if teardown fails. - Session logs must keep stdout and stderr separate. - The flag stays off by default to limit rollout risk. ## Model Used OpenAI GPT-5, Codex, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53bcf3897f |
feat: sync @-mentioned projects into remote sandboxes (confined sandbox transport, flag ON) (#10564)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The run layer must move project context into the sandbox that executes the agent > - Local sandboxes already stage referenced projects for @-mentions > - Remote confined sandboxes dropped the whole referenced set, so the agent lost needed files and paths > - This pull request keeps the confined sandbox transport aligned with the local behavior for referenced projects > - It does this behind a remote-only flag that defaults on, while SSH keeps the old drop-only path > - The benefit is that remote runs can read the same referenced project context that local runs already provide ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change touches server orchestration, sandbox transport, and observability. **Problem or motivation** A run can @-mention another project. Local targets stage each referenced project and give the agent a path. Remote confined sandboxes dropped the full referenced set, so the agent could not read those project files or paths. **Proposed solution** Enable referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`, which defaults on. Keep the SSH transport out of scope and keep it dropping referenced projects. Repoint each referenced workspace hint at its staged `project-<projectId>` sandbox directory. Publish `PAPERCLIP_WORKSPACES_JSON` on the confined sandbox lane. Count each per-project remote staging failure as a `staging` failure in the requested-vs-synced metrics. **Alternatives considered** Keep the remote path drop-only. That keeps the gap open. Move the change into SSH too. That expands scope beyond the target transport and adds risk. **Roadmap alignment** No matching item in `ROADMAP.md` showed up in this review. **Additional context** The change lands in three commits. The first commit opens the gate for the confined sandbox transport. The second commit repoints the workspace hints and publishes the workspace map. The third commit records per-project staging failure data. ## What Changed - Opened remote referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`. - Repointed referenced workspace hints to the staged `project-<projectId>` sandbox directories and published `PAPERCLIP_WORKSPACES_JSON`. - Counted per-project remote staging failures as first-class `staging` failures in the requested-vs-synced observability. ## Verification - The pushed ref `refs/heads/feat/sync-referenced-projects-remote-sandbox` resolves to the authorized submit SHA. - `git log --oneline origin/master..origin/feat/sync-referenced-projects-remote-sandbox` shows exactly the three expected commits. - The handoff reports server typecheck clean, adapter-utils typecheck clean, and the listed unit tests passing. - The handoff also reports no open review comments and no Greptile score yet. ## Risks - The change touches authorization and sandbox path handling, so regressions could block remote runs or expose the wrong project context. - The new flag defaults on, so any bug in the remote path affects normal remote use. - SSH stays out of scope, so the two transport paths must remain distinct. ## Model Used OpenAI GPT-5 via Codex. Tool use enabled. Context window not reported in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7083c275c8 |
refactor(sandbox): retire the dead noProfile flag from the exec path (#10461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The sandbox exec path starts agent commands and passes runtime options to the server and plugin layers > - This path kept a noProfile flag after the exec wrappers stopped sourcing a login profile > - The flag no longer changed behavior, so it left dead API surface in the protocol and runtime helpers > - This pull request removes that dead flag from the plugin protocol, the server drivers, and the managed-runtime helpers > - It also updates the tests and points the agent runtime README at the sandbox requirements file > - The benefit is a smaller and clearer exec-path contract with no behavior change ## Linked Issues or Issue Description - No public GitHub issue exists. ### What happened? The sandbox exec path kept a `noProfile` field after the exec wrappers stopped sourcing a login profile. ### Expected behavior The plugin protocol, server drivers, and managed-runtime helpers should not expose or forward a dead field. ### Steps to reproduce 1. Run a managed-runtime command through the sandbox exec path. 2. Inspect the protocol payload and runtime helper inputs. 3. Observe that `noProfile` is present even though it no longer changes behavior. ### Paperclip version or commit `60c7da86fc7a6c1dbf37bbcd86e25ecaaff01607` ### Deployment mode Built from source (pnpm dev / pnpm build) ### Additional context This pull request removes the dead field, updates the affected tests, and updates the README note for the sandbox profile path. ## What Changed - Removed noProfile from the plugin protocol and the server exec-path call sites. - Updated the managed-runtime helpers to use the narrower exec-path contract. - Updated the affected tests and added the README pointer to SANDBOX-REQUIREMENTS.md. ## Verification - `git grep -n "noProfile" -- packages/ server/` returns zero matches. - `tsc --noEmit` passed for `@paperclipai/adapter-utils`, `@paperclipai/plugin-sdk`, and `@paperclipai/server`. - `command-managed-runtime.test.ts` passed: 22/22. - `environment-runtime.test.ts` passed: 24/24. ## Risks - Low risk. The flag was already a no-op. - A hidden external caller may still send the removed field. ## Model Used - OpenAI GPT-5, tool-enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7797995038 |
perf(plugin-daytona): opt-in no-profile fast path for default-PATH execs (#10352)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - The Daytona adapter turns tasks into shell commands and manages execution overhead > - Many short-lived exec calls still pay for login-shell profile sourcing even when the binary already resolves on the sandbox default PATH > - That extra startup work adds latency on the hot path for repeated command execution > - This pull request adds an opt-in fast path that skips profile sourcing only when the caller explicitly requests it and the command does not need shell initialization > - The benefit is lower per-call latency for eligible commands without changing the conservative default behavior for commands that need the profile ## Linked Issues or Issue Description This change does not reference a public GitHub issue. It follows the same Daytona startup-speed work as merged PR #10335 and narrows the execution path for eligible commands while keeping the default login-shell behavior intact. ## What Changed - Added an optional `noProfile` flag to `PluginEnvironmentExecuteParams`. - Refactored Daytona login-shell script assembly so the profile and nvm sourcing block is omitted only on the explicit fast path. - Preserved environment prefixing, `cd`, shell quoting, `NONINTERACTIVE_GIT_ENV`, stdin handling, and `durationMs` behavior on both paths. - Added regression tests for the fast path omission, the preserved execution parameters, and the default profile-sourcing path. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run src/plugin.test.ts` - `pnpm --filter @paperclipai/plugin-sdk tsc --noEmit` - Reverted the guard locally to confirm the two behavior tests fail again, then restored the change. ## Risks - If a caller opts into `noProfile` for a command that depends on shell initialization, the command can fail to resolve its binary. - The API comment and opt-in design keep that risk narrow; the default path remains unchanged. ## Model Used OpenAI GPT-5 (Codex tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7f766526a6 |
feat(sandbox): add task-scoped egress grants (#10155)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9edde68373 |
feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - AI agents run in sandboxed execution environments (Kubernetes pods, Daytona workspaces, etc.) and need to sync files between the host and those environments — for workspace setup, asset delivery, and output retrieval > - The existing sync path for Kubernetes uses a base64-over-exec chunk loop: each ~4 MB chunk requires its own `execInPod` round-trip, so large syncs balloon into many exec calls with corresponding overhead > - `execInPod` supports piped stdin/stdout, meaning the full transfer can be done as a single exec that streams a raw `tar` archive over the data channel — one round-trip regardless of file size, with nothing base64-encoded and nothing buffered whole in memory on either side > - PR-1 (#10013, merged) added the `onEnvironmentSyncIn`/`onEnvironmentSyncOut` opt-in hook API to the sandbox provider interface and documented the protocol; PR-2 (#10028, merged) implemented these hooks for the Daytona provider > - This pull request implements the same two lifecycle hooks in the Kubernetes sandbox provider, so workspace/asset file sync streams through one `execInPod` per operation instead of the chunk loop > - The benefit is significantly fewer exec round-trips for large syncs and flat memory use on both host and pod, with security properties preserved: atomic replace, secret-mode enforcement, path confinement, TOCTOU-safe snapshot, and member-confinement on host-assembled archives from sandbox-authored tar output ## Linked Issues or Issue Description This is the third and final PR in a sequential series: - Refs #10013 — PR-1: opt-in sync hook API + provider docs (merged) - Refs #10028 — PR-2: native file-sync lifecycle hooks for Daytona provider (merged) **Feature:** Native single-exec file-sync lifecycle hooks for the Kubernetes sandbox provider. *Motivation:* The existing Kubernetes sync path encodes files as base64 and loops over `execInPod` one chunk at a time (~4 MB per exec). For large workspaces or asset sets this is slow and resource-intensive. The Kubernetes `execInPod` API supports piped stdin/stdout, enabling a raw-`tar` streaming transfer that needs only one exec regardless of file count or size and never buffers the whole payload in memory. *Proposed solution:* Implement `onEnvironmentSyncIn` and `onEnvironmentSyncOut` in the Kubernetes provider using a streaming `execInPod` with a tar pipeline — for syncIn the host builds the archive on disk and streams its raw bytes into the pod's stdin (`head -c <exact-size> | tar -x`, no base64); for syncOut in-pod `tar` writes to the exec's stdout and the host streams those bytes straight to a file. Path confinement, atomic replace, secret-mode enforcement, TOCTOU protection, and a streamed-bytes fail-closed guard are all enforced. ## What Changed - **New `src/file-sync.ts`** in `packages/plugins/sandbox-providers/kubernetes/` — `performSyncIn` and `performSyncOut` over an injected pod-exec closure, keeping transfer logic hermetically unit-testable - **New `execInPodStreaming` in `src/pod-exec.ts`** — a streaming exec primitive that binds a caller-supplied stdin readable and a stdout writable to the exec WebSocket data channel, added alongside the existing `execInPod` (which is unchanged). This lets a transfer stream raw bytes to/from disk instead of buffering the payload as a single string - **Updated `src/plugin.ts`** — registers `onEnvironmentSyncIn`/`onEnvironmentSyncOut`; resolves the `sandbox-cr` pod exactly like `onEnvironmentExecute` and delegates; `job` backend rejects file-sync calls explicitly (out of scope) - **syncIn path:** host builds the tarball to a temp file → streams its raw bytes over exec stdin, bounded in-pod by `head -c <exact-archive-size> | tar -x` (no base64 anywhere) → extract into a `/proc/self/fd`-pinned reserved `0700` staging dir → `chmod`-before-`mv -f` atomic replace per file (directory mappings use `followSymlinks`→`-h`) - **syncOut path:** in-pod validate + realpath-snapshot each source (closes the validation→copy TOCTOU window) → single-exec `tar -c` streamed over exec stdout → host streams that stdout straight to a temp file through a byte-counting transform → member-confined extraction of the sandbox-authored archive - **Security properties:** secret files land at requested mode with no widened window; every interpolated path is shell-quoted and confined lexically plus via in-pod `realpath`; the outbound stream is bounded by a **streamed-bytes disk guard** (`MAX_SYNC_OUTPUT_BYTES`, 8 GiB default, per-call overridable) that fails the transfer closed — writing no target file — if an untrusted pod emits more bytes than allowed. Neither host nor pod buffers the whole payload, so there is no in-memory size cap on the transfer - **No changes** to `execInPod`, `wrapCommandWithEnv`, or `FastUploadInterceptor` (the `environmentExecute` path is untouched) - **No dependency or lockfile changes** - **New tests** in `test/unit/file-sync.test.ts` (atomic-replace, `0600` secret mode, symlink preserve/deref, dir-mapping, exclude, path-confinement rejection, streamed-output guard fail-closed) and `test/unit/pod-exec.test.ts` (streaming stdin/stdout, caller-sink error fail-closed), plus extended `test/unit/plugin.test.ts` ## Follow-up: Legacy Job-Lease Base64 Fallback Fix Addresses the Greptile 4/5 blocking finding ("Handle existing job leases", `server/src/services/environment-runtime.ts`). Job leases provisioned before the `nativeFileSyncUnsupported` metadata flag existed carry `backend: "job"` but no flag, so `supportsSync()` treated them as native-capable and routed their sync to the pod-exec hook — which the job backend rejects (it has no exec channel) instead of using the byte-identical base64 fallback. The fix adds a belt-and-suspenders gate on the persisted `backend === "job"` field alongside the existing `nativeFileSyncUnsupported` flag check, so pre-existing job leases continue syncing via the base64 fallback after deployment. No behaviour change for `sandbox-cr` leases. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-kubernetes test` — 19 files / 182 tests green, including the existing `upload-interceptor` and `pod-exec` suites - `tsc --noEmit` in the kubernetes package — 0 errors - The sync hooks are opt-in; existing `environmentExecute` behaviour is unaffected and tested by the unchanged existing suites ## Risks - **Opt-in only:** `onEnvironmentSyncIn`/`onEnvironmentSyncOut` are registered conditionally; providers that do not register them fall back to the existing chunk loop. No regression risk on the existing path. - **Shell-injection surface:** all path interpolation uses shell-quoting; paths are additionally confined lexically and via in-pod `realpath` before use. - **TOCTOU on syncOut:** the in-pod snapshot validates and records file metadata before the tar call, closing the window between validation and copy. - **Archive member confinement:** host-side reassembly rejects any tar member whose resolved path escapes the target directory, preventing a malicious in-pod tar from writing outside the intended destination. - **Untrusted-output volume:** an over-large outbound stream trips the streamed-bytes disk guard and fails closed (no target written and the temp sink is swept) rather than filling host disk or memory; the guard bounds disk unconditionally and bounds memory insofar as WebSocket write-backpressure holds. ## Model Used Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`) — produced by a Claude-based AI agent using agentic tool use and multi-step code generation. 200K context window, extended reasoning, code execution and verification capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b247bf7150 |
feat(runtime): opt-in sandbox file-sync lifecycle hooks (API + provider docs) (#10013)
## Thinking Path > - Paperclip is an open-source AI-agent management platform; agents run tasks inside sandboxed environments (Daytona, Kubernetes, E2B, etc.) > - The control-plane ↔ sandbox file-transfer path flows through the `environmentExecute` seam in `protocol.ts` — the only verb available to plugins — which forces a base64-over-exec chunked loop for every file move: workspace files, assets, Codex home sync > - This transport is correct and safe, but it bypasses provider-native bulk/streaming APIs (Daytona `uploadFiles`, K8s `FastUploadInterceptor` / volume mounts), leaving significant throughput on the table for large workspaces > - The right fix is an opt-in seam extension: providers with faster native transfer declare two optional verbs; providers that do not opt in stay on the existing fallback with zero code or behavior change required > - This PR adds the first layer of that extension — two optional verbs (`environmentSyncIn` / `environmentSyncOut`) in the plugin SDK, the runtime plumbing to prefer the native path for the two clean destroy-then-replace cases, and a doc for the contract > - The core correctness invariant is byte-identical fallback: if no provider opts in, execution is exactly what ships today; `assertSyncOperationsConfined` enforces host-side path confinement for providers that do opt in > - No provider advertises the verbs yet → zero production behavior change; future PRs wire up Daytona and K8s providers against this contract ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Description follows the `feature_request` issue template: **Subsystem affected:** packages/plugins — plugin system; packages/adapter-utils — adapter runtime; server/ — EnvironmentRuntimeService **Problem or motivation:** Sandbox file transfers currently always use a base64-over-exec chunked loop regardless of what the underlying provider supports. For workspaces larger than a few MB this becomes the dominant wall-clock cost of every sandbox run, and it bypasses bulk/stream APIs that providers like Daytona already expose natively. **Proposed solution:** Add two optional, opt-in plugin hooks — `onEnvironmentSyncIn` / `onEnvironmentSyncOut` — to the plugin SDK. When a provider defines both hooks and both are advertised via the existing `supportedMethods` negotiation, the runtime prefers the native path for the two clean destroy-then-replace transfer cases; all other cases fall back to the existing byte-identical base64 transport. **Alternatives considered:** An unconditional verb would require every provider to implement or stub the verb. The opt-in / `METHOD_NOT_IMPLEMENTED` pattern (already used by `environmentExecute`) preserves backward compatibility with zero provider changes required. **Roadmap alignment:** Consistent with the ✅ "Cloud / Sandbox agents" and ✅ "Plugin system" milestones; extends the plugin seam rather than adding control-plane-level logic. **Additional context:** Searched open pull requests and issues for duplicate sandbox file-sync / native-transfer work; none found. ## What Changed - **`packages/plugins/sdk`** - `protocol.ts`: two new optional `HostToWorkerMethods` — `environmentSyncIn` / `environmentSyncOut` — plus generic `SyncOperation`, `SyncFileMapping`, and `SyncOutcome` types - `define-plugin.ts`: optional `onEnvironmentSyncIn` / `onEnvironmentSyncOut` fields on `PluginDefinition`; worker advertises each verb only when its hook is defined (else `METHOD_NOT_IMPLEMENTED`, mirroring `environmentExecute`) - `worker-rpc-host.ts`: route new verbs to plugin hooks - `index.ts`: re-export new public types - **`packages/adapter-utils`** - `command-managed-runtime.ts`: expose optional `syncIn` / `syncOut` on `CommandManagedRuntimeRunner` (available only when both verbs are advertised); add `assertSyncOperationsConfined` host-side path-confinement guard - `sandbox-managed-runtime.ts`: `SandboxManagedRuntimeClient` gains optional `syncIn` / `syncOut`; orchestrator prefers native path for default-provision asset inbound and workspace-download-into-fresh-dir outbound; all other paths keep the existing base64 fallback - `sandbox-file-sync.test.ts` (new): 234-line characterization suite — native-opt-in branch, fallback branch, `assertSyncOperationsConfined` escape-path rejection, `followSymlinks` → tar `-h` - `command-managed-runtime.test.ts`: negotiation + native-sync + confinement tests - **`server/src/services/environment-runtime.ts`**: `EnvironmentRuntimeService` delegates to `syncIn` / `syncOut`, gated on advertised support - **`server/src/services/environment-execution-target.ts`**: minor typing fix alongside the new verbs - **`doc/plugins/SANDBOX_FILE_SYNC_HOOKS.md`** (new): documents the full contract — opt-in / no-op guarantee, operation ordering, provider-may-tar, atomicity, `followSymlinks`, secret modes (0600, no window), path confinement, `operationId` opacity, resource bounds, shell-quoting ## Verification ```bash # SDK suite pnpm --filter packages/plugins/sdk test # Adapter-utils suite (includes new sandbox-file-sync characterization tests) pnpm --filter packages/adapter-utils test # Expected: 255 pass / 4 skip # Type-check across affected packages pnpm --filter packages/plugins/sdk typecheck pnpm --filter packages/adapter-utils typecheck # Server changed-file spot check: cd server && npx tsc --noEmit --skipLibCheck 2>&1 | grep -E "environment-(runtime|execution-target)" | head -20 ``` Key behavioral invariant to spot-check: with no provider opting in (the current state), run any sandbox task and confirm file-transfer behavior is byte-for-byte identical to what the pre-PR code produces. The characterization tests assert this at the unit level. ## Risks - **Zero production risk today**: no provider advertises `environmentSyncIn` / `environmentSyncOut`, so the new code paths are unreachable in production; all real traffic stays on the existing base64 fallback - **Path confinement**: `assertSyncOperationsConfined` rejects any `targetPath` that escapes the declared root — this is the primary security boundary for future providers. The test suite covers escape-path rejection - **Atomicity**: the contract delegates atomicity to providers; the doc explicitly calls out that directory-level ops are not guaranteed atomic - **Secret transport**: credential assets (e.g., Codex `auth.json`, directory mappings) continue to use the existing tar path — they do not go through the new verbs in any current provider > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: Anthropic Model: `claude-sonnet-4-6` (Claude Sonnet 4.6) Context window: 200 K tokens Capabilities: extended tool use, multi-file code generation, agentic reasoning via the Paperclip agent framework ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a8f0ebaa80 |
Refresh run config before reusing workspaces (#8797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs are assembled by the heartbeat service from agent config, project workspaces, environment config, secret bindings, skills, and runtime session state. > - The heartbeat service intentionally reuses adapter sessions, execution workspaces, and sandbox leases when that preserves useful state. > - Reuse becomes incorrect when the effective next-run config changes after a saved session, workspace, or lease was created. > - Stale reuse can make a later run appear pinned to old agent, environment, secret, instruction, or workspace settings. > - This pull request records non-sensitive fingerprints for the effective session, workspace, and lease config at run boundaries. > - When those fingerprints drift, Paperclip refreshes persisted runtime config or starts fresh execution instead of reusing stale state. > - The benefit is predictable next-run config freshness without storing raw secret values, full env maps, provider credentials, or private path details. ## Linked Issues or Issue Description - Refs #8058 - Related PRs checked during dedup search: #4968, #4155, #84, #8480. These cover nearby workspace/session routing or model-config freshness areas, but do not duplicate this effective run config fingerprinting path. ## What Changed - Added effective run config fingerprinting for session, workspace, and lease reuse decisions, with canonicalization that ignores generated runtime noise and redacts sensitive values. - Updated heartbeat reuse logic to compare stored and next-run fingerprints, reset stale saved sessions, refresh persisted workspace config snapshots, replace stale reused workspaces when required, and avoid stale sandbox lease reuse. - Included plain environment value drift via value hashes, without storing the raw env values. - Root-bound instruction content hashing so legacy direct absolute instruction paths are represented but not read for config fingerprints. - Batched secret/version metadata lookups for environment lease fingerprinting. - Added workspace operation/run result freshness metadata so operators can inspect non-sensitive decision categories. - Surfaced config freshness labels and next-run copy in the UI and docs. - Added focused coverage for fingerprint redaction, session reset decisions, workspace refresh/replace behavior, environment lease drift, and persisted workspace restoration. ## Verification - `git diff --check` - Sensitive-data scan before push: - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "(AWS_ACCESS_KEY_ID|AWS_SECRET_ACCESS_KEY|ghp_[A-Za-z0-9_]{20,}|github_pat_[A-Za-z0-9_]{20,}|sk-[A-Za-z0-9]{20,}|-----BEGIN (RSA |OPENSSH |EC |DSA )?PRIVATE KEY-----|AKIA[0-9A-Z]{16})"` - `git diff --unified=0 origin/master...HEAD | rg -n --pcre2 "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}"` - `pnpm exec vitest run server/src/__tests__/effective-run-config-fingerprints.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/environment-runtime.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm -r typecheck` - `pnpm --filter @paperclipai/db clean` - `pnpm test:run` - `pnpm build` - UI screenshots from Cutter: - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-01.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-02.png - https://artifacts.cutter.sh/8797/run-2f4827c-2026-06-30T18-57-25/preview/change-03.png ## Risks - Medium: overly broad fingerprints could start fresh sessions, workspaces, or sandbox leases more often than necessary. - Medium: missing a config category would allow stale reuse to persist for that category. - Medium: legacy direct absolute instruction paths are no longer content-hashed unless they are paired with an absolute managed instructions root. - Low data risk: fingerprint metadata stores hashes and category names, not raw secrets, raw env values, provider credentials, or private path details. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex CLI / Codex coding agent, tool-enabled with shell, Git, GitHub CLI, local test execution, and code editing. The exact deployed model variant and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.ing> |
||
|
|
8d9f9fd240 |
Add reusable sandbox custom images (#8794)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A growing part of that work runs in sandboxed environments rather than on the operator's local machine. > - Today sandbox providers can start fresh workspaces and run probes, but they do not have a shared contract for capturing and reusing prepared sandbox state. > - Operators need a way to set up tools, credentials, and project dependencies once, then reuse that prepared image for later agent runs. > - This pull request adds reusable sandbox custom images across the provider contract, server runtime, and board UI. > - It also keeps probes and sandbox copy flows aligned with pre-authenticated/custom-image environments. > - The benefit is faster, more reliable sandbox runs without repeatedly rebuilding the same environment setup. ## Linked Issues or Issue Description No public GitHub issue was found for this change. Inline feature request follows. ### Problem or motivation Sandboxed agents need reusable prepared runtime state so repeated runs do not require manual setup every time. Operators often need system packages, CLIs, SDKs, dependency caches, credentials, and project tooling available before an agent can work productively. ### Proposed solution Add a provider-level custom-image capability, server-side setup/capture lifecycle, Daytona/fake provider support, and board UI controls for creating, testing, selecting, and deleting custom images. ### Alternatives considered Leaving this as provider-specific setup outside Paperclip would keep the control plane blind to image state and would not give agents consistent environment metadata. Re-running setup commands for every lease is simpler, but slower and less reliable for interactive or credentialed setup. ### Roadmap alignment Checked `ROADMAP.md`; this aligns with the Cloud / Sandbox agents roadmap area and does not duplicate any related public issue or PR found by search. Additional context: - Subsystem affected: cross-cutting (`packages/db`, `packages/shared`, `packages/plugins`, `server`, `ui`). - Duplicate search: searched GitHub for `sandbox custom image` and `sandbox template environment`; no related public issues or PRs were found. ## What Changed - Added custom-image shared types, validators, constants, API paths, and database schema/migration. - Added server services/routes for custom-image templates and setup sessions, including runtime cleanup and provider metadata handling. - Extended plugin/sandbox provider capabilities for interactive setup, template capture, and template deletion. - Implemented custom-image support in the fake sandbox provider and Daytona provider. - Updated environment runtime/config handling so active custom images flow into leases, probes, and agent execution. - Added board UI controls and API client support for custom-image setup, capture, selection, status, and error states. - Hardened sandbox copy/probe behavior for insecure clipboard contexts and pre-authenticated sandbox images. - Added targeted coverage across shared validators, DB schema, server routes/services, provider plugins, adapter probes, and UI flows. ## Verification - `pnpm install --frozen-lockfile --ignore-scripts` - `pnpm vitest run packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/test.ts packages/adapters/codex-local/src/server/test.remote.test.ts packages/adapters/codex-local/src/server/test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm vitest run packages/db/src/environment-custom-images-schema.test.ts packages/shared/src/environment-custom-images.test.ts packages/shared/src/validators/plugin.test.ts server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/workspace-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` - `pnpm -r typecheck` - `pnpm build` - `rm -rf packages/db/dist && pnpm test:run` - Public-safety scan of the final diff found no internal Paperclip issue links, private instance URLs, or real secret patterns. ## Risks - Adds a database migration and new environment runtime tables, so migration ordering and rollback need care. - Provider implementations may differ in how reliably they can capture/delete images; unsupported providers surface capability-gated UI states. - Custom-image state can contain operator-prepared tooling and credentials inside the provider image, so providers must enforce their own access controls and cleanup semantics. - Broad surface area across shared contracts, server runtime, plugins, adapters, and UI means CI and Greptile review should be watched closely. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (`gpt-5`) via Codex CLI with tool use and code execution. Assisted with branch cleanup, conflict resolution, local verification, and PR preparation. Earlier branch implementation work was assisted by Paperclip-managed Claude/Codex agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd38c150b0 |
feat: reuse Daytona sandbox leases (#8513)
Add opt-in reusable Daytona sandbox lease support, including retryable pending cleanup handling.\n\nPR: https://github.com/paperclipai/paperclip/pull/8513 |
||
|
|
547463d3a2 |
refactor(environments): make execution environments instance-scoped (#8375)
## Thinking Path > - Paperclip is the control plane for AI-agent companies, so execution environment selection has to stay inspectable and predictable across companies, agents, and runs. > - The environment subsystem decides where an agent heartbeat actually runs and how remote sandbox state is realized and restored. > - That subsystem previously mixed company-scoped environment catalogs with issue-level environment stamping, so a reassigned issue could keep executing in the previous assignee's sandbox. > - That behavior breaks the control-plane contract: changing the assignee should change the executing agent/environment path unless there is an explicit current override. > - Fixing it cleanly required more than a narrow patch; the environment model had to move to instance scope with a single inherited default and per-agent override semantics. > - This pull request rewires the schema, server/API surface, runtime resolution, and UI around that model, then adds regression coverage for cross-company inheritance and per-agent isolation. > - The benefit is that environment choice now follows the approved instance/agent configuration path instead of stale issue state, while shared environments only need to be configured once per instance. ## Linked Issues or Issue Description - No directly matching public GitHub issue or PR was found while searching for this refactor. ### What happened? Reassigning work between agents with different execution environments could keep running in the previous sandbox because environment choice was stamped onto the issue and outranked the current assignee. The same subsystem also forced environment catalogs to be duplicated per company even though the underlying execution environments were instance-wide resources. ### Expected behavior Execution should resolve through the current instance and agent configuration path, with one instance-scoped environment catalog, one instance default, optional per-agent override, and no stale issue-level environment authority surviving reassignment. ### Steps to reproduce 1. Configure two agents to use different execution environments. 2. Assign an issue to the first agent so the issue records execution state in that environment. 3. Reassign the same issue to the second agent and run another heartbeat. 4. Observe that the pre-fix runtime can still sync or execute in the original sandbox instead of the second agent's environment. ### Paperclip version or commit Current `master` before this PR. ### Deployment mode Self-hosted server. ### Installation method Built from source (`pnpm dev` / `pnpm build`). ### Agent adapter(s) involved - Claude Code - Not adapter-specific (core bug in environment authority / resolution) ### Database mode External Postgres. ### Access context Both board reassignment and agent heartbeats were involved. ## What Changed - Moved environments and their default selection contract to instance scope in DB/shared types, including the migration that dedupes legacy per-company environments and seeds the instance local default. - Reworked environment CRUD/auth flows to use instance-scoped APIs and added route/service coverage for instance-level environment management. - Changed runtime resolution to prefer `agent default -> instance default -> built-in local`, removed issue-level environment stamping from the active execution path, and isolated sandbox/plugin leases by `(executionWorkspaceId, agentId)`. - Added environment env-var runtime precedence so environment-provided values act as the baseline for agent execution. - Moved the environment UI into instance settings and updated agent configuration surfaces to reflect inherit/override behavior. - Added regression coverage for instance-default inheritance across companies and for the new runtime resolution behavior. - Fixed a rebase-only duplicate `enableTaskWatchdogs` flag regression in instance settings types/validators/services so the branch typechecks cleanly on current `master`. - Updated stale server tests so CI matches the shipped instance-scoped environment contract. ## Verification - `git diff --check` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run server/src/__tests__/environment-runtime-driver-contract.test.ts server/src/__tests__/agent-permissions-routes.test.ts server/src/__tests__/environment-routes.test.ts server/src/__tests__/environment-instance-routes.test.ts server/src/__tests__/execution-workspace-policy.test.ts server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/instance-settings-routes.test.ts` ## Risks - The migration changes environment scope and dedupes existing rows, so installs with unusual legacy environment combinations should be reviewed carefully during upgrade. - Remote execution behavior now depends on instance-default inheritance semantics instead of issue-level stamping, so any remaining code paths that still assume issue-scoped environment authority would surface as follow-up bugs. - This PR includes both server/runtime behavior and UI relocation, so reviewers should watch for authorization edge cases around instance settings and environment management. > I checked [`ROADMAP.md`](ROADMAP.md). This work fits the existing Cloud / Sandbox agents direction as a bug-fix/refactor to current behavior, not a new parallel product surface. ## Model Used - OpenAI Codex coding agent in this Paperclip/Codex session; GPT-5-class tool-using model with code execution and shell access. The exact backend model ID is not exposed to the session runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4ad94d0bde |
feat(server): kubernetes execution integration for sandbox-provider plugins (stage 2/3) (#7938)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The execution subsystem runs those agents in environments (local, ssh, sandbox), and sandbox-provider plugins let an environment materialize per-run sandboxes > - Stage 1 (#5790) contributed a first-party Kubernetes sandbox-provider plugin, but the server core has no way to adopt it operationally: no per-run adapter selection, no way to force an instance onto sandboxed execution, no declarative adapter/model configuration, and the plugin must be installed by hand > - Without this, a multi-tenant or security-conscious deployment cannot guarantee that agent runs never execute on the host, and a single environment cannot serve agents with different harnesses > - This pull request adds the server + SDK integration: per-run adapterType on the lease protocol, an env-gated forced-Kubernetes execution policy with provisioning and a per-run allowlist guard, a declarative adapter registry and model list, in-cluster env passthrough for sandbox plugin workers, fail-safe auto-install of the bundled plugin, and the matching UI affordance > - The benefit is that sandbox-provider plugins become fully usable for Kubernetes execution: operators configure everything via environment variables and GitOps, while self-hosters who set none of the variables see exactly the behavior they have today ## Linked Issues or Issue Description Refs #5790 (stage 1 of 3: the Kubernetes sandbox-provider plugin package). No existing issue. Feature description: the server core lacks the integration seams to operate a sandbox-provider plugin as the mandatory execution path of an instance. This PR is stage 2 of 3 of the staged Kubernetes contribution; stage 3 will contribute the agent runtime images and their build pipeline. ## What Changed One line per piece: - `packages/plugins/sdk/protocol.ts`: optional `adapterType` on `PluginEnvironmentAcquireLeaseParams` so a provider can select the runtime image per run; existing providers simply ignore it - `server/services/environment-runtime.ts` + `environment-run-orchestrator.ts`: thread the agent's adapter type into both lease-acquiring drivers, including the heartbeat path (the two call sites have historically drifted, hence the pinned test) - `server/services/environments.ts`: `ensureKubernetesEnvironment` / `findKubernetesEnvironment`, an idempotent managed Kubernetes environment per company, identified by a metadata marker and refreshed (not recreated) on config change; `timeoutMs` rides on the config for slow cold-start leases - `server/services/execution-allowlist.ts`: pure (driver, provider, policy) -> allow/deny guard; `executionMode=kubernetes` only allows the kubernetes sandbox provider - `server/services/execution-policy-bootstrap.ts` + startup hook in `server/index.ts`: parse `PAPERCLIP_EXECUTION_MODE` / `PAPERCLIP_K8S_*`, persist `executionMode` into instance general settings, and provision the managed environment for every company; fails loud on misconfiguration - `server/services/heartbeat.ts`: when the policy forces Kubernetes, pin run selection to the managed environment (also overriding any persisted workspace environment id), refuse to fall back to local, and re-check the actually acquired environment against the allowlist as defense in depth - `server/services/adapter-registry-bootstrap.ts` + shared `AdapterRegistryEntry` type/validator: declarative `PAPERCLIP_ADAPTERS` registry (inline JSON or file) that reconciles adapter availability at startup and rides on the Kubernetes environment config - `server/services/adapter-models-env.ts` + `adapters/registry.ts`: `PAPERCLIP_ADAPTER_MODELS` lets an operator declare picker model lists the server cannot CLI-discover - `server/services/plugin-loader.ts`: pass `KUBERNETES_SERVICE_HOST/PORT(_HTTPS)` through to plugin workers that register environment drivers, so in-cluster API clients can be constructed; all other host env stays stripped - `server/app.ts`: fail-safe auto-install of the bundled kubernetes plugin at boot; no-ops when the bundle is absent and never blocks startup on error - `packages/shared` types/validators: `InstanceExecutionMode` on general settings (optional, strict schema) - `ui/lib/forced-kubernetes-environment.ts` + `AgentConfigForm`: when the policy is active, show a read-only Kubernetes environment instead of the environment picker and default new agents onto the managed environment - Tests for every new module plus the adapterType pin in `heartbeat-plugin-environment` and the managed-environment lifecycle in `environment-service` Everything is gated: with `PAPERCLIP_EXECUTION_MODE`, `PAPERCLIP_ADAPTERS`, and `PAPERCLIP_ADAPTER_MODELS` unset (and no bundled plugin present), every code path reduces to current behavior. The per-run `adapterType` is an optional SDK parameter that existing providers ignore. ## Verification - `cd server && npx tsc --noEmit`: clean (0 errors); `ui` typecheck also clean - Targeted suites all green (11 files, 90 tests): `npx vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-service.test.ts server/src/__tests__/environment-runtime.test.ts server/src/__tests__/environment-run-orchestrator.test.ts server/src/__tests__/plugin-database.test.ts server/src/services/execution-policy-bootstrap.test.ts server/src/services/execution-allowlist.test.ts server/src/services/adapter-registry-bootstrap.test.ts server/src/services/adapter-registry-bootstrap.reconcile.test.ts server/src/services/adapter-models-env.test.ts packages/shared/src/validators/adapter-registry.test.ts` - `npx vitest run ui/src/components/AgentConfigForm.test.ts`: green (6 tests) - Full `npx vitest run server/src/__tests__`: 2323 passed, 1 skipped; the only failures (heartbeat-process-recovery pid-retry, workspace-runtime symbolic-ref/git tests) reproduce identically on pristine `master` in the same environment, so they are machine-environment issues unrelated to this change; `server-startup-feedback-export` needed its `services/index.js` mock extended with the new export and is green - This integration has been running in production on a hosted multi-tenant deployment, where it executes agent runs across five different harnesses through the stage 1 plugin ## Risks - Low for existing deployments: every behavior is env-gated and the defaults preserve current semantics; the auto-install block is wrapped fail-safe and skips silently when the plugin bundle is absent - `executionMode` is a new optional field on a strict zod schema; absent input normalizes exactly as before - The forced policy intentionally fails runs loudly (rather than falling back to local) when no managed Kubernetes environment exists; this only affects instances that explicitly set `PAPERCLIP_EXECUTION_MODE=kubernetes` ## Model Used Claude Opus 4.8 (claude-opus-4-8, 1M context), extended thinking, agentic tool use via Claude Code. ## UI screenshots The UI change is a new read-only "Execution" section in `AgentConfigForm`, shown only when the instance execution policy forces Kubernetes (`executionMode=kubernetes`); there is no "before" state for it (the section did not exist, and instances without the forced policy render the existing picker unchanged). Captured from the new Storybook stories added in this PR (`Product/Agent Management`): Managed Kubernetes environment present (read-only display, no local/SSH picker):  No managed environment available yet (warning notice, no silent local fallback):  ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
486fb88a15 |
Add Cloudflare sandbox provider plugin (#5687)
> _Stacked on top of #5685 → #5686. Diff against master includes commits from earlier PRs in the stack — review focuses on the two new commits (`Extend sandbox callback bridge for Worker-hosted plugins` + `Add Cloudflare sandbox provider plugin`)._ ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Each agent runs in a sandbox environment, and operators choose which provider backs that sandbox — today E2B and Daytona are bundled with the platform > - Cloudflare Workers + Durable Objects + the Sandbox SDK offer a credible new option: globally distributed, cheap idle, and operator-deployable as a single Worker > - To plug it in, Paperclip needs (a) a provider plugin that speaks the `PaperclipPluginManifestV1` lifecycle and (b) a small operator-deployed Worker — the **bridge** — that adapts Paperclip's runtime RPCs to the Cloudflare Sandbox SDK > - The plugin extends the existing sandbox-callback-bridge with a `bridge.transport: "worker"` discriminator so the platform routes runtime RPCs through the Worker bridge instead of the in-process runner > - This pull request adds the plugin, the bridge Worker template, and the supporting adapter-utils + server hooks the new transport needs > - The benefit is that operators can run sandboxes on Cloudflare's edge with no new platform code beyond installing the plugin and deploying the Worker ## What Changed **Shared support (`Extend sandbox callback bridge for Worker-hosted plugins`):** - `packages/adapter-utils/src/sandbox-callback-bridge.{ts,test.ts}`: expose `expectedHostHeader` so plugin-side bridge clients can verify the canonical request envelope before forwarding. - `packages/adapter-utils/src/command-managed-runtime.{ts,test.ts}`: relax the always-fresh runner construction so callers can re-use a runner across exec calls (Worker-hosted bridges hold the runner inside a Durable Object). - `server/src/services/environment-runtime.ts` + `environment-runtime.test.ts`: route Worker-hosted bridges through the same env-shaping path as E2B and pin the `requestEnv` contract. - `server/src/services/plugin-environment-driver.ts`: thread an optional `issueId` through the runtime descriptor so bridges can scope leases to the originating issue (used by Cloudflare to map a sandbox to the issue/workflow for billing and audit). - `packages/plugins/sdk/src/protocol.ts`: add `issueId?` to `PluginEnvironmentDriverBaseParams` and the new `bridge.transport: "worker"` discriminator that the new plugin declares. - `server/__tests__/heartbeat-plugin-environment.test.ts`: pin the heartbeat path against the new runtime descriptor. **The Cloudflare plugin itself (`Add Cloudflare sandbox provider plugin`):** - `packages/plugins/sandbox-providers/cloudflare/`: plugin entry, manifest, plugin runtime (lifecycle + bridge client), config parsing, and Vitest coverage. Manifest declares `bridge.transport: "worker"` so the platform routes runtime RPCs through the bridge client. - `bridge-template/`: a Worker template the operator deploys with `wrangler`. Owns Durable Object-backed sessions (`sessions.ts`), exec/stream routes (`exec.ts`, `routes.ts`), and an HMAC auth layer (`auth.ts`) that pins the `Host` header surface. Includes the SDK-contract-correct exec implementation, lease recovery, and chunked stdout/stderr streaming. - Tests cover lease/session handoff (`bridge-template/src/exec.test.ts`, `routes.test.ts`), bridge client request shaping (`src/bridge-client.test.ts`), and end-to-end plugin behavior (`src/plugin.test.ts`) including streamed exec output. 27 tests in total. - `README.md` walks the operator through deploying the bridge Worker, registering the plugin, and configuring the runtime. ## Verification - `pnpm typecheck` - `pnpm exec vitest run --no-coverage packages/adapter-utils/src/sandbox-callback-bridge.test.ts packages/adapter-utils/src/command-managed-runtime.test.ts server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-plugin-environment.test.ts` - `(cd packages/plugins/sandbox-providers/cloudflare && pnpm test)` — 27 passing For an operator-side smoke test: 1. Deploy the bridge: `cd packages/plugins/sandbox-providers/cloudflare/bridge-template && wrangler deploy` 2. Register the plugin in your Paperclip instance, point its bridge URL at the deployed Worker, set the HMAC shared secret. 3. Create a sandbox environment whose provider is `cloudflare`, then run a Codex or Claude job against it. ## Risks - Adds a new `bridge.transport: "worker"` code path, but the existing E2B / Daytona transports go through the same shaped helpers and have explicit test coverage that pins their behavior unchanged. - The Worker bridge stores session state in a Durable Object; operator instances must be aware of the corresponding Cloudflare costs (DO requests, storage). Documented in the README. - The `issueId` plumbing is optional throughout — existing plugins that don't supply it continue to work. ## Model Used - Provider: Anthropic - Model: Claude Opus 4.7 (1M context) - Capabilities used: extended reasoning, tool use (Read/Edit/Bash/Grep) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots — N/A, no UI change - [x] I have updated relevant documentation to reflect my changes (plugin README, bridge-template README) - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
778e775c35 |
Add secrets provider vaults and remote import (#5429)
## Thinking Path > - Paperclip orchestrates AI-agent companies and needs secrets handling to work across local development, hosted operators, and governed agent execution. > - The affected subsystem is the company-scoped secrets control plane: database schema, server services/routes, CLI workflows, and the Secrets settings UI. > - The gap was that secrets were local-only and operators could not manage provider vaults or import existing remote references without exposing plaintext. > - This branch adds provider vault configuration plus an AWS Secrets Manager remote-import path while preserving company boundaries, binding context, and audit trails. > - I kept the PR to a single branch PR, removed unrelated lockfile/package drift, rebased the full branch onto the current `public-gh/master`, and addressed fresh Greptile findings. > - The benefit is a reviewable implementation of provider-backed secrets with focused tests covering provider selection, import conflicts, deleted secret reuse, rotation guards, and AWS signing behavior. ## What Changed - Added provider vault support for company secrets, including provider config storage, default vault handling, health checks, binding usage, access events, and remote import preview/commit. - Added an AWS Secrets Manager provider using SigV4 request signing, bounded request timeouts, namespace guardrails, cached runtime credential resolution, and external-reference linking without plaintext reads. - Added Secrets UI surfaces for vault management and remote import, plus CLI/API documentation for setup and operations. - Stabilized routine webhook secret binding paths and SSH environment-driver fixture bindings discovered during verification. - Addressed Greptile and CI findings: no lockfile/package drift, monotonic migration metadata, disabled-vault default races, soft-deleted secret hiding/recreate behavior, remove behavior with disabled vaults, soft-deleted external-reference re-import, non-active rotation guards, managed-secret soft deletion through PATCH, and per-call AWS SDK credential client churn. - Rebased this branch onto `public-gh/master` at `0e1a5828` and force-pushed with lease to keep this as the single PR for the branch. ## Verification - `git fetch public-gh master` - `git rebase public-gh/master` - `git diff --name-only public-gh/master...HEAD | grep '^pnpm-lock\.yaml$' || true` confirmed `pnpm-lock.yaml` is not in the PR diff. - Confirmed migration ordering: master ends at `0081_optimal_dormammu`; this PR adds `0082_dry_vision` and `0083_company_secret_provider_configs`. - Inspected migrations for repeat safety: new tables/indexes use `IF NOT EXISTS`; foreign keys are guarded by `DO $$ ... IF NOT EXISTS`; column additions use `ADD COLUMN IF NOT EXISTS`. - `pnpm -r typecheck` passed before the Greptile follow-up commits. - `pnpm test:run` ran the full stable Vitest path before the Greptile follow-up commits; it completed with 3 timing-related failures under parallel load: `codex-local-execute.test.ts`, `cursor-local-execute.test.ts`, and `environment-service.test.ts`. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/codex-local-execute.test.ts src/__tests__/cursor-local-execute.test.ts src/__tests__/environment-service.test.ts` passed on targeted rerun (`24/24`). - `pnpm build` passed before the Greptile follow-up commits. Vite reported existing chunk-size/dynamic-import warnings. - After Greptile follow-up commits: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` passed (`26/26`). - After Greptile follow-up commits: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/aws-secrets-manager-provider.test.ts src/__tests__/secrets-service.test.ts` passed (`39/39`). - After Greptile follow-up commits: `pnpm --filter @paperclipai/server typecheck` passed. - Captured Storybook screenshots from `ui/storybook-static` for visual review. - Latest PR checks on `5ca3a5cf`: `policy`, serialized server suites 1/4-4/4, `Canary Dry Run`, `e2e`, `security/snyk`, and `Greptile Review` pass; aggregate `verify` is still registering the completed child checks. - Greptile review loop continued through the latest requested pass; all Greptile review threads are resolved and the latest `Greptile Review` check on `5ca3a5cf` passed with 0 comments added. ## Screenshots Before: the provider-vault and remote-import surfaces did not exist on `master`; these are after-state screenshots from the Storybook fixtures.    ## Risks - Migration risk: this adds new secret provider tables and extends existing secret rows. The migrations were checked for monotonic ordering and idempotent guards, but reviewers should still inspect upgrade behavior carefully. - Provider risk: AWS support uses direct SigV4 requests. Automated tests cover signing, request timeouts, vault-config selection, namespace guardrails, pending-version archival, sanitized provider errors, and service-level cleanup paths. A real-vault AWS smoke test remains deployment validation for an operator with AWS credentials rather than an unverified merge blocker in this local branch. - UI risk: the Secrets page and import dialog are large new surfaces; screenshots are included above for reviewer inspection. - Verification risk: the full local stable test command hit parallel-load timing failures, although the exact failed files passed when rerun directly. - Operational risk: remote import intentionally avoids plaintext reads; operators must understand that imported external references resolve at runtime and may fail if AWS permissions change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent with local shell/tool use in the Paperclip worktree. Exact context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
5c2f9aba9d |
Run explicit-environment adapter tests on the requested target instead of falling back to the host (#5277)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - When a user clicks "Test" on a configured environment (SSH or sandbox), the agent-test route exercises the adapter against that target > - The route previously fell back to running the probe on the Paperclip host whenever an explicit environment target couldn't be resolved, with the test report still saying "passed" > - That hid two real failure modes: misconfigured environments looked green, and sandbox environments were never actually exercised > - This pull request acquires an ad-hoc lease and realizes a workspace for sandbox/plugin test environments, resolves a sandbox execution target wired to the environment runtime, and returns synthesized diagnostics instead of running a host probe when an explicit env target can't be resolved > - The benefit is the Test action surfaces the real environment state and never silently exercises the wrong machine ## What Changed - `server/routes/agents.ts`: acquire an ad-hoc lease and realize a workspace for sandbox/plugin test environments; resolve a sandbox execution target wired to the environment runtime - Return synthesized diagnostics (no host fallback) when an explicit env target can't be resolved - `server/services/environment-runtime.ts`: small adjustments to support the explicit-env-target case - Clarify test-route messages so they no longer claim a host fallback in explicit env flows - New `agent-test-environment-routes.test.ts` covers the guard and missing-environment path ## Verification - `pnpm vitest run --no-coverage server/src/__tests__/agent-test-environment-routes.test.ts` - `pnpm typecheck` clean - Manual: a deliberately misconfigured sandbox environment now reports diagnostics instead of a misleading host-pass ## Risks Medium — Test route behavior change. Explicit environments that previously appeared to pass via host fallback will now report their real state. This is the desired behavior, but operators should expect to see new failures for environments that were never actually working. ## Model Used Claude Opus 4.7 (1M context) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable — new tests cover guard + missing-env paths - [x] If this change affects the UI, I have included before/after screenshots — N/A (no UI) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
076067865f |
Migrate SSH environment callback to bridge (#5116)
> **Stacked PR (part 3 of 7).** Depends on: - PR #5114 - PR #5115 > Diff against `master` includes commits from earlier PRs in the stack — the new commit in this PR is the topmost one. ## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Agents executing on a remote SSH-backed environment need a way to call back into > the Paperclip control plane (run events, log streaming, signals) > - When the SSH host can't reach the Paperclip host (NAT, firewalls, or simply not > on the same network), the run silently fails or hangs — a recurring class of > failure during SSH testing > - In sandboxed environments we already solved this with a callback bridge that > tunnels back through the existing connection; SSH was the odd one out > - This PR migrates SSH execution to use the same callback bridge, so every > adapter's remote run uses one consistent reverse-channel. Per-adapter SSH glue > is deleted in favour of a shared `CommandManagedRuntimeRunner` built from the > SSH spec > - The benefit is fewer SSH-specific failure modes, a smaller code surface, and > one place to evolve the callback contract going forward ## What Changed - Added `createSshCommandManagedRuntimeRunner` in `packages/adapter-utils/src/ssh.ts` that adapts an SSH spec into a generic command-managed-runtime runner (with cwd, env, and timeout handling) - Removed `paperclipApiUrl` from `SshRemoteExecutionSpec`; the bridge URL now flows through the shared runner - Reworked `execution-target.ts` to use the SSH runner alongside sandbox runners via a unified `CommandManagedRuntimeRunner` interface - Simplified `remote-managed-runtime.ts` and `sandbox-managed-runtime.ts` to consume the shared runner abstraction - Deleted per-adapter SSH callback wiring from claude-local, codex-local, cursor-local, gemini-local, opencode-local, pi-local execute.ts files - Removed `environment-runtime-driver-contract.test.ts` (the contract is now enforced by `environment-execution-target.test.ts`) - Added/updated `execute.remote.test.ts` cases for each adapter to cover the SSH runner path ## Verification - `pnpm --filter @paperclipai/adapter-utils test` - `pnpm test -- execute.remote` (covers all six local adapters' SSH paths) - Manual QA: ran a claude-local agent against an SSH-backed environment, confirmed the agent successfully called back to `/api/agent-callback/*` endpoints during the run ## Risks - Refactor touches all six local adapters. If any adapter had subtle SSH-specific behaviour that wasn't captured in tests, it could regress. Mitigation: each adapter's `execute.remote.test.ts` was extended. - `paperclipApiUrl` removal from `SshRemoteExecutionSpec` is a breaking type change for any internal consumer. Verified no external plugins consume this type. - The new `CommandManagedRuntimeRunner` shape is a public surface in `@paperclipai/adapter-utils`; downstream plugins implementing custom runners may need updates, but no such plugins exist in this repo. ## Model Used - OpenAI GPT-5.4 (reasoning effort: high) via Codex CLI - Provider: OpenAI - Used to author the code changes in this PR ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots — N/A - [ ] I have updated relevant documentation to reflect my changes — N/A - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a7b45938b7 |
Let sandbox providers declare shell defaults (#5114)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - Agents execute in sandboxed remote environments served by pluggable
sandbox
> providers (E2B today, more later)
> - Today every sandbox command runs under `sh -lc` regardless of what
the
> provider's container actually ships
> - That misses bash-only shell init on E2B (which ships bash) and
prevents
> future providers from declaring a different default — there's no way
for a
> provider to say "I have bash, use it"
> - This PR adds a `shellCommand` field to sandbox execution targets so
providers
> can declare their preferred shell ("bash" for E2B), threads it through
the
> sandbox-managed-runtime client, callback bridge, and execution-target
shell
> helper, and validates the value at the lease-metadata boundary
> - The benefit is that sandbox commands run under the right shell on
the right
> provider, and adding new sandbox providers only needs to declare a
shell
> preference
## What Changed
- Added `packages/adapter-utils/src/sandbox-shell.ts` exporting
`preferredShellForSandbox(shellCommand)` (returns `"bash"` if input is
`"bash"`,
else `"sh"`)
- Added `shellCommand?: "bash" | "sh" | null` to
`AdapterSandboxExecutionTarget`
and `CommandManagedRuntimeSpec`; threaded it through
`runAdapterExecutionTargetShellCommand`,
`prepareAdapterExecutionTargetRuntime`,
and `startAdapterExecutionTargetPaperclipBridge`
- `createCommandManagedRuntimeClient`, `prepareCommandManagedRuntime`,
and
`createCommandManagedSandboxCallbackBridgeQueueClient` now take an
optional
`shellCommand` and use `preferredShellForSandbox` to pick the shell
- `startSandboxCallbackBridgeServer` accepts a `shellCommand` for its
server
startup, readiness probe, and stop hook
- E2B sandbox plugin declares `shellCommand: "bash"` in `leaseMetadata`
- `resolveEnvironmentExecutionTarget` reads `shellCommand` from lease
metadata
(validating against `"bash" | "sh" | null`)
- `environment-runtime.ts` adds `"shellCommand"` to
`INTERNAL_PLUGIN_SANDBOX_CONFIG_KEYS`
so the field round-trips through internal plugin config without leaking
to
external plugin metadata
- Updated tests in `command-managed-runtime.test.ts`,
`execution-target-sandbox.test.ts`, `sandbox-callback-bridge.test.ts`,
`environment-execution-target.test.ts`
## Verification
- `pnpm --filter @paperclipai/adapter-utils test`
- `pnpm --filter @paperclipai/server test --
environment-execution-target`
- `pnpm --filter @paperclipai/sandbox-providers-e2b test`
- Manual QA: boot a Paperclip instance, create an E2B-backed
environment, run a
claude_local agent against it, and confirm the run completes (verifies
bash
shell semantics flow through the callback bridge end-to-end)
## Risks
- E2B sandbox commands now run under `bash -lc` instead of `sh -lc`.
Bash is a
strict superset for the commands we issue (no busybox-only flags in our
shell
scripts), so risk is low. The shellCommand field is opt-in via lease
metadata —
providers that don't declare it stay on `sh`.
- New optional field on `CommandManagedRuntimeSpec` and
`AdapterSandboxExecutionTarget`.
Consumers ignoring the field retain previous behaviour (sh).
- Lease metadata now carries an additional field. Existing leases
without
`shellCommand` resolve to `null` and fall back to sh — backwards
compatible.
## Model Used
- OpenAI GPT-5.4 (reasoning effort: high) via Codex CLI
- Provider: OpenAI
- Used to author the code changes in this PR
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] If this change affects the UI, I have included before/after
screenshots — N/A (no UI changes)
- [ ] I have updated relevant documentation to reflect my changes — N/A
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
c0ce35d1fb |
Improve E2B plugin configuration UX and fix execution timeouts (#4802)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - E2B is a sandbox provider plugin that runs agent code in isolated cloud environments > - Operators configure E2B through the plugin settings page > - But the E2B API key configuration was unclear — the settings field description didn't explain that pasted keys are auto-saved as company secrets, and the fallback to the host `E2B_API_KEY` variable wasn't documented > - Additionally, long-running E2B sandbox commands were timing out because the plugin environment RPC driver used a fixed timeout, and environment commands competed for the single foreground command slot > - This PR clarifies the E2B configuration UX, fixes RPC timeouts for plugin environment execution, and runs E2B environment commands in background mode to avoid blocking the foreground slot > - The benefit is clearer E2B setup for operators and more reliable sandbox command execution ## What Changed - Updated E2B plugin manifest and settings UI to clarify API key configuration — field description now explains that pasted keys are saved as company secrets and documents the `E2B_API_KEY` host fallback - Added test coverage for the plugin settings page rendering - Fixed `plugin-environment-driver.ts` to pass the configured timeout through to RPC calls instead of using a hardcoded default - Updated `environment-runtime.ts` to propagate timeout from the environment lease to the plugin driver - Changed E2B sandbox command execution to use background handles so long-running agent commands don't block the foreground slot needed by the callback bridge ## Verification - `pnpm test` — all existing and new tests pass - `pnpm typecheck` — clean - Manual: navigate to plugin settings, verify E2B API key field shows the updated description text - Manual: run an E2B-backed agent task with a long-running command, verify it completes without RPC timeout ## Risks - Low risk. Configuration UX change is cosmetic. The timeout fix passes an existing value through instead of dropping it. Background command execution is a behavioral change but only affects E2B sandbox commands — the foreground slot is still available for bridge health checks. ## Model Used Codex GPT 5.4 high via Paperclip. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |