mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
dd6742f7d07a67c42ae8a5cd38a958f2e7e9636b
193
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dd6742f7d0 |
fix(sandbox): reset step store for long-lived bridge work
A run-time `sandbox.exec` span attached to a dead startup step. Two bridge startup steps start long-lived work inside their measured step body: a poll timer, socket handlers, and a callback-bridge worker loop. Node snapshots the active-step store on each async resource at creation time, so this work kept the ended step store. Each later run-time exec then read that store and opened its span under the ended step, and it copied a wrong `criticalPath: false` flag. Add an exported `runWithoutActiveStep` helper in startup-timing and wrap the long-lived poll timer, the socket handlers, and the callback-bridge worker loop of both bridge lanes. Each continuation now reads an empty store, so each run-time exec opens an unparented span with no stale `criticalPath` flag. Control flow is unchanged and the change adds no new dependency. Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a69a3d6b0 |
refactor(observability): stop writing detailed per-step timing to the run log (#10776)
## Thinking Path > - Paperclip tracks agent work for the team. > - The startup timing path already emits detailed spans. > - The run log repeats the same detailed timing. > - That makes the event larger than it needs to be. > - This pull request removes the redundant per-step timing fields from the run-log payload. > - The benefit is a smaller log with the same detail still available in traces. ## Linked Issues or Issue Description No public GitHub issue or open pull request matches this change. I described the enhancement below. **What existing behavior does this improve?** `run.startup.step` event output. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** `run.startup.step` writes `roundTrips`, `providerExecMs`, `providerGetMs`, `createRuntimeMs`, and `ensureSessionMs`. The same detail already exists in the spans. **Proposed behavior** `run.startup.step` keeps only `step`, `durationMs`, and `outcome`. The heartbeat lifecycle timestamps stay unchanged. **Reason and benefit** This change removes redundant data from the run log. It keeps the useful detail in trace spans. It also makes the payload smaller and easier to read. The revert path stays clear because the removed data has one producer chain. **Breaking changes** Yes. Consumers that read the removed fields must switch to span data or the remaining event fields. The heartbeat lifecycle timestamps do not change. **Additional context** I searched GitHub for related open PRs and issues. I found no match. ## What Changed - Removed the redundant per-step timing fields from the run-log payload. - Removed the now-dead producer chain that fed those fields. - Kept the span-level timing data and the heartbeat lifecycle timestamps. ## Verification - `adapter-utils` typecheck clean - `server` typecheck clean - `startup-timing` suite: 30 passed - `adapter-utils` `execute` suite: 99 passed - `server` `environment-execution-target` suite: 21 passed ## Risks - Consumers that still read the removed fields will need a code change. - This is a one-way-door data removal for the run log. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af6b32d82a |
feat(observability): add provider OTel spans - cache-hit flag, plugin tracer, Daytona pack/transfer spans (#10764)
## Thinking Path > - Paperclip uses pull requests to ship control-plane changes with review and audit. > - This change extends sandbox provider telemetry. > - The first PR added startup and exec spans. > - This PR adds provider spans for cache hits, plugin tracer flow, and Daytona file sync steps. > - The host must keep the trust boundary and reject unsafe span data. > - The result is fuller OTel coverage with bounded labels and safe parentage. ## Linked Issues or Issue Description **Problem or motivation** Sandbox provider telemetry lacks an explicit cache-hit signal, plugin span context, and file-sync pack and transfer spans. The current proxy for a cache hit is fragile. The plugin SDK also needs a safe tracer surface and a per-call parent context. **Proposed solution** Add an explicit `cacheHit` flag from the Daytona sandbox handle lookup. Expose `ctx.tracer` on the plugin context and pass a `traceparent` string with each call. Record provider spans on the host with allowlist clamping and capability checks. Wrap Daytona file sync pack and transfer steps in spans. **Alternatives considered** Keep the old `providerGetMs == 0` proxy. Reject that path because the cache-hit decision belongs at the handle lookup, not in a timing proxy. The host stays the trust boundary for worker span data. It clamps labels, rejects bad parent data, and gates the worker-to-host span RPC by capability. A security review completed before this PR opened. ## What Changed - Added an explicit `cacheHit` flag from the Daytona sandbox handle lookup. - Added `ctx.tracer` on the plugin context and a `traceparent` field in the per-call context. - Added host-side span recording with attribute clamping and capability checks. - Wrapped Daytona file sync pack and transfer steps in spans. - Added tests for the host trust boundary, the plugin tracer no-op path, and the Daytona span paths. ## Verification - `pnpm --filter @paperclipai/plugin-daytona exec vitest run` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run plugin-host-services environment-execution-target instrumentation plugin-worker-manager` - `tsc --noEmit` for server, plugin-sdk, and adapter-utils ## Risks - Span data now crosses the worker boundary, so the host allowlist and capability gate must stay strict. - The `traceparent` path must stay valid and host-owned. - Daytona file sync spans must not add new round trips. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes:` / `Closes:` / `Refs:` or described the issue in the PR body - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c09ea7112b |
feat(observability): add granular OpenTelemetry spans for sandbox startup and execution (#10758)
## Thinking Path > - Paperclip coordinates AI agent work and records execution data. > - The sandbox startup path and the host-to-sandbox execution path need clearer OpenTelemetry spans. > - The trace data now shows wait time, critical path data, and bounded labels. > - This pull request adds bounded spans and closed attribute helpers for sandbox startup and execution. > - The result is better trace data with low cardinality and no secret leakage. ## Linked Issues or Issue Description - No public GitHub issue exists for this branch. - Related PR: #10536 - This PR extends the sandbox OpenTelemetry work with exec spans, root timing, and bounded labels. ## What Changed - Added a closed span-attribute contract for sandbox startup data. - Added a bounded label helper for command and region values. - Threaded the active step context into the execution path. - Added a `sandbox.exec` span with true timestamps and bounded attributes. - Added skipped-step and root-span timing data with low-cardinality context. - Moved handshake sub-times and bridge batch data onto spans. ## Verification - `pnpm --filter @paperclipai/adapter-utils run typecheck` - `pnpm --filter @paperclipai/server run typecheck` - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target` - `git log --oneline origin/master..origin/feat/sandbox-startup-otel-spans` - `git diff origin/master...origin/feat/sandbox-startup-otel-spans --stat` ## Risks - Span attribute rules could still miss a future field if new code skips the shared helper. - Low-cardinality labels may hide some detail, but that is the intended tradeoff. - Tracing stays fail-open, so a missing tracer still hides data rather than stopping work. ## Model Used - OpenAI Codex, GPT-5, tool use enabled. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge The current head still has one failing required check: `e2e`. The PR stays unready for board handoff until that check passes. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee851fc364 |
feat(heartbeat): expose cache-adjusted run cost (#10349)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The heartbeat service is the control plane that runs agents through adapters and records each run's usage and cost in the finance/cost ledger > - Adapters report a provider cost (`costUsd`), but there is no way to represent the provider-billed cost *after* prompt-cache discounts, so cache-heavy runs are priced wrong and some paid runs end up exported with a zero/null cost > - A benchmark comparing Paperclip-orchestrated runs against direct harness invocation of the same tasks measured 1.5–3.1× higher apparent USD per pair, largely because cache-discounted billing was not represented in the exported cost data > - This pull request adds an optional `cacheAdjustedCostUsd` field to `AdapterExecutionResult` and a `resolveCacheAdjustedCostUsd` helper in the heartbeat service that prefers the explicit cache-adjusted figure and falls back to the reported `costUsd`, persisting it into the run's usage/ledger JSON > - The benefit is that paid runs are no longer exported as zero/null cost and cache-heavy runs can be priced correctly, so operators comparing orchestrated vs. direct costs see real numbers ## Linked Issues or Issue Description No single existing public issue covers this exactly; closely related cost-reporting issues: - Refs #8947 — hermes adapter never reports usage/cost to Paperclip, so budget limits never trigger - Refs #6716 — hermes_local cost/usage capture returns zero - Refs #3320 — expose per-run token counts in activity log and dashboard **Problem (feature-request form):** Adapters can only report a single `costUsd`. Providers with prompt caching bill less than the nominal token cost, and the heartbeat cost accounting has no field for the cache-adjusted billed amount. As a result, cache-heavy paid runs are either priced at the undiscounted figure or, when the adapter withholds the ambiguous number, exported as zero/null. **Proposed solution:** an explicit optional `cacheAdjustedCostUsd` on the adapter execution result, resolved server-side with a safe fallback to `costUsd`. **Alternatives considered:** recomputing cache discounts server-side from token counts (rejected: provider pricing tables drift and cache billing rules are provider-specific; the adapter is the source of truth). ## What Changed - `packages/adapter-utils/src/types.ts`: added optional `cacheAdjustedCostUsd?: number | null` to `AdapterExecutionResult`, with a doc comment on adapter expectations - `server/src/services/heartbeat.ts`: added exported `resolveCacheAdjustedCostUsd()` (explicit cache-adjusted value wins when a finite non-negative number; otherwise falls back to a finite non-negative `costUsd`; otherwise `null`), and consistently uses the resolved billed value for ledger cents, cost status, and run usage JSON - `server/src/__tests__/heartbeat-cost-accounting.test.ts`: added unit coverage for explicit precedence, fallback, invalid values, adjusted-only pricing, and discounted ledger billing ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-cost-accounting.test.ts` — 1 file, 7 tests passed - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed - `pnpm --filter @paperclipai/server typecheck` — passed - GitHub Actions on head `dc9c3830bf694b9afb3b27c5c8c36bff38e7fdcb` — all 26 checks clean/skipped; one unrelated E2E checkout-contention flake passed on its single failed-job rerun - Greptile review on the same head — 5/5 with zero unresolved threads ## Risks - Low risk: the field is optional and additive; when absent, behavior falls back to the existing `costUsd` path - Ledger/usage JSON gains a new optional `cacheAdjustedCostUsd` key — consumers that strictly validate keys would need to tolerate it (usage JSON is already open-shaped) - No migrations, no API-breaking changes > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Implementation: Anthropic Claude Fable 5, model ID `claude-fable-5`, standard context window, agentic coding mode with shell/file tool use - PR preparation and verification: OpenAI Codex on GPT-5 (the runtime did not expose a more specific serving snapshot or context-window value), reasoning mode with shell and GitHub tool use ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (doc comment on the new field; no user-facing docs affected) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c185e64b77 |
feat(ui): chat-style task view behind an experimental flag (#10606)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators spend most of their time on the issue detail page. They talk to the assigned agent there through comments. > - The current page reads as a ticket form. The thread sits below properties, the composer sits mid-page, and live agent activity renders as dense transcript logs. > - Talking to an agent is a conversation. A chat-first layout matches that mental model better than a ticket form. > - A layout change this large must not disrupt current users. It needs a safe opt-in path and full parity with the existing thread features. > - This pull request adds a chat-style task view behind a new "Chat-Style Tasks" experiment toggle. The flag is off by default and the existing page is unchanged when it is off. > - The benefit is a focused, readable conversation with the agent: live tool activity folds into compact summaries, the composer stays at the bottom, and properties, plan, and artifacts move into header tabs. ## Linked Issues or Issue Description Refs #49 (chat with agents is a much-wanted feature). Related PRs found in the dedup search: - #4489 — an earlier, closed attempt to promote the conversation to the primary surface on issue detail. This PR is a fresh, flag-gated take on the same goal. - #8837 — an open PR that proposes a two-column task layout. It restructures the same page but keeps the ticket paradigm; this PR is orthogonal because it is opt-in and chat-first. **Subsystem affected** UI (issue detail page). **Problem or motivation** The issue detail page presents agent conversations as a ticket: properties first, thread below, composer in the middle of the page, and raw transcript noise during live runs. Users who mainly converse with their agents must scroll past chrome to follow the conversation, and live activity is hard to read. **Proposed solution** An opt-in chat-style view of the issue detail page, gated by a new "Chat-Style Tasks" experiment toggle in Settings → Experimental. With the flag on, the thread fills the center pane, the composer docks to the bottom of the viewport, Properties / Plan / Artifacts become header tabs, live turns show a status pill with the current tool action and elapsed time, and settled turns collapse to a "Worked · N tools" summary that expands into per-tool rows. With the flag off, nothing changes. **Alternatives considered** Restyling the existing layout in place (rejected: too disruptive without an opt-out), and a separate chat page beside the issue page (rejected: splits the task's single source of truth). A per-request lab page (`/task-chat-lab`, dev-only) was kept for design iteration instead. **Roadmap alignment** ROADMAP.md "CEO Chat" wants lighter conversations that still resolve to real work objects. This PR keeps the core task-and-comments model — it only changes presentation, opt-in — so it does not duplicate that planned work. ## What Changed - New `enableTaskChatRedesign` instance setting, exposed as a "Chat-Style Tasks" experiment card in Settings → Experimental (shared feature catalog, validators, server instance-settings service, and UI settings page). - New `ui/src/components/task-chat/` component family: chat thread with turn grouping, agent reply bubbles, live status pill, collapsible turn summaries with per-tool rows, plan tab with a sticky CTA action bar, inline interaction cards, per-request mode chips, and a bottom-docked composer. - A shared tool taxonomy (`tool-taxonomy.ts`) maps tool names to verbs and icons; the status pill, tool rows, and the classic transcript view all use it. - A transcript adapter converts stored run logs into chat turns; it dedupes tool-call updates by `toolUseId` so tool counts match the expanded rows, and it keeps a tool row's first real name when later generic updates arrive. - Composer: posts on Cmd/Ctrl+Enter, supports image paste with object-URL thumbnail previews (revoked on clear/unmount), and uploads through the issue attachments route. - `IssueDetail.tsx`: with the flag on, pane tabs move to the header bar, the header is not sticky, and the chat fills the center; with the flag off, the previous layout renders unchanged. - Motion tokens for the new animations live in `ui/src/index.css` with a `motion-tokens.ts` catalog and a test that keeps the two in sync (the catalog now also covers the shared enter/exit/swap tokens that the decision/quicklook block declares). - A dev-only `/task-chat-lab` page with fixtures and a tweak panel for motion tuning. ## Verification - `pnpm typecheck` — clean across the workspace. - `pnpm check:token-gates` — 3/3 CLEAN. - `cd ui && pnpm vitest run` — 3,344 of 3,345 tests pass locally. The one failure is `IssueProperties.test.tsx` monitor-row time formatting, which is timezone-sensitive: it also fails on unmodified `origin/master` in a non-UTC timezone and passes with `TZ=UTC`. It is not related to this change. - `cd server && pnpm vitest run src/__tests__/instance-settings-service.test.ts` — 21/21 pass (covers the new setting). - Manual: start the dev server, open Settings → Experimental, enable "Chat-Style Tasks", and open any issue. The thread fills the page, the composer docks to the bottom, and Properties / Plan / Artifacts appear as header tabs. Assign an agent and comment to watch a live run: the status pill shows the current tool action with elapsed time, and the finished turn folds into a "Worked · N tools" summary. Disable the toggle and confirm the classic page is unchanged. - Visual snapshot baselines are intentionally not updated: per `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - The flag-off path goes through the same `IssueDetail.tsx` file, so a regression there would affect current users. Mitigation: the classic markup renders through the same components as before behind explicit flag conditionals, and the full UI suite passes. - The transcript adapter interprets stored run-log formats, including legacy entries without `toolUseId`. Malformed logs degrade to generic tool rows rather than crashing. - The new view changes no server behavior other than one additive instance setting; it is additive and default-off. Overall risk with the flag off is low. ## Model Used - Claude (Anthropic), model id `claude-fable-5` (Claude Fable 5), extended thinking enabled, agentic tool use (file editing, shell, test execution) via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6401f4f78c |
fix(sandbox): bundle git copy-back against the merge-base so diverged/reset workspaces still import (#10601)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - An agent that runs in a sandbox has its workspace copied back to the host when the run ends, so its work persists — the copy-back ships a git bundle of the sandbox's commits > - The bundle is created as a thin delta, `git bundle create HEAD --not <baseSha>`, which records `baseSha` (the host workspace HEAD captured at export) as a prerequisite the host must already hold > - That assumption breaks when the sandbox HEAD has diverged from `baseSha`, or when a shared host workspace no longer holds `baseSha` at import time — then `git fetch` on the host hard-fails and the entire run is lost even though the agent finished its work > - This pull request bundles against the merge-base of `baseSha` and the sandbox HEAD (with a full-bundle fallback), which the host can satisfy in those cases > - The benefit is that copy-back no longer discards a completed run's work over a base the host can't reconcile ## Linked Issues or Issue Description **What happened?** A sandbox agent run completed its work, then failed during workspace finalize: ``` git -C <host workspace> fetch --force <git-delta.bundle> refs/…/export:refs/…/imported error: Could not read <baseSha> fatal: revision walk setup failed error: git-delta.bundle did not send all necessary objects ``` The run is reported as `adapter_failed` even though the agent produced output. The copy-back bundle names the host workspace's recorded HEAD (`baseSha`) as a prerequisite, but the host cannot satisfy it. **Steps to reproduce** Two independent triggers, both reproduced in tests: 1. The sandbox's HEAD has diverged from `baseSha` — e.g. the sandbox carries a local-only branch that forked from an older commit than the host's current HEAD. 2. The shared host workspace no longer holds `baseSha` at import time (it was reset / re-realized between export and import). In either case `git fetch` of the thin bundle fails with a missing prerequisite. **Expected behavior** Copy-back imports the sandbox's work as long as the host holds any common ancestor, instead of hard-failing and discarding the run. **Paperclip version** Current `master`. **Deployment mode** Any deployment running agents in sandbox environments with workspace sync (notably shared-workspace clones and custom images that carry a local-ahead branch). ## What Changed - `buildRemoteGitDeltaBundleScript` now computes `bundle_base = git merge-base <baseSha> HEAD` and bundles `HEAD --not <bundle_base>`. The merge-base is an ancestor of `baseSha`, so any host that holds `baseSha` (or an ancestor of it — e.g. after a reset) can satisfy the prerequisite, and the bundle stays a delta rather than a full-history transfer. - When `baseSha` is absent from the sandbox, or no merge-base exists, it falls back to a full, self-contained bundle (no prerequisites) so the import can always complete. - The existing empty-bundle no-op (no new commits) and the ordinary fast-forward path are unchanged; the `cat-file` base check no longer aborts the script under `set -e`. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — new cases: a diverged sandbox HEAD imports when the host holds only the merge-base (not `baseSha`), and the full-bundle fallback imports into a host that shares no history; existing thin-delta and empty-bundle cases still pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Standalone shell repro confirmed the old thin bundle fails with "Repository lacks these prerequisite commits" in both trigger cases, and the merge-base bundle imports successfully. ## Risks - Low. For the common case (sandbox HEAD descends from `baseSha`) the merge-base is `baseSha`, so the bundle is byte-for-byte the same delta as before. The change only alters behavior when the old code would have hard-failed. - This makes the copy-back import succeed on a diverged base; the subsequent reconciliation of divergent histories (`integrateImportedGitHead`) is unchanged and still owns how the imported head is merged into the host branch. Where a workspace's history has genuinely diverged (e.g. a stale custom image carrying a local-only branch), a clean re-clone/re-capture is still the right operational fix — this change prevents work loss, it does not reconcile intentional divergence. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — extended thinking, agentic tool use (file edits, shell repro, vitest/tsc runs). No other models involved. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2e4774ac90 |
fix(adapter-utils): correct confirmation wake semantics (#10588)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent wakes include a default execution contract > - That contract tells agents how issue-thread continuation policies behave > - The current text says `wake_assignee` resumes a confirmation only after acceptance > - The server actually wakes for every non-expired resolution and reserves acceptance-only behavior for `wake_assignee_on_accept` > - This pull request makes the default prompt match the server contract and strengthens the recovery follow-up regression case > - The benefit is that agents choose the correct continuation policy and recovery tests cover normalized agent name keys ## Linked Issues or Issue Description Related work: Refs #5473, Refs #5060, and Refs #10562. **What happened?** The default local-agent prompt described `wake_assignee` as acceptance-only for `request_confirmation`. This conflicts with the server. The server wakes on every non-expired resolution. A recovery follow-up test also used an already-normalized execution agent name key, so it did not exercise the normalization seam. **Expected behavior** The prompt must state that `wake_assignee` resumes after acceptance or rejection. It must direct acceptance-only flows to `wake_assignee_on_accept`. The recovery regression must use a display-style agent name key and prove that the follow-up path still works after normalization. **Steps to reproduce** 1. Read the default local-agent prompt in `packages/adapter-utils/src/server-utils.ts`. 2. Compare its confirmation continuation text with `queueResolvedInteractionContinuationWakeup` in `server/src/routes/issues.ts`. 3. Observe that the prompt gives acceptance-only semantics to `wake_assignee`. 4. Inspect the recovery hand-back test and observe that its execution name key is already normalized. **Paperclip version or commit** `7301fae942` **Deployment mode** Local dev. The prompt and test behavior are not deployment-specific. ## What Changed - Corrected the default agent prompt for `wake_assignee` and `wake_assignee_on_accept`. - Added focused prompt assertions for both the required and obsolete text. - Changed the recovery follow-up fixture to use a display-style agent name key. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts -t 'keeps the default local-agent prompt action-oriented'` passed: 1 test. - `pnpm exec vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts -t 'defers recovery hand-back wakes until the resolving run exits'` passed: 1 test. - `pnpm --filter @paperclipai/adapter-utils typecheck` passed. - `pnpm --filter @paperclipai/server typecheck` passed. - `git diff --check origin/master...HEAD` passed. ## Risks - Low risk. The production change updates prompt text only. - Agents that followed the old text may now choose `wake_assignee_on_accept` for acceptance-only flows. - The server test change only broadens an existing regression fixture. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact serving model ID and context-window size are not exposed to the agent. The model used reasoning, repository tools, tests, Git, and GitHub CLI access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ee1275f11 |
fix(adapters): persist ACPX process identity for hot restart (#9838)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work. > - Local agent heartbeats need durable process identity so the server can supervise them. > - The ACPX runtime owns the child process used by `codex_local` sessions. > - ACPX did not expose the child PID and start time to the Paperclip adapter. > - Warm ACPX runtimes can also serve a later heartbeat without a new spawn event. > - A hot restart could therefore classify a live Codex run as lost because its heartbeat row had no process identity. > - This pull request forwards ACPX spawn identity, reuses it for compatible warm heartbeats, and fails closed when identity cannot be persisted. > - The benefit is reliable hot-restart adoption for eligible local Codex runs. ## Linked Issues or Issue Description No matching public GitHub issue was found. **What happened?** A `codex_local` heartbeat could run through ACPX without a persisted `processPid` or `processStartedAt`. A Paperclip hot restart then had no durable identity for the live ACP child. Recovery could classify the run as `process_lost` even while the child was still alive. **Expected behavior** ACPX reports the real child PID and start time before the first prompt. A compatible warm runtime reports the same known identity to each later heartbeat that reuses the child. ACPX stops the child if the identity is invalid or persistence fails. Hot-restart recovery can then adopt the live run. **Steps to reproduce** 1. Start a `codex_local` heartbeat through the ACPX execution lane. 2. Keep the run active during a Paperclip hot restart. 3. Inspect the heartbeat row before this change. 4. Observe that the process identity can be null and recovery cannot adopt the live child. **Reproduced on** - Paperclip `master` before this change. - Linux source deployment. - `codex_local` with ACPX `0.12.0`. ## What Changed - Add an awaited `onAgentSpawn` lifecycle hook to the patched ACPX runtime. - Forward the ACP child PID and start time through the adapter `onSpawn` callback. - Keep a mutable callback sink for cached runtimes so a later respawn updates the current heartbeat. - Reuse the last known process identity when a compatible warm heartbeat reuses the existing child. - Kill the ACP child and fail session startup when the PID is invalid or identity persistence rejects. - Add ACPX and heartbeat recovery tests for callback ordering, warm reuse, failure cleanup, durable row identity, and hot-restart adoption. - Document the one-time drain required when an installed pre-fix run already lacks process metadata. ## Verification - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-execute-escalated" pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 89 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-recovery-escalated" pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts` — 92 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-remote-smoke-escalated" pnpm exec vitest run packages/adapter-utils/src/acpx-engine/remote-spawn-smoke.test.ts` — 3 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-ci-repro-escalated" pnpm exec vitest run server/src/__tests__/heartbeat-dependency-scheduling.test.ts` — 6 passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - Reverse and forward dry-run application of `patches/acpx@0.12.0.patch` — passed. - `git diff --check` — passed. - `git diff --exit-code origin/master...HEAD -- pnpm-lock.yaml` — passed. - `git diff --exit-code origin/master...HEAD -- .github/workflows` — passed. ## Risks - Runtime risk is low to moderate. ACPX now awaits process-identity persistence during child startup. - ACPX kills the child when persistence fails. This prevents an unsupervised process, but it makes that heartbeat fail visibly. - A compatible warm heartbeat reuses the identity of the existing ACP child. Regression tests verify that identity is persisted before the next prompt. - The change updates the vendored ACPX patch. Package installation must apply that patch. - There are no schema, migration, public API, UI, workflow, or lockfile changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex used GPT-5.3-Codex for the earlier implementation. - OpenAI Codex used GPT-5 for the lifecycle-hook revision and the current fail-closed review fix. The runtime did not expose a more specific snapshot ID or context-window size. Both runs used reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53bcf3897f |
feat: sync @-mentioned projects into remote sandboxes (confined sandbox transport, flag ON) (#10564)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The run layer must move project context into the sandbox that executes the agent > - Local sandboxes already stage referenced projects for @-mentions > - Remote confined sandboxes dropped the whole referenced set, so the agent lost needed files and paths > - This pull request keeps the confined sandbox transport aligned with the local behavior for referenced projects > - It does this behind a remote-only flag that defaults on, while SSH keeps the old drop-only path > - The benefit is that remote runs can read the same referenced project context that local runs already provide ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change touches server orchestration, sandbox transport, and observability. **Problem or motivation** A run can @-mention another project. Local targets stage each referenced project and give the agent a path. Remote confined sandboxes dropped the full referenced set, so the agent could not read those project files or paths. **Proposed solution** Enable referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`, which defaults on. Keep the SSH transport out of scope and keep it dropping referenced projects. Repoint each referenced workspace hint at its staged `project-<projectId>` sandbox directory. Publish `PAPERCLIP_WORKSPACES_JSON` on the confined sandbox lane. Count each per-project remote staging failure as a `staging` failure in the requested-vs-synced metrics. **Alternatives considered** Keep the remote path drop-only. That keeps the gap open. Move the change into SSH too. That expands scope beyond the target transport and adds risk. **Roadmap alignment** No matching item in `ROADMAP.md` showed up in this review. **Additional context** The change lands in three commits. The first commit opens the gate for the confined sandbox transport. The second commit repoints the workspace hints and publishes the workspace map. The third commit records per-project staging failure data. ## What Changed - Opened remote referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`. - Repointed referenced workspace hints to the staged `project-<projectId>` sandbox directories and published `PAPERCLIP_WORKSPACES_JSON`. - Counted per-project remote staging failures as first-class `staging` failures in the requested-vs-synced observability. ## Verification - The pushed ref `refs/heads/feat/sync-referenced-projects-remote-sandbox` resolves to the authorized submit SHA. - `git log --oneline origin/master..origin/feat/sync-referenced-projects-remote-sandbox` shows exactly the three expected commits. - The handoff reports server typecheck clean, adapter-utils typecheck clean, and the listed unit tests passing. - The handoff also reports no open review comments and no Greptile score yet. ## Risks - The change touches authorization and sandbox path handling, so regressions could block remote runs or expose the wrong project context. - The new flag defaults on, so any bug in the remote path affects normal remote use. - SSH stays out of scope, so the two transport paths must remain distinct. ## Model Used OpenAI GPT-5 via Codex. Tool use enabled. Context window not reported in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b4a7a12985 |
feat: make recovery updates quieter (#10542)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The recovery subsystem restores work after an agent run stops or loses state > - Recovery notices currently use the same visual weight as normal work comments > - Recovery agents can also post long narratives that obscure the useful hand-off > - The server must identify recovery output because agents cannot set presentation controls > - This pull request adds compact recovery notices, structured action references, and brief recovery prompts > - The benefit is a quieter issue thread that still keeps recovery state inspectable ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: `server/`, `packages/shared`, and `packages/adapter-utils`. **Problem or motivation** Recovery notices and recovery-run comments can dominate an issue thread. Operators must scan routine recovery narration before they find the work hand-off. **Proposed solution** Give routine recovery output a compact system-notice presentation. Derive the presentation on the server so agents cannot hide arbitrary comments. Keep the successful missing-state summary fully visible because that comment is the recovery deliverable. **Alternatives considered** The UI could detect recovery text. That approach is fragile and does not provide structured action references. Agents could also set presentation directly, but that would weaken the current board-only security boundary. **Roadmap alignment** This change refines the completed “Self-healing runs & automatic recovery” and “Enforced Outcomes” roadmap areas. It does not add a competing roadmap capability. **Additional context** The scope covers shared comment validation, server recovery notices, agent-comment derivation, and recovery prompt text. No database migration is needed because presentation data already uses JSON. ## What Changed - Add the `compact` issue-comment presentation density to shared constants, types, and validation. - Give recovery escalation, waiting, and in-place notices compact titles and structured recovery-action metadata. - Use recovery-action metadata for notice deduplication, with the legacy text marker as a compatibility fallback. - Derive compact presentation for comments from recovery-scoped runs while preserving the board-only presentation boundary. - Keep successful missing-state recovery summaries fully visible. - Ask recovery participants to record outcomes in `resolutionNote` and keep source-issue comments brief. - Add shared, route, service, and prompt tests for the new behavior and exceptions. ## Verification - `pnpm -r typecheck` - Focused Vitest coverage: 320 tests passed across shared validators, adapter prompts, issue comments, recovery actions, and heartbeat recovery. - Full server phase: 292 files passed, 3,094 tests passed, and 2 tests skipped. - Full UI phase: 386 files passed and 3,182 tests passed. - `pnpm build` - Known master baseline: `cli/src/__tests__/secrets.test.ts` expects `pass`, but the current implementation returns `warn` when strict secret mode is disabled for Postgres. This branch does not change CLI secrets code. ## Risks - Low migration risk. The presentation column is JSON and needs no database migration. - Recovery-run detection depends on the persisted run context snapshot. - Structured metadata becomes the primary deduplication key. The existing body marker remains as a fallback for older comments. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The runtime did not expose the context-window size. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9f7565f4ce |
feat: activate OpenTelemetry spans on the sandbox start path (#10536)
## Thinking Path > - Paperclip moves agent work through sandboxed execution and control-plane services. > - The sandbox start path now has a no-op span seam. > - This change turns that seam on when OTLP export is configured. > - It keeps the default path unchanged when export is off. > - The result is structured startup traces with low-cardinality attributes and explicit parent links. > - The benefit is better observability without changing normal behavior. ## Linked Issues or Issue Description No public GitHub issue exists for this change. ### Problem The sandbox start path has a tracer seam, but it stays a no-op unless the OTLP export path is active. ### Proposed solution Enable the server tracer on sandbox bring-up, open a root span, parent each startup boundary to that root, and keep the export path opt-in behind `OTEL_EXPORTER_OTLP_ENDPOINT`. ### Alternatives considered - Keep the start path as a no-op. I rejected that path because it leaves sandbox start opaque when OTLP export is already configured. - Add broad attributes for commands and paths. I rejected that path because the span allowlist must stay low-cardinality. ### Roadmap alignment This follows the current OTel sandbox-start work and keeps the default path unchanged. ## What Changed - Add a root sandbox startup span and child spans for each named startup boundary. - Keep concurrent bridge spans parented to the root span. - Inject the server tracer through the adapter deps without OpenTelemetry imports in the engine. - Attach host-received provider duration attributes only when the values are finite. - Keep span attributes inside the allowlist and keep command, path, id, and error text out of span data. ## Verification - The pushed branch already passed `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - The pushed branch already passed `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/`. - The pushed branch already passed `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-execution-target.test.ts src/__tests__/instrumentation.test.ts`. - The pushed branch already passed `pnpm --filter @paperclipai/server exec tsc --noEmit`. ## Risks - OTel export changes trace volume when the endpoint is set. - The allowlist limits trace detail, so new fields need care. - The change stays no-op when OTLP export is off. ## Model Used - OpenAI GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5cffd5c72e |
feat(sandbox): declare and collect the advisory read-write intent (#10521)
## Thinking Path > - Paperclip is the control plane for AI agent companies > - Sandbox providers move files between the host and the sandbox > - The runtime needs advisory data for which sync paths are writable > - The host already knows that intent when it prepares the sync map > - Daytona can record that intent for later use > - This pull request adds the advisory access field and collects writable paths > - The benefit is that later runtime work can use the data without changing current flow ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description below follows the feature-request template fields. ### Problem or motivation An agent runs inside an ephemeral sandbox. Some sandbox paths keep agent changes; other paths do not. Today the runtime has no declared signal for which sync destinations the agent may change and keep. A later feedback wrapper needs this signal to give the agent real-time feedback when a write lands on a non-persistent path. ### Proposed solution Add an optional advisory `access: "rw" | "ro"` field to the sync file-mapping types. The host sets `rw` for the workspace, git-history, and asset destinations, and `ro` for referenced-project trees. An absent value defaults to `ro`. The Daytona provider collects the `rw` target directories into a per-lease writable set for later use. This change adds the metadata and the collection only. No execution path reads the writable set yet, so runtime behavior does not change. ### Alternatives considered Derive a static writable set in provider code. This alternative is weaker: the layer that authors each sync destination already knows the intent, so a per-operation declaration is more accurate and does not hard-code a path list. ### Roadmap alignment This is the first step toward an advisory sandbox feedback wrapper. The wrapper is best-effort and adds no security. The ephemeral sandbox stays the only boundary. ### Additional context The field is advisory metadata. It does not change the file transfer and adds no security. ## What Changed - Added an optional `access` field to the sync file mapping types. - Set the workspace, git history, and asset destinations to writable. - Set referenced project destinations to read only. - Added a writable set store in the Daytona provider. - Recorded the parent directory of each writable sync mapping during sync in. ## Verification - The handoff reports these checks before push. - `pnpm --filter @paperclipai/adapter-utils exec vitest run sandbox-managed-runtime.test.ts` - `pnpm --filter @paperclipai/plugin-sdk exec tsc --noEmit` - `vitest run src/plugin.test.ts` in the Daytona package - `tsc --noEmit` in the Daytona package ## Risks - Low risk. - The new field is advisory. - The writable set store is best effort and in memory. - A cold store falls back to the workspace baseline. ## Model Used OpenAI Codex, GPT-5, tool use, code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in the PR - [x] I have not referenced internal or instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [ ] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
740554acc6 |
feat: add a no-op OpenTelemetry span seam for the sandbox startup path (#10522)
## Thinking Path > - Paperclip helps people run and govern AI agent work > - Sandbox startup needs a safe place to add telemetry spans without forcing OpenTelemetry on every run > - This change adds a no-op span seam, so the startup path can accept a tracer later and still stay inert now > - The server gets a lazy tracer accessor, and the adapter timing helper gets an injected tracer hook > - The change keeps the default path free of OpenTelemetry and keeps the existing startup event path unchanged > - The benefit is a future-safe seam with no runtime change today ## Linked Issues or Issue Description This PR addresses a feature gap in the sandbox startup path. ### Problem Sandbox startup has no safe span seam. A direct OpenTelemetry import would load telemetry packages on every run. ### Proposed Solution Add a lazy tracer accessor in the server. Add an injected no-op tracer seam in startup timing. ### Alternatives Import OpenTelemetry directly in the startup path. Reject that path because the default startup flow must stay inert. ## What Changed - Added a lazy startup tracer accessor in `server/src/instrumentation.ts`. - Added an injected startup tracer seam in `packages/adapter-utils/src/acpx-engine/startup-timing.ts`. - Kept the startup event path unchanged. - Kept `adapter-utils` free of OpenTelemetry imports. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` - `pnpm exec vitest run server/src/__tests__/instrumentation.test.ts` - `tsc --noEmit` for `@paperclipai/adapter-utils` and `@paperclipai/server` ## Risks Low risk. The default tracer is a no-op, so the runtime path stays inert until a later change injects a real tracer. ## Model Used OpenAI GPT-5. Tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and found none - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d5b9f6c8c9 |
perf(sandbox): coalesce git-workspace stage-sync into one confined syncIn (#10488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cloud and sandbox agents need less round-trip overhead at workspace start > - The workspace start path now uploads the git history and the working tree overlay in two separate sync operations > - That adds extra mkdir, guard, upload, and rename work for the same workspace bytes > - This pull request merges those two uploads into one guarded sync operation > - The benefit is one sync round trip for the workspace bytes with the same guard and login-shell contract ## Linked Issues or Issue Description No public GitHub issue exists for this change. ### Problem A git-backed workspace start uploads the git-history clone and the working-tree overlay as two separate host-to-sandbox sync operations. ### Proposed Solution Merge the two uploads into one sync operation. Keep both tar mappings under the same confinement guard. Run the git-history extract first, then the overlay extract, then optional cleanup. ### Alternatives Considered Keep two sync operations. That keeps the current round-trip cost and duplicates the guard, upload, and rename work. ### Roadmap Alignment This change supports the roadmap work on cloud and sandbox agents by reducing workspace start overhead. ## What Changed - Merge the git-history tar and the working-tree overlay tar into one sync operation. - Run the git-history extract first, then the overlay extract, then optional remove-deleted-paths cleanup. - Keep both temporary tar targets under the runtime directory so the first extract cannot delete the second tar before use. - Keep the symlink-escape guard, login-shell contract, and file-mapping checks on the merged file set. ## Verification - `tsc --noEmit` on `@paperclipai/adapter-utils` - `@paperclipai/adapter-utils` full project test suite: 354 pass / 4 skipped - Orchestrator proof: one merged operation with both tars and two ordered extract commands - Orchestrator proof: the confinement guard covers both tar mappings - Daytona plugin `plugin.test.ts`: 94 pass ## Risks - Any regression in the extract order could change workspace start behavior. - Any regression in the merged guard could block valid uploads or miss an escape attempt. ## Model Used OpenAI GPT-5, tool use enabled, current Codex session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
674d71548a |
perf(sandbox): fold dead sandbox start round trips (bridge dirs + handle-cache seed) (#10485)
## Thinking Path > - Paperclip is the control plane for AI agents. > - Sandbox startup uses bridge directories and a Daytona workspace handle. > - The cold path made repeat directory creation calls and one avoidable handle fetch. > - Those calls add delay but do not change state. > - This pull request folds the bridge directory setup into one exec, removes redundant process-session setup, and seeds the Daytona handle cache at acquire. > - The benefit is fewer deterministic host-to-sandbox round trips and faster cold starts. ## Linked Issues or Issue Description - Problem: Cold sandbox start does extra directory creation work and re-fetches a handle it already has. - Expected result: The startup path should create each directory once and reuse the fresh handle. - Related PRs I found on GitHub: #9280, #9293. ## What Changed - Added `makeDirs` to the bridge queue client and used one `mkdir -p` exec for the callback bridge directories. - Removed the two upfront `mkdir` execs for the process-session bridge stdin and events directories. - Seeded the Daytona sandbox handle cache at acquire so realize can reuse the fresh handle. - Reset the process-scoped cache in the compatibility test so the second sync run sees the expected exec count. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - 351 passed, 4 skipped. - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run` - 91 passed. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` - clean. - I checked `ROADMAP.md` for sandbox round-trip work. I found no duplicate planned core work for this change. ## Risks - Low risk. The change removes redundant calls and adds cache seeding. - A wrong cache scope would hide the handle. The seed now checks the lease scope and fails loudly. - The daytona package `tsc --noEmit` still depends on SDK types that are not installed in this isolated workspace. CI covers that path. ## Model Used - OpenAI GPT-5, tool-using, with code execution in the current workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a267e0328 |
feat(sandbox): stage referenced projects into the run sandbox (#10469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs can reference more than one project > - Each referenced project must land in its own sandbox tree so one run does not mix files across projects > - The anchor workspace must keep its own git history and overlay rules > - Fail closed on sync and confinement errors, and keep the other projects alive > - This pull request stages referenced projects into isolated project directories under the run sandbox root > - The benefit is safer multi-project runs with clear failure isolation ## Linked Issues or Issue Description No public issue exists. Related PR: #10448. ## What Changed - Thread additional referenced-project sources through the run prepare path and the runtime layers. - Stage each referenced project into its own `project-<projectId>` directory under the runtime root. - Keep the anchor workspace history and overlay semantics unchanged. - Fail closed on confinement or sync errors for one project, and keep the other projects running. - Keep the path inert by default behind the multi-project workspace-sync kill-switch. - No documentation update was needed for this code-only runtime change. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/sandbox-file-sync.test.ts src/command-managed-runtime.test.ts src/sandbox-managed-runtime.test.ts src/remote-managed-runtime.test.ts` - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` - The branch points at `871378e4064a95adc8e4647e442ec1359768de61` on `origin/feat/stage-referenced-projects-into-sandbox`. - GitHub CI is green on PR #10469. ## Risks - A referenced project can skip if confinement or sync fails. - A new runtime tree layout can affect tools that assume one project root. - The kill-switch keeps the path inert until operators enable it. ## Model Used OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes or confirmed no docs update was needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7083c275c8 |
refactor(sandbox): retire the dead noProfile flag from the exec path (#10461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The sandbox exec path starts agent commands and passes runtime options to the server and plugin layers > - This path kept a noProfile flag after the exec wrappers stopped sourcing a login profile > - The flag no longer changed behavior, so it left dead API surface in the protocol and runtime helpers > - This pull request removes that dead flag from the plugin protocol, the server drivers, and the managed-runtime helpers > - It also updates the tests and points the agent runtime README at the sandbox requirements file > - The benefit is a smaller and clearer exec-path contract with no behavior change ## Linked Issues or Issue Description - No public GitHub issue exists. ### What happened? The sandbox exec path kept a `noProfile` field after the exec wrappers stopped sourcing a login profile. ### Expected behavior The plugin protocol, server drivers, and managed-runtime helpers should not expose or forward a dead field. ### Steps to reproduce 1. Run a managed-runtime command through the sandbox exec path. 2. Inspect the protocol payload and runtime helper inputs. 3. Observe that `noProfile` is present even though it no longer changes behavior. ### Paperclip version or commit `60c7da86fc7a6c1dbf37bbcd86e25ecaaff01607` ### Deployment mode Built from source (pnpm dev / pnpm build) ### Additional context This pull request removes the dead field, updates the affected tests, and updates the README note for the sandbox profile path. ## What Changed - Removed noProfile from the plugin protocol and the server exec-path call sites. - Updated the managed-runtime helpers to use the narrower exec-path contract. - Updated the affected tests and added the README pointer to SANDBOX-REQUIREMENTS.md. ## Verification - `git grep -n "noProfile" -- packages/ server/` returns zero matches. - `tsc --noEmit` passed for `@paperclipai/adapter-utils`, `@paperclipai/plugin-sdk`, and `@paperclipai/server`. - `command-managed-runtime.test.ts` passed: 22/22. - `environment-runtime.test.ts` passed: 24/24. ## Risks - Low risk. The flag was already a no-op. - A hidden external caller may still send the removed field. ## Model Used - OpenAI GPT-5, tool-enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d6e235cbcf |
perf(sandbox-providers): drop nvm sourcing from exec wrappers (#10443)
## Thinking Path > - Paperclip keeps agent work on a controlled execution plane. > - Sandbox exec wrappers run on the hot path for agent commands. > - The current change removes the explicit `nvm.sh` load step from those wrappers. > - The sandbox image already restores PATH through profile startup. > - This pull request keeps profile sourcing where the wrapper still needs it and drops only the `nvm.sh` load step. > - The result is a smaller command path with the same node and agent CLI resolution. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Problem: The sandbox exec wrappers spent extra time sourcing `nvm.sh` before each command. The sandbox image already restores PATH in `/etc/profile.d/00-restore-env.sh`, so that explicit `nvm.sh` work was redundant. Proposed solution: Remove the `nvm.sh` source step from all six wrappers. Keep the profile sourcing that the provider still needs for PATH setup. Alternatives considered: Keep the existing shell setup and accept the launch cost. That keeps the current behavior, but it leaves the hot path slower than needed. Roadmap alignment: This change keeps the sandbox command path small and predictable. It does not change the adapter contract or the node resolution rules. ## What Changed - Removed `nvm.sh` sourcing from all six sandbox exec wrappers. - Kept profile sourcing where the provider still needs it for PATH setup. - Switched Modal to a non-login shell because the script now sources profiles itself. - Updated wrapper tests to assert that built commands do not source `nvm.sh`. ## Verification - Local TypeScript typecheck passed in each changed package. - Focused provider tests passed for Daytona, E2B, Modal, exe-dev, Cloudflare bridge, and adapter-utils. - One Daytona test failure is pre-existing and unrelated to this change. ## Risks - This change alters shell startup for sandbox exec paths. - A provider that depends on implicit shell setup may need a follow-up. - The current tests cover command shape, but they do not cover every runtime shell path. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ff7e1caf3 |
perf(adapter-utils): content-hash-skip process-session remote script write + lock inbound round-trip count (#10377)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The sandbox runtime depends on efficiently staging files into remote environments > - The process-session bridge was rewriting a static remote script on every start, even when the remote copy already matched > - That caused unnecessary round trips and made the startup path noisier than it needed to be > - This pull request adds a hash-skip path for the process-session script write and keeps the existing bridge-entrypoint hash gate on the same helper > - It also locks in the reduced inbound round-trip count so the collapse from the prior PR cannot silently regress > - The benefit is fewer remote execs on warm starts and a stronger guard against performance regressions ## Linked Issues or Issue Description - Refs: https://github.com/paperclipai/paperclip/pull/10354 ### Feature request-style description **Subsystem affected** - Cross-cutting (packages/adapter-utils and sandbox staging helpers) **Problem or motivation** - The process-session bootstrap path was rewriting a static remote script on every start even when the remote file already matched the host content. - That added avoidable remote execs and latency to warm starts. - The inbound staging collapse from the previous PR also needed a regression lock so it could not silently drift back to extra round trips. **Proposed solution** - Route the process-session remote script upload through the existing hash-skip helper used by the sandbox callback bridge entrypoint. - Keep the bridge-entrypoint behavior unchanged by delegating it to the same helper. - Add a regression test that asserts the inbound staging path still collapses to the expected round-trip count. **Alternatives considered** - Keep the existing unconditional write path and accept the extra execs on warm starts. Rejected because it preserves avoidable overhead. - Add a second specialized helper just for process-session scripts. Rejected because the bridge-entrypoint logic already solved the same problem and should stay aligned. **Roadmap alignment** - This fits the roadmap items around cloud/sandbox agents and enforced outcomes by reducing bootstrap waste and preventing performance regressions. - It does not introduce new product surface area, telemetry, or user-facing workflow changes. **Additional context** - The implementation preserves the existing upload/verify/rename behavior when the remote content differs and only skips the write when the hash already matches. - PR #10354 collapsed the inbound staging path; this PR preserves that win. ## What Changed - Added a shared hash-skip helper for remote text-file synchronization so unchanged content skips the write path after a single remote hash check. - Switched the process-session remote script upload to use that helper, preserving the existing write/verify/rename behavior when the remote content differs. - Refactored the sandbox callback bridge entrypoint sync to use the same helper without changing its observable behavior. - Added a mocked-native-runner regression test that locks the inbound staging path to one `client.syncIn` round-trip per step and zero direct write/run execs. ## Verification - `tsc --noEmit` - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - Full adapter-utils vitest sweep: 338 passed / 4 skipped ## Risks - Low risk: the helper changes when a write occurs, not the script contents or the remote execution surface. - A hash-check failure now fails loudly instead of silently rewriting, which is safer but could surface provider-side issues earlier than before. - The regression test is intentionally specific to the current collapsed staging path, so future architectural changes will need test updates. ## Model Used OpenAI Codex (GPT-5, tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
341993ebae | feat(sandbox-runtime): route all inbound staging through client.syncIn (Codex home -> native uploadFiles; delete usesCustomProvision gate) (#10354) | ||
|
|
7797995038 |
perf(plugin-daytona): opt-in no-profile fast path for default-PATH execs (#10352)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - The Daytona adapter turns tasks into shell commands and manages execution overhead > - Many short-lived exec calls still pay for login-shell profile sourcing even when the binary already resolves on the sandbox default PATH > - That extra startup work adds latency on the hot path for repeated command execution > - This pull request adds an opt-in fast path that skips profile sourcing only when the caller explicitly requests it and the command does not need shell initialization > - The benefit is lower per-call latency for eligible commands without changing the conservative default behavior for commands that need the profile ## Linked Issues or Issue Description This change does not reference a public GitHub issue. It follows the same Daytona startup-speed work as merged PR #10335 and narrows the execution path for eligible commands while keeping the default login-shell behavior intact. ## What Changed - Added an optional `noProfile` flag to `PluginEnvironmentExecuteParams`. - Refactored Daytona login-shell script assembly so the profile and nvm sourcing block is omitted only on the explicit fast path. - Preserved environment prefixing, `cd`, shell quoting, `NONINTERACTIVE_GIT_ENV`, stdin handling, and `durationMs` behavior on both paths. - Added regression tests for the fast path omission, the preserved execution parameters, and the default profile-sourcing path. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run src/plugin.test.ts` - `pnpm --filter @paperclipai/plugin-sdk tsc --noEmit` - Reverted the guard locally to confirm the two behavior tests fail again, then restored the change. ## Risks - If a caller opts into `noProfile` for a command that depends on shell initialization, the command can fail to resolve its binary. - The API comment and opt-in design keep that risk narrow; the default path remains unchanged. ## Model Used OpenAI GPT-5 (Codex tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1cd09ed555 |
perf(heartbeat): reuse task sessions for issue-scoped timer wakes and bound control-plane write retries (#10350)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents make progress in heartbeats: the server wakes an agent session, it does a slice of work on an issue, records a disposition, and exits > - Benchmarking identical coding tasks run as Paperclip-orchestrated agent pairs vs invoking the same agent harness directly measured a 1.8–2.2× wall-clock slowdown for the Paperclip pairs, dominated by per-heartbeat orchestration overhead rather than model time > - Two contributors stood out: (1) since PF-4 (#4838) every `heartbeat_timer` wake starts a brand-new task session, so continuation work on a specific issue repays the full session-start and re-orientation cost on every heartbeat; (2) in degraded environments agents burn many tool calls retrying the same failing control-plane write before giving up > - This pull request reuses the task session for issue-scoped timer wakes (keeping the PF-4 fresh-session rule only for unscoped exploratory wakes, which were the original context-bloat case) and adds a bounded-retry rule to the wake prompt and core skill: after 2 consecutive failures of the same control-plane write, stop retrying it for the rest of the heartbeat and rely on the adapter/runtime status channel > - The benefit is materially less wall-clock and token overhead per heartbeat while preserving the context-bloat protection PF-4 was added for ## Linked Issues or Issue Description Refs #4838 (merged PF-4 change whose reset rule this refines), Refs #5287, Refs #1907 (related timer-heartbeat session work). No public GitHub issue exists for the slowdown itself; bug-report fields: - **What happened:** Agent pairs orchestrated through Paperclip heartbeats complete identical task sets 1.8–2.2× slower (wall-clock) than the same harness invoked directly. Profiling attributed the gap to per-heartbeat orchestration overhead: every timer wake discards the task session (full session start + re-orientation), and in degraded environments agents repeatedly retry the same failing control-plane write. - **Expected behavior:** Heartbeat orchestration should add minimal wall-clock overhead on top of the underlying harness; issue-scoped continuation work should not pay a fresh-session tax each interval. - **Steps to reproduce:** Run a fixed benchmark task set once through Paperclip issue heartbeats and once via direct harness invocation with the same model/config; compare wall-clock totals. - **Version/commit:** master @ |
||
|
|
030dd9d15c |
feat(adapter-utils): provider-delegable syncIn seam with ordered post-upload commands (#10340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter/runtime layer has to move files into sandboxes safely and efficiently > - The current sync-in path needs a provider-delegable seam so providers can use their native upload transport when available > - The upload contract also needs ordered post-upload commands so extracted content can be finalized fail-fast after transfer > - The fallback path still has to preserve current behavior when the provider does not expose native sync verbs > - This pull request adds the contract and runtime seam for provider-delegable sync-in, plus the single-stream collapse flag > - The benefit is fewer round trips, a cleaner provider-owned upload path, and a compatible fallback for existing runners ## Linked Issues or Issue Description No public GitHub issue exists for this change. The underlying feature request is described below in the repository's feature-request format. ### Problem or motivation Paperclip needs a sync-in path that lets each provider choose the best available upload transport instead of forcing the harness to orchestrate uploads the same way every time. The runtime also needs a way to describe ordered post-upload commands so providers can finalize extracted content fail-fast after transfer. ### Proposed solution Extend the sync contract with ordered post-upload commands, forward that contract through the plugin and environment runtime layers, and make client syncIn always available. When a provider advertises native sync verbs, the client should delegate to that transport; otherwise it should fall back to the existing tarball/write/extract behavior and then run the post-upload commands in order. ### Alternatives considered Keeping upload orchestration entirely host-side would avoid a contract change, but it would block provider-specific transport optimizations and keep the harness responsible for a path the provider can do more efficiently. A separate post-upload API would add another surface without improving the existing sync flow. ### Roadmap alignment This work aligns with the broader runtime and adapter roadmap because it improves provider integration without changing the external product model. It is an additive contract change that preserves backward compatibility for providers that do not expose native sync verbs. ### Additional context The fallback path still needs to preserve existing observable behavior, including command ordering, cwd confinement, and fail-fast execution. The single-stream progress flag is part of the same transport improvement so smaller writes can collapse to a single round trip when the runner supports it. ## What Changed - Added ordered `postUploadCommands` support to the sync operation contract and SDK mirror. - Plumbed the sync-in contract through the plugin and environment runtime layers. - Implemented a runtime client `syncIn` path that delegates to native provider transport when available, otherwise uses the generic tarball/write/extract fallback. - Preserved fail-fast execution of ordered post-upload commands in the fallback path. - Flipped the sandbox runner's single-stream stdin progress flag to collapse small `writeFile` operations to a single round trip. - Added and updated tests for contract forwarding, fallback behavior, cwd rejection, fail-fast behavior, and single-stream collapse. ## Verification - `pnpm --filter @paperclipai/plugin-sdk exec vitest run protocol.postupload.test.ts` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run environment-sync-negotiation.test.ts` - `pnpm --filter @paperclipai/adapter-utils exec vitest run command-managed-runtime.test.ts` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target.test.ts` - Local typecheck and targeted suite runs reported in the handoff passed before PR creation. ## Risks - The new fallback path could diverge from the previous inline upload behavior if the tarball/extract contract changes. - Provider-native sync handling may expose provider-specific edge cases if a runner advertises sync verbs but does not fully honor the contract. - The single-stream flag changes transport behavior for small uploads, so regressions would likely show up as round-trip or upload failures. ## Model Used OpenAI Codex (GPT-5, tool-using coding agent). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a3b293e26d |
perf(adapter-utils): parallelize the two Daytona sandbox bridge setups (#10334)
## Thinking Path > - Paperclip runs AI-agent work through adapter-backed execution paths. > - Sandbox-backed local adapters pay startup overhead in two host-side bridge setups. > - Those setups were happening serially even though most of the work is independent. > - The remaining dependency is only the merged env that must reach the process-session launch. > - This pull request overlaps the independent setup work while preserving that launch-time dependency. > - The benefit is lower end-to-end startup latency without changing execution semantics. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Problem / motivation: - Sandbox-backed local adapter startup waited for two largely independent bridge setups in series, so users paid roughly the sum of both setup times. - The only hard dependency is the merged Paperclip env that must reach the process-session launch. Proposed solution: - Start the paperclip callback bridge and the process-session bridge concurrently. - Keep the env merge as the single sequencing point before process-session launch. - Stop whichever bridge started if either startup path fails. - Keep the concurrent bridge telemetry duration-only so shared runner counters are not double-counted. Alternatives considered: - Keep the bridges serial, which preserves the current telemetry shape but leaves the startup latency unchanged. - Move env merging earlier, which would complicate launch ordering and risk changing runtime semantics. Roadmap alignment: - This is the approved Daytona start-speedup work, focused on bridge startup concurrency rather than broader runtime behavior. ## What Changed - Started the paperclip callback bridge and the process-session bridge concurrently in `execute.ts`. - Added a memoized env finalizer so the process-session launch still waits for the merged paperclip env at the correct moment. - Updated `execution-target.ts` so the process-session bridge can accept a deferred env resolver and consume it only at launch time. - Added tests that cover the overlapped launch path and the failure-cleanup path when one bridge start fails. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/adapter-utils test` - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` - `git log --oneline origin/master..origin/perf/parallelize-daytona-bridge-setups` - `git diff --stat origin/master...origin/perf/parallelize-daytona-bridge-setups` ## Risks - A regression in the launch-time env merge could change what the process-session bridge sees at startup. - The concurrent start/cleanup logic could leak a bridge if the stop paths were incorrect, which is why the failure cleanup test matters. - This is low-to-moderate risk because the change is isolated to adapter-utils startup plumbing and is covered by targeted tests. ## Model Used OpenAI Codex, GPT-5, tool-using code execution mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
197718bc00 |
feat(acpx): per-step round-trip + provider-latency attribution for sandbox startup (#10222)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-backed runs need more precise startup observability so operators can see where time is spent before an adapter is invoked > - Aggregate startup timing hides which boundary is actually slow, especially in remote execution where the bottleneck can move between host round-trips, provider-boundary latency, and handshake phases > - This pull request keeps the existing startup timing channel additive while attributing the latency to the specific startup steps that caused it > - The benefit is better diagnosis of sandbox startup regressions without changing control flow or introducing a schema migration ## Linked Issues or Issue Description Refs: #10204 This PR extends the existing sandbox run-startup timing observability with per-step round-trip and provider-latency attribution for the Daytona startup path. It keeps the event payload additive and free-form, and it leaves the control flow, database schema, and external adapter interfaces unchanged. ## What Changed - Added per-step round-trip counting for the host-to-sandbox execute seam - Added provider-boundary duration accumulation for the Daytona execute and re-fetch steps - Split the ACP handshake timing into `createRuntimeMs` and `ensureSessionMs` while preserving the warm-handle skip - Kept the startup timing payload additive and did not add a schema migration ## Verification - `adapter-utils` acpx-engine and startup-timing suites: pass - Daytona plugin suite: pass with mocked SDK and injected-clock duration assertions - `server` environment-execution-target suite: pass - `tsc --noEmit` for adapter-utils and server: pass - Git validation: fetched `origin/feat/sandbox-start-step-timing-attribution`, confirmed it matches the authorized submit SHA `53c573266e618af05c57a0720aaa9d9e0452de61`, and confirmed the branch contains only the expected commit on top of `origin/master` - Searched GitHub for duplicate or related PRs/issues; found one closely related merged PR and no open duplicate on this branch - Checked `ROADMAP.md`; the broad sandbox-agent roadmap section does not call out this specific startup-timing attribution work as a duplicate ## Risks - Low risk: the change is additive and only enriches existing timing data - Downstream consumers that assume aggregate-only startup timing may need to tolerate the additional per-step fields - The finer spawn/initialize/session split remains a follow-up in the external ACP client because that hook is not available here yet ## Model Used OpenAI GPT-5, tool-using coding agent ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2568bdecc4 |
fix(adapter-utils): keep sandbox proxy sockets within Linux path limit (#10221)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work > - Local process adapters can enforce network allowlists through a Unix-socket proxy inside the Linux sandbox > - The proxy socket was created under `os.tmpdir()`, which can be a deeply nested run-specific directory > - Linux Unix-domain socket paths are limited to 107 usable bytes, so deep temporary paths can be silently truncated during bind > - A truncated proxy socket resets every confined outbound connection, including the model provider API, before the agent can do work > - This pull request creates proxy sockets under the short `/tmp` path when possible and validates the path length before bind > - The benefit is reliable sandboxed egress under deep `TMPDIR` values and an explicit error instead of an opaque connection-reset storm ## Linked Issues or Issue Description ### Pre-submission checklist - [x] I searched existing open and closed issues and this is not a duplicate. - [x] I can reproduce this on upstream `master`. - [x] I confirmed the error originates in Paperclip's local-process sandbox proxy. ### What happened? Sandboxed local-process agents using `networkScope: "allowlist"` could lose all proxied egress when `TMPDIR` was deeply nested because `proxy.sock` exceeded Linux's Unix socket path limit. The truncated bind surfaced as repeated connection resets, including for provider API traffic. ### Expected behavior Proxy creation should use a Linux-safe short path and fail explicitly if no safe socket path can be created. ### Steps to reproduce 1. Set `TMPDIR` to a path long enough that `<TMPDIR>/paperclip-network-sandbox-XXXXXX/proxy.sock` exceeds 107 bytes. 2. Build or run a local-process sandbox with `networkScope: "allowlist"`. 3. Make an HTTP or HTTPS request through the generated sandbox proxy. 4. Observe connection resets from the overlong Unix socket path on affected code. ### Paperclip version or commit Upstream `master` before this change. ### Deployment mode Other — Linux local-process adapters using Bubblewrap network allowlisting. ### Installation method Built from source. ### Agent adapter(s) involved Codex and any other local-process adapter using the shared sandbox utility. ### Operating system Linux. ### Additional context GitHub search found no duplicate public issue or pull request for this defect. ## What Changed - Add a Linux Unix-socket byte-length guard that reports the unsafe path before binding. - Create network proxy temporary directories under `/tmp` first, with `os.tmpdir()` as a validated fallback. - Use the safe temporary-directory helper at both allowlist proxy creation sites while preserving trusted URL handling. - Add regression coverage that sets a deliberately deep `TMPDIR` and verifies the working socket remains within 107 bytes under `/tmp`. ## Verification - `node /srv/paperclip/home/paperclipai/paperclip/node_modules/vitest/vitest.mjs run packages/adapter-utils/src/local-process-sandbox.test.ts` — 8 passed, 4 environment-gated tests skipped. - `git diff --check origin/master...HEAD` — passed. - Greptile Review — passed after reviewing 2 files with 0 comments and no unresolved threads; this repository integration did not emit a separate numeric confidence score. - Full GitHub CI matrix — passed after one rerun of an unrelated ACPX `ENOTEMPTY` cleanup flake. - Package typecheck was attempted, but the existing shared install cannot resolve `acpx/runtime`; the failure is outside the changed files and is expected to be covered by CI's clean dependency install. ## Risks - Low risk: the change is limited to Linux sandbox proxy temporary-directory selection and validation. - Systems without a writable `/tmp` fall back to `os.tmpdir()` only when the resulting socket path is safe; otherwise startup now fails loudly instead of producing connection resets. - No schema, API, UI, migration, or documentation behavior changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.3-codex`, tool-enabled coding agent with managed reasoning and code execution; context-window size is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e00818574 |
fix(runtime): support in-place workspace realization (#10230)
## Thinking Path > - Paperclip is the control plane people use to coordinate AI agents and their execution environments. > - Environment realization decides where an agent runs and which filesystem and toolchain are authoritative. > - Copy-based realization is unsafe for container-anchored tasks because absolute paths such as `/app` can point outside the synchronized tree and task-specific binaries may be absent. > - That mismatch can let an agent successfully verify work in a phantom writable path while sync-back silently discards the result. > - Existing task environments already provide the authoritative filesystem and toolchain, so they should be executed in place rather than copied. > - Copy mode still needs explicit confinement rules so aliases target the synchronized workspace and unsynchronized writable paths fail visibly. > - This pull request adds typed realization metadata, propagates the authoritative root through orchestration, and teaches Codex to honor it. > - The benefit is that container-anchored tasks operate on verifier-visible state with the intended tools, while copy mode remains safe and backward compatible. ## Linked Issues or Issue Description No public GitHub issue exists for this defect. GitHub duplicate searches for in-place execution, workspace realization, and authoritative workspace roots found no related pull request to link. ### What happened? Environment-backed agent runs were always realized through a copied workspace. Tasks anchored to absolute container paths could therefore write outside the synchronized tree, and task-provided toolchains were unavailable in the copy. A run could report success even though sync-back discarded its output. ### Expected behavior Existing task environments should run against their real authoritative root and toolchain. Copy-mode runs should map declared absolute aliases into the synchronized tree and reject writable paths that cannot be restored. ### Steps to reproduce 1. Run a Codex task environment whose required files live under `/app` or `/workspace` and whose required binary exists only in the task container. 2. Observe that copy realization changes the effective filesystem/toolchain or permits writes outside the synchronized root. 3. Complete and verify the task inside the agent sandbox. 4. Observe that the verifier cannot see out-of-tree artifacts or that task-specific commands were unavailable. ### Reproduction context - Paperclip commit: `3a16b91217483d2c233926de5b7f7bc3a1077924` - Deployment: built from source in a task-container execution environment - Adapter: Codex local - Database: not database-related - Access context: agent execution ## What Changed - Added typed `copy | in_place` workspace-realization metadata, authoritative roots, confined aliases, and outbound restore paths to shared execution-target contracts. - Selected in-place realization for existing task environments and skipped archive prepare/restore when the authoritative environment is used directly. - Propagated the authoritative root into adapter context so Codex uses it for cwd and `PAPERCLIP_WORKSPACE_*` semantics, including ACP execution. - Bound copy-mode aliases such as `/app` to the synchronized workspace and rejected writable out-of-tree paths without explicit restore mappings. - Added focused regression coverage while preserving existing copy-mode archive restore behavior. ## Verification - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/execute.remote.test.ts server/src/__tests__/environment-run-orchestrator.test.ts` — 48 passed, 4 skipped. - `pnpm -r typecheck` — passed. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY -u AWS_SESSION_TOKEN pnpm test:run` — passed across all general and serialized Vitest shards. - `pnpm build` — passed. - Codex `k=1` acceptance run completed July 24, 2026 at 23:54:30 UTC with 4 completed, 0 exceptions, and mean reward 1.0: `build-cython-ext`, `openssl-selfsigned-cert`, `prove-plus-comm`, and `sqlite-db-truncate` each received terminal grade 1.0 against real task-environment paths and toolchains. ## Risks - In-place mode deliberately exposes the authoritative task root to the adapter; incorrect environment metadata could point execution at the wrong root. Typed metadata and focused orchestration tests cover selection and propagation. - Copy-mode writable-path validation is stricter and may reject previously accepted unsafe configurations. The rejection is intentional and produces a visible error instead of silently losing output. - The acceptance run is focused on four Codex task-environment workloads, not a broad cross-adapter benchmark. Existing copy-mode archive tests and the full repository suite remain green. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact model ID and context-window size were not exposed to this runtime. Capabilities used: extended reasoning, repository editing, shell execution, test/build execution, Git, GitHub CLI, and Paperclip API tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b996b71a38 |
Deduplicate wake-payload issue descriptions and compact resume deltas (#10216)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Heartbeat wake payloads and the task-context markdown are the two channels that deliver an issue's brief into an agent's prompt > - #10151 fixed wake-prompt-only adapter lanes waking without the issue description by adding it to the structured wake payload > - That left the description delivered twice per prompt on lanes that also inject the task-context markdown, and re-delivered in full on every resume wake, permanently bloating persistent-session context > - This pull request makes the task markdown the single description carrier on lanes that use it, and omits the description from non-assignment resume deltas on all lanes while keeping it for assignment-shaped and recovery wakes > - The benefit is that every lane receives the brief exactly once when it needs it, and long-lived sessions stop re-paying the full brief in tokens on every wake ## Linked Issues or Issue Description Refs #10151 Related prior work: #2883, #8402 (earlier description-delivery attempts referenced by #10151). I searched the PR list for open work on wake-payload description handling and found none besides the merged #10151. **Bug:** After #10151, adapters that inject the `Paperclip task context` markdown (ACPX engine lanes, claude-local CLI, hermes server and gateway) receive the issue description twice in a single prompt — once in the wake prompt's `Issue description:` block and once in the task markdown. Separately, resume deltas re-send the full description (up to 12k characters) on every wake even though the persistent session already received it. **Expected behavior:** The description appears exactly once per prompt on every lane, and resume deltas only carry it when the resuming session may not have seen the brief (assignment-shaped or recovery wakes), leaving an explicit fetch breadcrumb otherwise. **Reproduction:** Wake a claude-local or ACPX agent on an issue with a description and inspect the assembled prompt: the description text appears in both the wake-payload block and the task-context block. Wake the same session again via a comment: the full description is present again in the resume delta. **Affected version:** Current `master` (with #10151 merged). **Deployment mode:** Adapter-backed heartbeat execution, local and sandboxed lanes. ## What Changed - `renderPaperclipWakePrompt` accepts `suppressIssueDescription`; the four task-markdown lanes pass it so the task markdown stays the single, uncapped description carrier there. - Non-assignment resume deltas omit the description and emit `- issue description: omitted from this resume delta; fetch the issue if you need the latest brief`. Assignment-shaped reasons (`issue_assigned`, `issue_reopened_via_comment`, `issue_recovery_action_restored`, `issue_tree_restored`) and recovery wakes still deliver the full brief. - `buildPaperclipTaskMarkdown` gains `includeDescription`; the server now also publishes `context.paperclipTaskMarkdownCompact` (description stripped, directives and wake comment kept), and the new `selectPaperclipTaskMarkdown` helper picks the right variant under the same resume rules, falling back to the full markdown when no compact variant exists (version skew safety). - The wake prompt's description block now carries the same user-authored trust framing the task markdown already had. ## Verification - `npx vitest run packages/adapter-utils/src/server-utils.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/acp.test.ts server/src/__tests__/heartbeat-context-summary.test.ts` — 137 tests passed, including new coverage for suppression, resume omission plus breadcrumb, assignment-shaped resume inclusion, compact-variant building, variant selection, and an end-to-end ACPX prompt-assembly test asserting the description appears exactly once on fresh wakes and not at all on comment resumes. - `npx vitest run` in `packages/adapters/hermes` — 59 tests passed, including a gateway execute-level test asserting the brief is sent exactly once on fresh runs and not re-sent on stable-session resumes. - `tsc --noEmit` in `packages/adapter-utils`, `packages/adapters/claude-local`, `packages/adapters/hermes` — clean; `server` matches the `master` baseline exactly (pre-existing plugin-sdk resolution errors only, none in touched files). - Pre-existing failures confirmed identical on clean `master`: claude-local `execute.remote.test.ts` / `test.probe.test.ts`, adapter-utils `mcp-isolation.integration.test.ts` (requires a newer local Claude CLI). ## Risks - Behavioral shift, prompt-only: a resumed session woken by a comment on an issue it never handled (rare — assignment wakes normally precede comment wakes) would not get the inline description; the breadcrumb plus the standard issue-fetch path covers it. - Additive context key (`paperclipTaskMarkdownCompact`); older adapters ignore it and newer adapters fall back to the full markdown when it is absent, so mixed-version deployments degrade to current behavior. - No schema, migration, or API changes; the structured wake-payload JSON shape is unchanged. - Known follow-up deliberately out of scope: openclaw embeds the raw wake-payload JSON (which still contains the description) in prompt text for machine parsing. The hermes-gateway lane is handled: it detects stable-session resumes (issue/agent session-key strategy plus a stored prior session id), compacts the task markdown, and omits the description from its prompt-embedded JSON copy. > This is a focused correctness/efficiency fix to existing wake plumbing and does not overlap with planned roadmap feature work. ## Model Used - Anthropic Claude Fable 5 (`claude-fable-5`), extended thinking enabled, with repository tool use, shell execution, and local test execution via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details (execution-workspace branch, same convention as merged #10202) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (code-level docs; no user-facing docs affected) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3c004d1b8 |
Fix issue descriptions in structured wake payloads (#10151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Heartbeat wake payloads provide the scoped task context an agent needs before it can act safely > - Assignment wakes already loaded the issue description for task markdown, but the structured wake-payload builder dropped it > - Agents reading `PAPERCLIP_WAKE_PAYLOAD_JSON` could therefore see a missing brief while also being told no fallback fetch was needed > - Long descriptions also need a bounded representation so wake environments and prompts remain safe > - This pull request carries the description through the server and adapter contract, and marks truncated descriptions as requiring fallback fetch > - The benefit is that agents receive the actual brief instead of inventing requirements from the title ## Linked Issues or Issue Description Fixes: #5844 Fixes: #2882 Related prior attempts: #2883 and #8402. This change adds focused regression coverage and enforces the missing long-description fallback invariant. **Bug:** Issue-assignment wake payloads omitted the issue description from the structured payload even when the issue had a populated description. **Expected behavior:** The structured wake payload includes the issue description. If the description must be truncated for payload size, `fallbackFetchNeeded` is `true`. **Reproduction:** Assign an issue with a description to an agent and inspect `PAPERCLIP_WAKE_PAYLOAD_JSON`; before this change, `issue.description` was absent while `fallbackFetchNeeded` could remain `false`. **Affected version:** Reproduced on current `master` before this patch. **Deployment mode:** Adapter-backed heartbeat execution, including local Codex agents. ## What Changed - Include `issues.description` in the server wake-payload query and supplied issue summaries. - Bound inline descriptions at 12,000 characters and force fallback fetch when truncation occurs. - Preserve and render description metadata through shared adapter normalization and prompt rendering. - Add focused tests for long-description fallback and exact brief-string rendering. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-agent-session-message.test.ts packages/adapter-utils/src/server-utils.test.ts` — 81 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Low risk: the payload shape is additive. - Very long descriptions are truncated at 12,000 characters; the payload explicitly requests a fallback fetch for the full brief. - Prompt size increases by the issue-description length for scoped wakes, bounded by the same limit. > This is a focused correctness fix and does not overlap with planned roadmap feature work. ## Model Used - OpenAI GPT-5.4 via Codex CLI, with reasoning, repository tool use, shell execution, and test execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e4662b1f9d |
fix: preserve control plane access in Codex sandbox (#10152)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work > - Local coding adapters can be confined with filesystem and network sandbox policies > - Codex confinement proxied only user-allowlisted hosts, so it denied Paperclip's own run API and managed MCP endpoints > - The same proxy returned untyped plaintext denials, which MCP clients could treat as a fatal unexpected content type > - Host-only `PAPERCLIPAI_CMD` values could also leak into agents even when the referenced checkout module did not exist > - This pull request gives the sandbox an explicit trusted-URL channel for Paperclip endpoints, returns structured JSON errors, and removes the inherited host CLI pointer > - The benefit is confined Codex agents retain control-plane access without broadening the operator's external network allowlist or crashing on policy denials ## Linked Issues or Issue Description Refs #3802 Related foundation: #9504 ### What happened? A `codex_local` agent configured with `networkScope: "allowlist"` and an external-only allowlist could not reach its own Paperclip API or Paperclip-managed MCP endpoints. The sandbox proxy returned a `403` plaintext response without `Content-Type`, and an inherited `PAPERCLIPAI_CMD` could point at a missing checkout-local CLI module. ### Expected behavior Paperclip's run-scoped API and managed MCP endpoints remain reachable regardless of the user external allowlist. Policy denials are valid structured JSON responses with an explicit media type, and host-only CLI pointers are not inherited by agent processes. ### Steps to reproduce 1. Configure a `codex_local` agent with `networkScope: "allowlist"` and `networkAllowlist: ["api.openai.com"]`. 2. Run the agent and request its Paperclip issue API or a Paperclip-managed MCP endpoint. 3. Observe the sandbox proxy deny the request with an untyped plaintext `403` response. ### Environment - Paperclip commit: `f49a3f99` originally exhibited the defect; fix is based on current `master`. - Deployment: local source build on Linux. - Adapter: Codex. - Database: not database-related. - Access context: agent bearer/run-scoped credentials. ## What Changed - Added internal trusted URL rules to the local network allowlist proxy and supplied the Codex run API plus managed MCP endpoints. - Returned JSON error envelopes with `Content-Type` and `Content-Length` for HTTP and CONNECT policy denials. - Removed inherited `PAPERCLIPAI_CMD` from child process environments while preserving explicitly constructed runtime variables. - Added focused proxy and environment sanitizer regression tests. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/local-process-sandbox.test.ts packages/adapter-utils/src/server-utils-env.test.ts --reporter=verbose` - 2 test files passed; 8 tests passed; 4 platform-dependent tests skipped. - `git diff --check` - Package typecheck was attempted; it reaches unrelated current-`master` type drift in untouched files (`spawnCwd` in `adapter-utils`, and staged-runtime ACP types in `codex-local`). ## Risks - Low risk: trusted access is restricted to exact HTTP(S) hostname and port pairs derived from Paperclip-provided URLs. - Invalid or non-HTTP trusted URL values are ignored rather than broadening access. - Denial response bodies change from plaintext to structured JSON; status codes remain unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agent (exact model ID and context window are not exposed to this runtime), reasoning and terminal/tool execution enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
762ce5b4ef |
feat(acpx): per-step timing observability for sandbox run-startup (#10204)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-backed runs need observable startup behavior so operators can see where time is spent before an adapter is invoked > - The current startup path only surfaced aggregate timing, which makes it hard to identify the slow boundary in the bring-up sequence > - That gap matters because sandbox startup latency is often dominated by one specific step, and aggregate timing hides the bottleneck > - This pull request adds per-step startup timing events for the named sandbox bring-up boundaries > - The benefit is more precise observability with no control-flow change and no schema migration ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting (multiple of the above) ### Problem or motivation Sandbox run startup only exposed aggregate timing. That makes it hard to identify which bring-up boundary is responsible for slow starts, especially in remote or sandboxed execution where the bottleneck can move between workspace setup, skill reconciliation, bridge setup, and adapter handshake. ### Proposed solution Emit a structured timing event for each named startup boundary before the adapter is invoked, so the existing run-event stream carries per-step duration data. This keeps the event path additive and lets operators see which step dominates startup latency without changing control flow or introducing a schema migration. ### Alternatives considered - Keep only the aggregate startup duration: simpler, but it hides the bottleneck and makes regression analysis much harder. - Add a new telemetry sink or schema field: rejected because the existing run-event payload already carries structured event data and does not need a new storage path. - Log unstructured text for each step: rejected because it is harder to query and aggregate than a structured `step` + `durationMs` event. ### Roadmap alignment This fits the roadmap direction around cloud / sandbox agents and enforced outcomes by improving observability for sandboxed execution without changing the control plane model. The roadmap section is broad, but it does not call out this specific startup-timing work as a planned duplicate. ### Additional context This PR is intentionally additive. It records timing for the named startup boundaries in the existing event stream and leaves the bridge, database shape, and adapter invocation order unchanged. ## What Changed - Added a `measureStartupStep` helper that times a startup step, emits one structured `run.startup.step` event, and rethrows failures after recording duration - Wrapped the seven sandbox bring-up boundaries in `execute.ts` so the structured timing covers each named step before adapter invocation - Added unit coverage for the helper and integration coverage for the startup-step events in the adapter-utils execute path - Kept the event path additive, with no bridge change and no database migration ## Verification - `tsc --noEmit` for `@paperclip/adapter-utils` - `pnpm test` in `packages/adapter-utils` equivalent suite coverage: 292 passed, 4 skipped - Adjacent server event/log-store suites: `run-log-store.test.ts` and `heartbeat-run-log.test.ts` passed (11 total) - Git validation: fetched `origin/feat/sandbox-startup-step-timing`, confirmed it matches the authorized submit SHA, and confirmed `origin/master..origin/feat/sandbox-startup-step-timing` contains the expected single commit - Searched GitHub for duplicate or related open PRs/issues and found no overlapping open items - Checked `ROADMAP.md`; the roadmap covers sandboxed environments generally, but does not call out this specific startup-timing observability work as a planned duplicate ## Risks - Low risk: the change is additive and only emits additional structured events - If downstream consumers assume startup events are aggregate-only, they may need to ignore or account for the new `run.startup.step` entries - Timing is measured via the injected clock and event emission happens in a `finally`, so failures still report duration before rethrowing ## Model Used OpenAI GPT-5, tool-using coding agent ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
caae2778f0 |
fix: deliver plugin agent session turns and replies (#10137)
## Thinking Path > - Paperclip manages agent execution through heartbeat runs and adapter-specific sessions > - Plugins can open an agent session and send a conversational message through the host service > - The host previously stored that message only in opaque wake payload metadata, so local adapters never saw it in their CLI prompt > - The host also forwarded run log chunks but did not expose the persisted final assistant text as the session reply > - This pull request defines both sides of the session contract in the shared wake renderer and terminal run event > - The benefit is that local adapters receive the actual conversational turn and plugins receive one canonical final reply ## Linked Issues or Issue Description Related context: Refs #629 and Refs #2880 describe adjacent `claude_local` final-text visibility failures. They concern issue comments rather than plugin agent sessions, but exercise the same need for a canonical persisted run summary. Companion consumer change: paperclipai/paperclip-gateway#3. Bug description: - **Observed:** calling the plugin host's `agents.sessions.sendMessage()` with `prompt: "hello"` woke a `claude_local` agent, but the generated CLI prompt omitted `hello`. On completion, the session emitted log chunks and a generic `Run completed` done event, so callers could not reliably recover the assistant reply. - **Expected:** the prompt becomes the user-supplied conversational turn for that agent session, and the successful terminal event carries the run's canonical final user-facing assistant text. - **Reproduction:** create a plugin agent session for a local adapter, call `sendMessage()` with a non-empty prompt, inspect the adapter prompt and terminal session event. - **Affected baseline:** `b517b887a` on `master`, local trusted deployment with plugin host services and `claude_local`; `codex_local` shared the wake-rendering gap because both use the common Paperclip wake prompt renderer. ## What Changed - Added a typed `agentMessage` wake payload rendered by the shared adapter prompt path used by `claude_local`, `codex_local`, and other local adapters. - Labeled session content as user-supplied and explicitly non-authoritative: it cannot expand authorization, permissions, task scope, or company boundaries. - Preserved ordinary heartbeat behavior by omitting the section when no agent-session message exists. - Added canonical `finalText` to terminal heartbeat status events from the already-persisted run summary/result/message. - Defined successful `AgentSessionEvent.message` as the canonical final user-facing reply (or `null`) and forwarded it on the terminal `done` event. - Added host, wake-renderer, normal-heartbeat, and terminal-reply regression coverage. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-agent-session-message.test.ts server/src/__tests__/heartbeat-run-status-payload.test.ts server/src/__tests__/plugin-agent-sessions.test.ts server/src/__tests__/heartbeat-run-summary.test.ts` — 87 passed. - `pnpm -r typecheck` — passed across all 31 workspaces. - `pnpm build` — passed. - `pnpm test:run` — 2,860 passed, 1 skipped, 3 unrelated failures: two existing macOS temp-path alias assertions (`/tmp` vs `/private/tmp`) in workspace branch-containment tests and one reproducible auto-port runtime-service adoption failure. The same three failures reproduce when the two files run alone; none touch this change. - Live Slack verification intentionally remains operator-gated because it requires rebuilding/restarting the host. ## Risks - User-controlled chat text now reaches the model prompt, which is an intentional prompt-injection surface. The renderer labels it as untrusted conversational content, while the existing plugin/session company checks and caller authorization remain unchanged. - `finalText` is added to company-scoped heartbeat status events. It is derived from the same persisted summary/result/message already used for run comments; no raw stdout or secrets are added. - Consumers that ignore the new field remain compatible, and successful runs without usable final text still emit `message: null`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex (GPT-5), agentic reasoning with repository/tool use and code execution; context-window size is not surfaced in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b517b887ad | fix(acpx): decouple host proxy spawn cwd from in-sandbox remoteCwd (#10122) | ||
|
|
d36ea13e08 |
feat(acpx): stage once per remote ACP session - compatible resume reuses staged runtime, no cross-session credential reuse (#10089)
## Thinking Path > - Paperclip coordinates autonomous agent work across isolated company-scoped sessions > - The ACP remote lane stages workspaces and managed home state inside the sandbox so sessions can resume safely > - If a compatible resume re-staged everything every time, it would waste work and risk inconsistent session reuse behavior > - If an incompatible resume reused the wrong staged runtime, it could cross session boundaries or leak credentials > - This pull request keeps the session fingerprint as the scoping key and adds a staged-runtime cache keyed to that fingerprint > - Compatible resumes now reuse the already staged runtime while incompatible fingerprints stage fresh > - The benefit is faster safe resumes without weakening session isolation or credential separation ## Linked Issues or Issue Description ### Problem or motivation The ACP remote lane currently needs to preserve safe resume behavior without repeatedly re-staging work that is already valid for the same session. The failure mode to avoid is letting one session reuse another session's staged workspace or credentials. ### Proposed solution Keep the session fingerprint as the scoping key and add a staged-runtime cache keyed to that fingerprint. When the fingerprint matches, reuse the already staged workspace and managed home. When the fingerprint changes, stage fresh. ### Alternatives considered - Always restage on resume: safest mechanically, but wastes work and breaks the compatible-resume optimization. - Reuse without fingerprint scoping: too risky because it could cross session boundaries. ### Roadmap alignment This is a narrow implementation change for the ACP remote lane and does not duplicate any broader roadmap item I could find in `ROADMAP.md`. ### Additional context The change preserves the existing session fingerprint contents and codex auth copy-back cadence while adding tests for compatible reuse, incompatible fresh staging, no cross-session credential reuse, and failed-turn eviction. ## What Changed - Added a staged-runtime cache in the ACP remote execution path keyed by the session fingerprint. - Reused the existing staged workspace and managed home for compatible resumes. - Kept incompatible fingerprints on the fresh staging path. - Preserved the existing session fingerprint contents and codex auth copy-back cadence. - Added tests for compatible reuse, incompatible fresh staging, no cross-session credential reuse, and failed-turn eviction. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` (74/74 pass, including the active-turn lease regression) - `pnpm exec tsc -p packages/adapter-utils/tsconfig.json --noEmit` - Verified the PR touches only `packages/adapter-utils/src/acpx-engine/execute.ts` and `packages/adapter-utils/src/acpx-engine/execute.test.ts` ## Risks - A cache eviction bug could cause an unavailable or partially staged runtime to be reused. - If the fingerprint scoping regressed, a session could incorrectly reuse another session's state. - The change is isolated to the ACP remote lane, but it still affects resume behavior for that path. ## Model Used OpenAI Codex, GPT-5, reasoning-capable coding agent, tool-enabled session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e1cc63328a |
feat(acpx): per-adapter managed-home seed (x3) + Codex auth copy-back for remote ACP lane (#10073)
## Thinking Path > - Paperclip is a control plane that orchestrates AI agents and adapter execution for human operators. > - Agents run across local and remote execution contexts, and reliability in remote sessions depends on consistent adapter bootstrapping. > - The ACP path must prepare per-adapter runtime homes so managed credentials and config are available in sandboxed runner environments. > - Before this change, the new remote ACP lane did not yet consistently stage managed-home paths for all affected adapters or restore Codex auth state on teardown. > - We added a shared per-adapter seam in the ACP engine, then wired Codex, Claude, and Gemini remote lanes to seed managed homes and remap to in-sandbox locations. > - Codex additionally reuses the existing atomic auth restore flow to copy auth back on teardown, matching CLI behavior. > - This improves remote runner parity with existing CLI behavior and avoids credential drift in shared-code-path executions. ## Linked Issues or Issue Description - This change continues the ACP remote managed-home work by completing per-adapter remote bootstrapping and Codex auth restore behavior for the remote ACP lane. - It specifically covers: `acpx-engine`, `codex-local`, `claude-local`, and `gemini-local`. - Related prior work in this repo: PR #10070. ## What Changed - Add a per-adapter remote managed-home seam (`prepareRemoteManagedHome`) in `acpx-engine` and thread it through ACP execution options. - Implement Codex ACP preparation to: - stage `CODEX_HOME` (auth/config/skills) into a sandboxed remote home, - repoint `CODEX_HOME` to the in-sandbox path, - and wire teardown copy-back through the existing managed-auth restore path. - Implement Claude ACP preparation to seed a sanitized `config-seed` into `CLAUDE_CONFIG_DIR` and remap that directory to the sandbox root. - Implement Gemini ACP preparation to seed `~/.gemini/skills`, set `HOME` to managed runtime root, and preselect API-key auth in `settings.json`. - Keep local and runner-less ACP→CLI behavior unchanged by only invoking the remote managed-home seam when running in remote ACP mode. - Preserve existing authorization and activity boundaries in the shared engine and adapter layers. ## Verification - `git log --oneline origin/master..origin/feat/acp-remote-managed-home-seed` confirms only the expected 4 commits. - `tsc --noEmit` is clean in `adapter-utils`, `codex-local`, `claude-local`, and `gemini-local`. - Vitest selection used during validation passed (120 tests across the ACP-related suites). ## Risks - If remote sandbox teardown occurs after token rotation but before restore timing, Codex credentials can become stale and require re-auth on next startup. - Partial provisioning of managed-home assets would cause adapter bootstrap failures in runner-backed ACP sessions. - This change is scoped to execution-path behavior; it should not affect CLI behavior. ## Model Used None — human-authored. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8a3cc86531 |
feat(acpx): stage workspace + route in-sandbox cwd for remote ACP lane (#10070)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The shared ACP engine
(`packages/adapter-utils/src/acpx-engine/execute.ts`) is responsible for
launching local and remote agent processes via the ACP protocol
> - On runner-backed remote sandbox (Daytona) targets, `buildRuntime`
never crossed the CLI's staging seam: it never called
`prepareAdapterExecutionTargetRuntime`, left `runtimeRootDir: null` in
both the paperclip and process-session bridges, and handed the agent the
**HOST filesystem path** as the `session/new` cwd — meaning
Claude/Gemini silently operated on a path that does not exist inside the
sandbox (Codex additionally crashes on its HOST home path, addressed in
a follow-up PR)
> - The fix must cross the staging seam for remote sandboxes, thread the
real `runtimeRootDir` through both bridges, and bind the in-sandbox
workspace path as the session cwd — without touching local ACP runs or
the runner-less ACP→CLI fallback
> - This pull request introduces a `stageAcpRemoteRuntime` helper that
calls `prepareAdapterExecutionTargetRuntime` for runner-backed remote
runs, captures `{ workspaceRemoteDir, runtimeRootDir, assetDirs,
restoreWorkspace }`, and reuses the in-sandbox `workspaceRemoteDir` as
the single `sessionCwd` across `session/new`, the fingerprint,
compatibility check, persistence, `ensureSession`, the process-session
bridge cwd, and the error path
> - The benefit is that remote ACP runs now operate in the correct
in-sandbox cwd and receive a non-null `runtimeRootDir` in both bridges —
fixing silent wrong-cwd degradation for Claude/Gemini on Daytona
targets; this is PR 1 of 3 and seeds no credential material
## Linked Issues or Issue Description
No public GitHub issue exists for this change. Inline description
follows the feature request template:
### Subsystem affected
packages/adapters — agent adapter implementations
### Problem or motivation
The shared ACP engine
(`packages/adapter-utils/src/acpx-engine/execute.ts`) never crossed the
CLI's staging seam on runner-backed remote sandbox targets. It never
called `prepareAdapterExecutionTargetRuntime`, always passed
`runtimeRootDir: null` to both the paperclip and process-session
bridges, and handed the agent the **HOST filesystem path** as the ACP
`session/new` cwd. As a result, Claude and Gemini silently operated on a
cwd that does not exist inside the sandbox; Codex crashed with a fatal
error on the HOST `CODEX_HOME` path (that crash is in a follow-up PR).
### Proposed solution
Gate on `usesRunnerBackedSandbox` (`kind === "remote" && transport ===
"sandbox" && runner`). For runs that pass the gate, call
`prepareAdapterExecutionTargetRuntime` via a new `stageAcpRemoteRuntime`
helper that ships the workspace into the sandbox and captures `{
workspaceRemoteDir, runtimeRootDir, assetDirs, restoreWorkspace }`.
Thread the real `runtimeRootDir` into both bridges. Bind a single
`sessionCwd` (the in-sandbox `workspaceRemoteDir`) and use it at every
cwd-keyed session site (`session/new`, fingerprint, compat, persist,
`ensureSession`, process-session bridge, error path) so a warm/resumable
session is reused rather than invalidated. For local runs and the
runner-less ACP→CLI fallback, `sessionCwd` resolves to the HOST cwd —
byte-identical to the previous behavior.
### Alternatives considered
Patching each per-adapter bridge individually — rejected because the bug
is in the shared engine layer and the fix belongs there so all three
adapters (Codex, Claude, Gemini) benefit without per-adapter
duplication.
### Roadmap alignment
Internal correctness fix enabling remote ACP to work as designed; no new
user-facing features. This is PR 1 of 3 in a sequential chain: PR 1
(this PR) stages the workspace and routes the cwd; PR 2 adds per-adapter
home seeding and copy-back; PR 3 wires session-lifecycle restore.
## What Changed
- **`packages/adapter-utils/src/acpx-engine/execute.ts`** — added
`stageAcpRemoteRuntime` helper that calls
`prepareAdapterExecutionTargetRuntime` for runner-backed remote
sandboxes and returns `{ sessionCwd, runtimeRootDir, stagedRuntime }`;
`buildRuntime` now uses this helper to derive `sessionCwd` (in-sandbox
`workspaceRemoteDir` for remote; HOST cwd unchanged for
local/runner-less) and threads the real `runtimeRootDir` to both the
paperclip bridge and process-session bridge; `stagedRuntime` is stashed
for the follow-up credential PR
- **`packages/adapter-utils/src/acpx-engine/execute.test.ts`** — new
engine-level unit tests: staging seam crossed with no credential asset,
non-null `runtimeRootDir` in both bridges, in-sandbox `session/new` cwd,
warm-handle reuse after the cwd change, local-unchanged; 60 tests green
- **`packages/adapters/codex-local/src/server/acp.test.ts`** — new
per-adapter test: runner-backed remote asserts `ensureSession` cwd ==
`workspaceRemoteDir`; runner-less sandbox falls back to CLI
- **`packages/adapters/claude-local/src/server/acp.test.ts`** — same
per-adapter coverage for Claude
- **`packages/adapters/gemini-local/src/server/acp.test.ts`** — same
per-adapter coverage for Gemini
## Verification
- `adapter-utils` typecheck clean; `codex/claude/gemini-local` typecheck
clean
- Engine units (`acpx-engine/execute.test.ts`): 60/60 green — staging
seam crossed with no credential asset, non-null `runtimeRootDir` to both
bridges, in-sandbox `session/new` cwd, warm-handle reuse after cwd
change, local unchanged
- Per-adapter ACP test suites (`codex/claude/gemini-local`
`acp.test.ts`): 103 tests green — runner-backed remote asserts
`ensureSession` cwd == `workspaceRemoteDir`; runner-less sandbox falls
back to CLI
- CI green on PR (in progress)
## Risks
This is PR 1 of 3 in a strictly sequential chain; it seeds **no
credential material** (no `assets`, no `installCommand`). The
per-adapter home seeding is deferred to PR 2, which consumes the
`stagedRuntime` object stashed here. The `restoreWorkspace` callback is
carried on `stagedRuntime` for PR 3's session-lifecycle wiring (see the
`stageAcpRemoteRuntime` function comment). Local ACP runs and the
runner-less ACP→CLI fallback are untouched — `sessionCwd` resolves to
the HOST cwd for those paths, preserving existing behavior. The
`stageAcpRemoteRuntime` helper is gated on `usesRunnerBackedSandbox`, so
there is no regression risk for local or CLI-lane runs.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`) via Paperclip ACPX engine —
extended context, tool use enabled, co-authored with Paperclip agent
orchestration.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`,
`feat/...`) and contains no internal Paperclip ticket id or
instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Harold Kim <harold@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
b247bf7150 |
feat(runtime): opt-in sandbox file-sync lifecycle hooks (API + provider docs) (#10013)
## Thinking Path > - Paperclip is an open-source AI-agent management platform; agents run tasks inside sandboxed environments (Daytona, Kubernetes, E2B, etc.) > - The control-plane ↔ sandbox file-transfer path flows through the `environmentExecute` seam in `protocol.ts` — the only verb available to plugins — which forces a base64-over-exec chunked loop for every file move: workspace files, assets, Codex home sync > - This transport is correct and safe, but it bypasses provider-native bulk/streaming APIs (Daytona `uploadFiles`, K8s `FastUploadInterceptor` / volume mounts), leaving significant throughput on the table for large workspaces > - The right fix is an opt-in seam extension: providers with faster native transfer declare two optional verbs; providers that do not opt in stay on the existing fallback with zero code or behavior change required > - This PR adds the first layer of that extension — two optional verbs (`environmentSyncIn` / `environmentSyncOut`) in the plugin SDK, the runtime plumbing to prefer the native path for the two clean destroy-then-replace cases, and a doc for the contract > - The core correctness invariant is byte-identical fallback: if no provider opts in, execution is exactly what ships today; `assertSyncOperationsConfined` enforces host-side path confinement for providers that do opt in > - No provider advertises the verbs yet → zero production behavior change; future PRs wire up Daytona and K8s providers against this contract ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Description follows the `feature_request` issue template: **Subsystem affected:** packages/plugins — plugin system; packages/adapter-utils — adapter runtime; server/ — EnvironmentRuntimeService **Problem or motivation:** Sandbox file transfers currently always use a base64-over-exec chunked loop regardless of what the underlying provider supports. For workspaces larger than a few MB this becomes the dominant wall-clock cost of every sandbox run, and it bypasses bulk/stream APIs that providers like Daytona already expose natively. **Proposed solution:** Add two optional, opt-in plugin hooks — `onEnvironmentSyncIn` / `onEnvironmentSyncOut` — to the plugin SDK. When a provider defines both hooks and both are advertised via the existing `supportedMethods` negotiation, the runtime prefers the native path for the two clean destroy-then-replace transfer cases; all other cases fall back to the existing byte-identical base64 transport. **Alternatives considered:** An unconditional verb would require every provider to implement or stub the verb. The opt-in / `METHOD_NOT_IMPLEMENTED` pattern (already used by `environmentExecute`) preserves backward compatibility with zero provider changes required. **Roadmap alignment:** Consistent with the ✅ "Cloud / Sandbox agents" and ✅ "Plugin system" milestones; extends the plugin seam rather than adding control-plane-level logic. **Additional context:** Searched open pull requests and issues for duplicate sandbox file-sync / native-transfer work; none found. ## What Changed - **`packages/plugins/sdk`** - `protocol.ts`: two new optional `HostToWorkerMethods` — `environmentSyncIn` / `environmentSyncOut` — plus generic `SyncOperation`, `SyncFileMapping`, and `SyncOutcome` types - `define-plugin.ts`: optional `onEnvironmentSyncIn` / `onEnvironmentSyncOut` fields on `PluginDefinition`; worker advertises each verb only when its hook is defined (else `METHOD_NOT_IMPLEMENTED`, mirroring `environmentExecute`) - `worker-rpc-host.ts`: route new verbs to plugin hooks - `index.ts`: re-export new public types - **`packages/adapter-utils`** - `command-managed-runtime.ts`: expose optional `syncIn` / `syncOut` on `CommandManagedRuntimeRunner` (available only when both verbs are advertised); add `assertSyncOperationsConfined` host-side path-confinement guard - `sandbox-managed-runtime.ts`: `SandboxManagedRuntimeClient` gains optional `syncIn` / `syncOut`; orchestrator prefers native path for default-provision asset inbound and workspace-download-into-fresh-dir outbound; all other paths keep the existing base64 fallback - `sandbox-file-sync.test.ts` (new): 234-line characterization suite — native-opt-in branch, fallback branch, `assertSyncOperationsConfined` escape-path rejection, `followSymlinks` → tar `-h` - `command-managed-runtime.test.ts`: negotiation + native-sync + confinement tests - **`server/src/services/environment-runtime.ts`**: `EnvironmentRuntimeService` delegates to `syncIn` / `syncOut`, gated on advertised support - **`server/src/services/environment-execution-target.ts`**: minor typing fix alongside the new verbs - **`doc/plugins/SANDBOX_FILE_SYNC_HOOKS.md`** (new): documents the full contract — opt-in / no-op guarantee, operation ordering, provider-may-tar, atomicity, `followSymlinks`, secret modes (0600, no window), path confinement, `operationId` opacity, resource bounds, shell-quoting ## Verification ```bash # SDK suite pnpm --filter packages/plugins/sdk test # Adapter-utils suite (includes new sandbox-file-sync characterization tests) pnpm --filter packages/adapter-utils test # Expected: 255 pass / 4 skip # Type-check across affected packages pnpm --filter packages/plugins/sdk typecheck pnpm --filter packages/adapter-utils typecheck # Server changed-file spot check: cd server && npx tsc --noEmit --skipLibCheck 2>&1 | grep -E "environment-(runtime|execution-target)" | head -20 ``` Key behavioral invariant to spot-check: with no provider opting in (the current state), run any sandbox task and confirm file-transfer behavior is byte-for-byte identical to what the pre-PR code produces. The characterization tests assert this at the unit level. ## Risks - **Zero production risk today**: no provider advertises `environmentSyncIn` / `environmentSyncOut`, so the new code paths are unreachable in production; all real traffic stays on the existing base64 fallback - **Path confinement**: `assertSyncOperationsConfined` rejects any `targetPath` that escapes the declared root — this is the primary security boundary for future providers. The test suite covers escape-path rejection - **Atomicity**: the contract delegates atomicity to providers; the doc explicitly calls out that directory-level ops are not guaranteed atomic - **Secret transport**: credential assets (e.g., Codex `auth.json`, directory mappings) continue to use the existing tar path — they do not go through the new verbs in any current provider > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: Anthropic Model: `claude-sonnet-4-6` (Claude Sonnet 4.6) Context window: 200 K tokens Capabilities: extended tool use, multi-file code generation, agentic reasoning via the Paperclip agent framework ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d31a28828b |
fix(acpx): support Windows agent spawning (#9980)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local Claude, Codex, Gemini, and custom ACP adapters run through the shared embedded ACPX engine > - That engine wrapped every local agent command in a generated Bash script to inject environment variables and filter child stderr > - Windows cannot directly spawn that Bash wrapper, and npm/pnpm ACP binaries are exposed through `.cmd` shims there > - ACPX 0.12 already supports per-session child environment variables, so the wrapper is unnecessary > - This pull request registers agent commands directly, injects env through ACPX session options, captures child stderr in-process, and adds a real Node ACP spawn smoke on Ubuntu and Windows > - The benefit is one cross-platform spawn path with a reusable smoke test instead of parallel shell-wrapper implementations ## Linked Issues or Issue Description Fixes #9941. Refs #9428 and #9771. **What happened** ACPX-backed local agents failed to start on Windows because Paperclip registered a generated POSIX `.sh` wrapper as the agent command. Windows also needs the `.cmd` npm/pnpm shim when resolving built-in ACP binaries, and symlink creation can fail with `EPERM` for seeded auth/skill files. **Expected behavior** The same ACPX engine path should spawn a real ACP agent on Windows and Linux, forward Paperclip/runtime env without mutating `process.env`, preserve filtered/unfiltered child stderr behavior, and fall back to copies where Windows symlinks are unavailable. **Steps to reproduce** Run a local ACPX adapter on Windows with the prior wrapper path. ACPX attempts to spawn the generated `.sh` file and the agent never initializes. **Deployment mode** Local Paperclip adapters using `packages/adapter-utils/src/acpx-engine/`. ## What Changed - Removed generated Bash agent/env wrappers and registered local commands directly with ACPX. - Passed the resolved child environment through ACPX `sessionOptions.env`, including resume retry paths. - Added a minimal `acpx@0.12.0` package patch exposing child stderr callbacks and allowing documented uppercase env-map keys in persisted session options. - Moved stderr tee/filter behavior in-process: raw stderr remains in the per-run file while benign `nes/close` noise is omitted from live stderr. - Preferred `.cmd` ancestor binaries on Windows and added `EPERM` copy fallbacks for Codex auth seeding and Gemini skill materialization. - Added a real Node ACP echo-agent spawn smoke that can run directly on any supported platform. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapter-utils/src/acpx-engine/spawn-smoke.test.ts` — 57 passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `node --test scripts/acpx-patch-packaging.test.mjs scripts/release-lib.test.mjs` — 10 passed. - Full canary release dry run under Node 24.18.0 / npm 11.16.0 — passed in an isolated scratch clone. - `git diff --check` — passed during implementation verification. - One-time GitHub Actions proof: [Ubuntu ACPX spawn smoke](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774579), [Windows ACPX spawn smoke](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774558), and [Canary Dry Run](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774497) passed on head `f345ac69f2`; the dedicated smoke jobs are intentionally not retained in the recurring PR workflow. ## Risks - The ACPX stderr callback and env persistence exemption are carried as a pnpm dependency patch until ACPX exposes/fixes those behaviors upstream. - Child stderr is synchronously appended to preserve ordering and failure diagnostics; unusually high-volume agent stderr could briefly block the Node event loop. - The Windows-specific `.cmd` resolution and symlink `EPERM` branches are proven by the standalone smoke test and the linked one-time `windows-latest` run rather than a permanent CI gate. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.4 via Codex CLI, medium reasoning, repository/tool execution enabled; context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5ed0b74b34 |
fix(runtime): scope PAPERCLIP_ env-binding strip to reserved keys (#9974)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs get their environment from user/adapter/project/routine env bindings resolved by the server heartbeat, plus `PAPERCLIP_*` runtime vars (identity, wake, workspace, API access) injected by the harness > - The heartbeat stripped **every** `PAPERCLIP_`-prefixed binding before resolution, so legitimately user-named keys (e.g. cloud provider token bindings like `PAPERCLIP_CLOUD_PROD_PROVIDER_RAILWAY_*`) were silently dropped and never reached the run env > - At the same time, several adapters honored an explicitly configured `PAPERCLIP_API_KEY` over the harness-minted run token, which is exactly the one key config must never control > - This pull request replaces the blanket prefix strip with a precise three-rule policy: never accept `PAPERCLIP_API_KEY` from config, always let harness-assigned runtime vars win, and let every other `PAPERCLIP_*`-named user binding flow through > - The benefit is that user secrets with a `PAPERCLIP_`-style name work like any other binding, while runtime identity and API credentials stay fully harness-controlled ## Linked Issues or Issue Description **Bug description** (no public issue exists): - **What happened:** Env bindings whose key starts with `PAPERCLIP_` (e.g. a cloud provider token a user deliberately named `PAPERCLIP_CLOUD_PROD_PROVIDER_RAILWAY_TOKEN`) were silently stripped by the server before secret resolution, so the spawned agent never received them. No error, no access event — the variable just never appeared. - **Expected behavior:** A user-named `PAPERCLIP_*` binding should reach the run env unless the harness itself uses that key. Only `PAPERCLIP_API_KEY` should be categorically rejected, and harness-assigned runtime vars (`PAPERCLIP_RUN_ID`, `PAPERCLIP_AGENT_ID`, wake/workspace vars, …) should always win over config. - **Steps to reproduce:** Configure an agent/project env binding named `PAPERCLIP_<ANYTHING>` (plain or secret_ref), run a heartbeat, and inspect the spawned process env — the key is absent. - **Deployment mode:** local server, any local adapter. Related prior PRs (different, save-time/API-layer blanket-ban approach; this PR supersedes that direction with a runtime allow-except-reserved policy): Refs #8239, Refs #8439. ## What Changed - `server/src/services/heartbeat.ts`: the pre-resolution strip now removes only `PAPERCLIP_API_KEY` (hard denylist) instead of every `PAPERCLIP_`-prefixed binding; other `PAPERCLIP_*` keys flow into binding resolution. Low-trust inline-sensitive-env checks now also cover those keys. - `packages/adapter-utils/src/server-utils.ts`: new `isForbiddenConfigEnvKey()` helper; the shared `refreshPaperclipWorkspaceEnvForExecution` merge drops `PAPERCLIP_API_KEY` from config and keeps harness-assigned `PAPERCLIP_*` keys authoritative. - `packages/adapter-utils/src/acpx-engine/execute.ts`: removed the explicit-`PAPERCLIP_API_KEY`-from-config allowance; the run token (`authToken`) is now always applied; config `PAPERCLIP_API_KEY` is ignored. - All local adapters (`claude-local`, `codex-local`, `cursor-local`, `gemini-local`, `grok-local`, `opencode-local`, `pi-local`) plus `cursor-cloud`, `hermes`, and the server `process` adapter: removed `hasExplicitApiKey`-style allowances so the harness token always wins, and guarded the remaining unguarded env-merge loops (claude-local inline loop, process adapter) with the same policy. - Tests updated/added: heartbeat binding-strip test now asserts the three-rule policy; adapter-utils merge tests assert the `PAPERCLIP_API_KEY` ban and `PAPERCLIP_*` pass-through; acpx engine tests moved credential fixtures to `authToken` and assert config `PAPERCLIP_API_KEY` is ignored while other `PAPERCLIP_*` config keys forward and still bust the session fingerprint on rotation. ## Verification - `pnpm vitest run packages/adapter-utils/src/server-utils.test.ts packages/adapter-utils/src/acpx-engine/execute.test.ts` — 127 passed - `pnpm vitest run server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-local-environment.test.ts server/src/__tests__/claude-local-execute.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/cursor-local-execute.test.ts server/src/__tests__/gemini-local-execute.test.ts` — 68 passed - Adapter package execute suites and the server tests touching API-key fixtures (`heartbeat-run-log`, `redaction`, `effective-run-config-fingerprints`, `agent-permissions-routes`) — green. Three pre-existing sandbox/SSH fixture failures reproduce identically on clean `master` on this host and are unrelated. - `pnpm --filter <pkg> typecheck` for server, adapter-utils, and all nine touched adapter packages — all pass. ## Risks - Behavioral change: a deployment that relied on configuring a static `PAPERCLIP_API_KEY` in adapter config env loses that override — by design; the harness-minted run token is now the only source. When no run token exists, no API key is injected at all. - `PAPERCLIP_*`-named user bindings now reach binding resolution and run envs; a key that collides with a harness runtime var is still discarded at merge time, so runtime identity/wake/workspace vars cannot be spoofed. - Low risk otherwise: no migrations, no API surface changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic Claude 5 family, Mythos-class tier), extended thinking enabled, agentic tool use (file edits, shell, test runner) via Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7ffeafad1f |
refactor(codex-local): relocate Codex auth-merge scripts + decision predicate into the adapter (#9785)
## Thinking Path > - Paperclip is an open source platform for managing AI agent teams, with adapter packages providing concrete runtime environments (e.g. `codex-local` runs a local OpenAI Codex sandbox) > - `adapter-utils` is the shared utilities package — it should hold only generic, adapter-agnostic primitives (sandbox lifecycle helpers, merge logic, type definitions) usable by every adapter > - Three files lived in `adapter-utils/src/` that are entirely Codex-specific: `codex-auth-merge-extract.sh` (shell script that extracts auth tokens), `codex-auth-merge-decision.cjs` (CJS decision helper), and `codex-auth-merge-scripts.ts` (TypeScript factory that wires them into the inbound provision seam added in PR #9778) > - Their presence in a "generic" package violates the adapter isolation principle and requires a cross-package build step to copy `.sh`/`.cjs` files into `codex-local/dist/server/` at build time > - This PR completes Phase 2 of the inbound-seam refactor: move all three files to `packages/adapters/codex-local/src/server/`, rebase the TypeScript imports, update `execute.ts` to import locally, move the Codex-auth tests into a new `codex-local` test file, and fix packaging so `codex-local` copies its own scripts to `dist/server/` > - Script bytes are identical after the move; no change to inbound auth-merge behavior or to which bytes cross the sandbox boundary > - The benefit is a clean ownership boundary: `adapter-utils` retains only generic runtime code, and the structural "Codex-free core" test in the adapter-utils suite validates this invariant going forward ## Linked Issues or Issue Description No pre-existing public GitHub issue. Describing the underlying problem inline: **Problem:** `packages/adapter-utils/src/` contains three files (`codex-auth-merge-extract.sh`, `codex-auth-merge-decision.cjs`, `codex-auth-merge-scripts.ts`) consumed exclusively by the `codex-local` adapter. Their presence in a generic utilities package violates adapter isolation and requires a cross-package build step (copy `.sh`/`.cjs` into `codex-local/dist/server/`). The inbound provision seam landed in PR #9778 routed these through `adapter-utils`; this PR finishes the relocation. **Related:** Refs #9778 (Phase 1 — generic asset-lifecycle-seam, now merged). ## What Changed - Moved `codex-auth-merge-extract.sh`, `codex-auth-merge-decision.cjs`, and `codex-auth-merge-scripts.ts` from `packages/adapter-utils/src/` → `packages/adapters/codex-local/src/server/` (script bytes are unchanged; `codex-auth-merge-scripts.ts` imports rebased to `adapter-utils` subpaths for `shellQuote` and `SandboxManagedRuntimeAssetProvision`) - `packages/adapters/codex-local/src/server/execute.ts`: updated import of `buildCodexAuthInboundProvision` from `adapter-utils` → local `./codex-auth-merge-scripts` - New `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` — 5 Codex-auth test rows extracted from `adapter-utils/src/workspace-restore-merge.test.ts`; the generic test file stays and no longer references Codex - `packages/adapters/codex-local/package.json`: build now copies `.sh`/`.cjs` from `src/server/` into `dist/server/` directly - `packages/adapter-utils/package.json`: removed the now-obsolete cross-package copy step for those files ## Verification ```sh # Type-check both affected packages pnpm --filter @paperclipai/adapter-utils --filter @paperclipai/adapter-codex-local typecheck # Codex auth-merge suite (moved tests — 5/5 pass) pnpm --filter @paperclipai/adapter-codex-local test -- --reporter=verbose codex-auth-merge # Full adapter-utils suite including the structural "core free of Codex literals" test (243 passed, 4 pre-existing skips) pnpm --filter @paperclipai/adapter-utils test # Confirm build copies scripts to dist/server pnpm --filter @paperclipai/adapter-codex-local build ls packages/adapters/codex-local/dist/server/*.sh packages/adapters/codex-local/dist/server/*.cjs ``` All of the above were run locally and passed before this PR was opened. ## Risks Low risk — pure file relocation: - No change to `.sh` or `.cjs` script bytes; the same content reaches the sandbox boundary as before - No behavioral change to inbound auth-merge logic from the sandbox's perspective - `adapter-utils`' structural "core free of Codex literals" test now validates that the relocation is complete and will catch any regression - Only `codex-local` consumed these files from `adapter-utils`; no other package in the monorepo imported them from there ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) — Anthropic, 200k context window, tool use, agentic reasoning. Used to author the refactor. `Co-authored-by: Paperclip <noreply@paperclip.ing>` trailer present on all commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cf5ba4bbea |
feat(adapter-utils): generic per-asset lifecycle-contribution seam (#9778)
## Thinking Path
> - Paperclip's sandbox managed runtime is responsible for provisioning
the agent's execution environment — it extracts a home directory asset
into the sandbox before the adapter runs.
> - The sandbox runtime core was directly branching on the adapter key
(`codex`) to decide which merge scripts to stage and which merge-extract
command to run, coupling generic infrastructure to a specific adapter's
credential-merge protocol.
> - This makes it harder to add, remove, or modify per-adapter asset
provisioning without touching the runtime core; it also prevents other
adapters from contributing staged files or a custom extract command at
all.
> - The fix is to move the adapter-specific knowledge into the adapter
itself: the asset descriptor gains optional `provision` (stageFiles +
extractCommand) and `restore` contribution fields that any adapter can
populate, and the runtime core consumes them generically.
> - This pull request introduces those contribution fields, wires the
Codex adapter's inbound credential-merge as a `provision` contribution,
and removes the adapter-specific branching from the runtime core.
> - The benefit is a clean seam: the runtime core is now
adapter-agnostic for asset provisioning, the inbound behavior is
unchanged (same merge matrix, same scripts), and other adapters can
attach custom staged files or extract commands without modifying shared
infrastructure.
## Linked Issues or Issue Description
No pre-existing public GitHub issue. Describing the problem inline per
the feature template:
**Problem or motivation**
The sandbox managed-runtime asset provisioning in
`sandbox-managed-runtime.ts` branched directly on the adapter key
(`codex`) to decide which merge scripts to stage and which shell command
to use during asset extraction. This tight coupling prevents other
adapters from customizing their provisioning without modifying the
runtime core, and it means the runtime core must import and know about
adapter-specific merge scripts.
**Proposed solution**
Add an optional `provision` contribution (array of `stageFiles` entries
+ an `extractCommand` string) and an optional `restore` contribution to
the asset descriptor returned by adapters. The runtime core now consumes
these generically — if a `provision` contribution is present, it stages
those files and uses the supplied command; otherwise it falls back to
the default `tar -xf` extraction. The Codex adapter populates the
`provision` contribution where it previously depended on core branching.
**Alternatives considered**
Keeping the adapter-specific logic in the core as a documented
exception; rejected because it makes the seam inextensible.
**Roadmap alignment**
Decoupling — removes a latent coupling between the runtime core and a
specific adapter.
## What Changed
- Added `provision` contribution field (`stageFiles: Array<{src, dest}>`
+ `extractCommand: string`) to the `SandboxManagedRuntimeAsset`
descriptor type in `adapter-utils`.
- Added `restore` contribution field (hook for post-restore logic,
populated in a later phase) to the descriptor.
- Removed adapter-key branching (`if adapterKey === 'codex'`) from the
runtime core in `sandbox-managed-runtime.ts`; the core now reads
`provision.stageFiles` and `provision.extractCommand` generically.
- Extracted Codex-specific merge-script paths and the merge-extract
command into `codex-auth-merge-scripts.ts` in `adapter-utils`; the Codex
adapter's `execute.ts` now attaches them as a `provision` contribution
when it builds its managed-home asset descriptor.
- Updated `execution-target.ts` to pass the extended asset type through
to the adapter call site so the new fields are load-bearing end-to-end.
- Added seam-proving unit tests in `sandbox-managed-runtime.test.ts`:
contribution-less asset uses the default path; a non-adapter asset
round-trips the generic provision+restore seam; a structural assertion
verifies the runtime core carries no Codex-specific string literals.
- Added one test in `workspace-restore-merge.test.ts` confirming the
inbound merge matrix is unaffected.
## Verification
```bash
# Unit tests (20 pass):
npx vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/adapter-utils/src/workspace-restore-merge.test.ts
# Type-check both affected packages:
cd packages/adapter-utils && npx tsc --noEmit
cd packages/adapters/codex-local && npx tsc --noEmit
# Structural: runtime core carries no adapter string literals
grep -n 'codex\|auth\.json' packages/adapter-utils/src/sandbox-managed-runtime.ts
# Expected: zero matches
```
## Risks
**Low risk.** This is a behavior-preserving refactor: the inbound
provisioning output (which files get staged, which command runs) is
identical to before, now driven by the adapter-supplied contribution
instead of core branching. The existing inbound merge matrix tests are
the regression guard. No change to which bytes cross the sandbox
boundary. The SSH transport is untouched.
## Model Used
- **Provider:** Anthropic
- **Model:** Claude Sonnet 4.6 (`claude-sonnet-4-6`)
- **Context window:** 200 K tokens
- **Capabilities used:** tool use (file read/edit, bash execution,
Paperclip API), extended reasoning over multi-file TypeScript refactor
- **Mode:** agentic (Paperclip ACPX platform)
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Harold Kim <harold@paperclip.ing>
|
||
|
|
5a5c918705 |
fix(acpx): configure Codex models at startup (#9700)
Move Codex ACPX model, reasoning effort, and fast-mode settings into CODEX_CONFIG startup config so arbitrary model IDs avoid ACP session picker validation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d32ed88443 |
fix(recovery): route recovery by failure cause (#9634)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies and their work. > - Its recovery subsystem detects stranded issue execution and decides whether to retry, escalate, or request operator intervention. > - The existing recovery path used a mostly generic owner ladder and generic execution contract, so transient failures could wake a manager who then performed the deliverable instead of repairing and returning the task. > - Provider quota failures also entered the same takeover path even when the correct action was to wait for capacity and retry the original assignee. > - Recovery actions already retain the source owner and evidence needed to choose a cause-specific route, render a scoped contract, and measure whether work was handed back. > - This pull request adds a cause-keyed recovery playbook, propagates its contract through every built-in adapter, and makes resolved recovery actions return work to the original owner by default. > - The benefit is bounded self-recovery that preserves task ownership, avoids needless management takeover, and makes recovery outcomes observable. ## Linked Issues or Issue Description No matching public GitHub issue was found. Related recovery work was reviewed but is not duplicated here: #9630 restores bounded recovery continuations, #8807 changes one assignee-ranking case, and #9404 records runtime-failure transition evidence. This change instead introduces cause-specific routing and recovery contracts across the recovery lifecycle. ### What happened? When an issue became stranded, recovery generally selected an owner through the same fallback ladder and rendered the normal execution contract. That made the recovery wake look like ordinary deliverable work, even when the correct action was to retry the original agent, repair its runtime, or wait for a provider quota reset. ### Expected behavior Recovery should select a response by failure cause, tell the recipient to recover rather than complete the deliverable, suppress takeover wakes for provider quota waits, and return repaired work to its original assignee unless the recovery owner explicitly completes it. ### Actual behavior Recovery could escalate transient failures to management, omit the cause-specific next action from the wake, and leave the recovery owner assigned after the runtime problem was resolved. ### Impact The generic path creates avoidable management work, ownership churn, and budget consumption while obscuring whether recovery successfully returned work to the responsible agent. ## What Changed - Added cause-keyed routing for process loss, missing disposition, provider quota limits, Codex output inactivity, workspace validation failures, and fallback recovery causes. - Added recovery-scoped wake rendering that replaces the generic execution contract with the failure summary, original assignee, attempt count, next action, and cause-specific playbook instruction. - Propagated the structured recovery contract through all built-in adapter execution paths, including Hermes local and gateway adapters. - Added provider-quota wait monitoring so capacity failures schedule the original assignee instead of enqueueing a takeover wake. - Added hand-back behavior and `handed_back` / `owner_completed` outcome accounting when recovery actions are resolved. - Added focused routing, renderer, quota-monitor, and hand-back regression coverage plus implementation-spec documentation. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/issue-recovery-actions.test.ts` - 4 test files passed; 194 tests passed. - Targeted `pnpm --filter ... typecheck` across `@paperclipai/adapter-utils`, `@paperclipai/shared`, `@paperclipai/server`, `@paperclipai/ui`, and all nine changed adapter packages. - 13 affected workspace packages passed typecheck. - `pnpm check:token-gates` - All UI token gates passed. ## Risks - Recovery routing behavior changes for stranded work, so an incorrectly classified cause could select a different recipient than before; fallback causes retain the existing management ladder. - Provider quota detection depends on structured failure evidence and conservative text matching; unmatched failures continue through fallback recovery. - Adapter prompt plumbing changes across built-ins, covered by shared renderer tests and compile-time call signatures. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with exact model ID `gpt-5.6-sol`, using reasoning, tool use, and code execution. The runtime does not expose its configured context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ae2c30f2f |
feat(skills): import skills from projects (#9620)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies. > - Company skills make reusable agent behavior discoverable and editable from one place. > - Projects already contain skill directories, but operators had to import each skill path manually. > - Copying those skills would break the desired write-through workflow between Skill Studio and the source project. > - The server therefore needs a safe preview/select/import contract that only accepts rediscovered, workspace-contained candidates. > - The UI needs a guided project picker that explains reference semantics, handles conflicts, and remains usable on mobile. > - This pull request adds that end-to-end project skill import flow with authorization, tenant-scope, traversal, and symlink regression coverage. > - The benefit is faster bulk onboarding while keeping project files as the single source of truth. ## Linked Issues or Issue Description **Feature request** **Problem:** Importing several skills already stored in a Paperclip project requires operators to discover and submit each local path individually. This is slow, hides which well-known directories were searched, and makes conflict/already-imported states difficult to evaluate before mutation. **Proposed solution:** Add an “Import skills from project” flow that previews skills from well-known directories, lets operators selectively import eligible candidates, and stores local-path references so Skill Studio edits write through to the project files. **Alternatives considered:** Copying files into company-managed skill storage was rejected because it creates divergent copies. Trusting client-supplied paths was rejected because imports must be constrained to server-rediscovered, workspace-contained candidates. **Additional context:** GitHub duplicate search found no existing issue or PR for this exact workflow. Refs #3799 for related skill-import inventory behavior; this PR does not claim to close that issue. ## What Changed - Extend `scan-projects` with backward-compatible preview and selective-import modes, typed validation, candidate statuses, and OpenAPI coverage. - Discover project skills under `skills`, `.agents/skills`, `.claude/skills`, `.codex/skills`, `.cursor/skills`, `.opencode/skills`, and `.gemini/skills`. - Re-discover selections server-side, enforce company/project/workspace scope, and reject traversal or symlink escapes before creating `local_path` references. - Add the Skills-page menu entry and responsive project import dialog with project selection, grouped candidates, select all/deselect all, conflicts, empty/error/403 states, and import results. - Add route, service, and component regressions for preview authorization, cross-tenant selections, traversal/symlink safety, selection counts, grouping, and result semantics. ### Screenshots **Choose a project**  **Review discovered skills**  **Mobile selection footer**  **Import result**  ## Verification - `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts ui/src/pages/skills/ImportSkillsFromProjectDialog.test.tsx` — 3 files, 81 tests passed. - `pnpm check:token-gates` — all token gates clean. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - Security review passed after adding tenant-scope and unauthorized-preview regressions; UX re-review approved desktop/mobile surfaces; QA passed all seven acceptance areas including write-through editing, deduplication, conflicts, empty state, and permission denial. ## Risks - Files remain referenced in project workspaces, so moving or deleting a source directory can make an imported skill unavailable; the UI explicitly communicates the reference behavior. - New well-known directory scans may discover more candidates than older versions, but preview mode prevents mutation until the operator confirms a selection. - The endpoint remains backward compatible: omitting `mode` preserves the prior full-import behavior. - No schema migration or telemetry event changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 4.8 with tool use/code execution assisted with the UI implementation and UX polish. OpenAI Codex CLI with tool use/code execution assisted with server implementation, security fixes, regression coverage, integration, and PR preparation; the runtime did not expose Codex's exact backing model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3a727bf780 | fix(codex): warn when sandbox auth is shadowed (#9259) | ||
|
|
0ecae2cd7e |
fix(adapters): forward resolved adapter env to local agents, keep runtime env authoritative (#9617)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local coding adapters (Claude, Codex) run each agent heartbeat in a spawned child process; Paperclip resolves adapter-configured env — including `secret_ref` bindings — into the env for that process > - The resolved adapter env was not being forwarded reliably: config env could overwrite Paperclip's own runtime vars, and a warm/resumable ACP session could keep serving stale env because its fingerprint ignored the resolved env > - This mattered because a configured API key or other secret could be silently absent from the agent shell, and a resumed session would never pick up an updated value — while a config binding could also override runtime identity/wake vars > - This pull request keeps Paperclip-managed `PAPERCLIP_*` runtime env authoritative over config, and folds a stable hash of the applied adapter env into the session fingerprint so an env change forces a fresh launch > - The benefit is that adapter-configured env (plain values and resolved secrets) reliably reaches the agent process, updates are picked up on the next launch, and runtime identity can never be clobbered by config ## Linked Issues or Issue Description No public GitHub issue exists, so the underlying bug is described inline following the bug-report template. **What happened** Env keys configured on a local adapter (plain values and `secret_ref` bindings, resolved server-side into plain strings) did not reliably reach the spawned agent process. Two distinct gaps: (1) when merging config env into the process env, a config key in the reserved `PAPERCLIP_*` namespace could overwrite a Paperclip-managed runtime variable (identity, wake, workspace, API access); (2) a warm-handle / resumable ACP session computed its reuse fingerprint from `secretManifestHash` only, which misses plain-value edits and same-version secret rotations — so a resumed session kept serving stale env and never re-launched with updated values. **Expected behavior** Non-`PAPERCLIP_*` adapter env (plain + resolved secret values) is forwarded to the agent process; a change to any applied forwarded value invalidates a warm/resumable session so the next launch sources the latest env; configured `PAPERCLIP_*` entries can never override Paperclip runtime env, while an explicitly configured `PAPERCLIP_API_KEY` (stable per-run config) is still honored and its rotation also busts the session. **Steps to reproduce** Configure an adapter with an env key (e.g. a `secret_ref` API key) and a resumable ACP session. On resume, the updated env value is not sourced; separately, a `PAPERCLIP_*` config key overrides the runtime value. **Deployment mode** Local adapters (Claude / Codex) via the shared adapter-utils execution path. ## What Changed - `packages/adapter-utils/src/server-utils.ts`: add `isPaperclipRuntimeEnvKey` and, in `refreshPaperclipWorkspaceEnvForExecution` (used by all local adapters), skip a `PAPERCLIP_*` config key when Paperclip has already assigned it this run; all other keys still forward. - `packages/adapter-utils/src/acpx-engine/execute.ts`: apply the same `PAPERCLIP_*` non-override rule (via the shared helper) when merging config env, capture the applied config env in `resolvedAdapterEnv`, and fold a stable `adapterEnvHash` of it into the session fingerprint so an env change forces a fresh launch. Per-wake `PAPERCLIP_*` runtime vars are assigned earlier and never enter that map, so they stay out of the hash; stable configured `PAPERCLIP_*` values (e.g. an explicit `PAPERCLIP_API_KEY`) are included so rotating one busts the session. - Added unit/integration tests for plain + secret forwarding, `PAPERCLIP_*` non-override, the explicit-API-key path, fingerprint refresh-on-env-change vs. stable-across-wakes, and rotation of a configured `PAPERCLIP_API_KEY`. ## Verification - `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapter-utils/src/server-utils.test.ts` → 2 files, 111 tests passing (includes the new cases). - Tests assert: forwarded plain/secret values appear in the spawned wrapper `.env`; a `PAPERCLIP_*` config key does not override the runtime value; changing an applied forwarded env value (including a rotated `PAPERCLIP_API_KEY`) changes `configFingerprint`, while a new wake with the same config env keeps it stable. ## Risks Low risk. Behavior change is limited to (a) config env no longer overriding `PAPERCLIP_*` runtime vars — a security-positive tightening — and (b) a resumable session re-launching when its applied config env changes, which is the intended fix. Per-wake `PAPERCLIP_*` churn is deliberately excluded from the fingerprint so normal sessions still resume across heartbeats. Existing sessions get a new fingerprint once on first deploy (the added `adapterEnvHash` field), which is expected. No secret values are logged (key-name redaction in the adapter plus manifest-driven redaction of `meta.env`). ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended reasoning with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7947308276 |
fix(codex): classify refresh auth failures (#9598)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip supports the Codex local adapter, which runs OpenAI Codex CLI sessions on behalf of agents > - Codex uses OAuth refresh tokens to maintain long-running authenticated sessions > - When a refresh fails, the failure has distinct root causes: a refresh token was already reused in a parallel request, the token expired by TTL, or the token was invalidated/revoked by the provider > - Without classifying these failure modes, all refresh auth errors surface identically — operators cannot distinguish retryable transient collisions from permanent invalidations, and run logs carry no actionable diagnosis > - This pull request adds structured classification (`refresh_token_reused`, `refresh_token_expired`, `refresh_token_invalidated`) of Codex refresh-token auth failures across the CLI quota-probe, ACP auth path, and execute path > - The benefit is that these distinct failure modes can be surfaced in run logs and acted on appropriately — transient reuse can be retried; true invalidations require re-auth ## Linked Issues or Issue Description <!-- Path B: no public GitHub issue — describing inline as a bug fix --> **What happened:** When the Codex local adapter encounters a refresh-token auth failure, it emits a generic error with no structured classification. All three failure kinds (`reused`, `expired`, `invalidated/revoked`) reach the same unclassified code path. **Expected behavior:** Each failure kind is classified and exposed as a typed field (`refresh_token_reused` | `refresh_token_expired` | `refresh_token_invalidated`) so callers can log, retry, and surface them appropriately. **Steps to reproduce:** 1. Run a Codex agent session with a reused or expired OAuth refresh token. 2. Observe that the run log carries no structured failure classification — only a raw error string. **Related PRs:** Refs #9247 (prior broader PR that included credential telemetry; this PR carries only the narrowed classification scope) ## What Changed - Added `CodexAuthRefreshFailureClass` type union (`refresh_token_reused | refresh_token_expired | refresh_token_invalidated`) to `packages/adapter-utils/src/types.ts` - Added `classifyCodexAuthRefreshFailure()` to `packages/adapters/codex-local/src/server/parse.ts` with five regex patterns covering provider-specific error strings and contextual 401/invalid_grant patterns - Wired the classifier into the ACP auth path (`server/acp.ts`), execute path (`server/execute.ts`), and CLI quota-probe (`cli/quota-probe.ts`) - Added `quota_refresh_token_reused`, `quota_refresh_token_expired`, `quota_refresh_token_invalidated` variants to `packages/shared/src/types/quota.ts` - Added classification unit tests (`parse.test.ts`, `quota-spawn-error.test.ts`, `acp.test.ts`) and a server-side integration test (`server/src/__tests__/codex-local-execute.test.ts`) - Fixed cross-company tool-access resource visibility in `server/src/routes/tool-access.ts` - Stabilized `heartbeat-retry-scheduling.test.ts` (CASCADE cleanup), `heartbeat-run-log.test.ts`, and `quota-windows.test.ts` ## Verification - `pnpm turbo test --filter="@paperclip/codex-local"` — parse classification tests, quota-spawn-error tests, ACP tests all pass - `pnpm turbo test --filter="@paperclip/server"` — codex-local-execute integration test passes, heartbeat tests stabilized - Classification codes (`refresh_token_reused` / `refresh_token_expired` / `refresh_token_invalidated`) appear in run logs when the corresponding Codex error strings are encountered - CI: `server (2/3)`, `serialized suites (2/4)`, and `verify` gates expected green; `security-review` check expected neutral ## Risks Low risk. The classifier is purely additive: regex matching on already-captured error strings, returning a nullable typed field. Callers that do not inspect the classification field are unaffected. No execution paths, retry logic, or existing error surfaces changed. ## Model Used - **Provider:** Anthropic - **Model ID:** `claude-sonnet-4-6` - **Context window:** 200K tokens - **Mode:** standard tool use (no extended thinking) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b79f744a8d |
Fix Codex auth merge host-unusable fail closed (#9276)
## Thinking Path > - Paperclip is the open source platform people use to manage AI agents for work > - The Codex adapter runs agent tasks in isolated sandbox environments on the user's machine > - When a Codex sandbox is reused across agent runs, its home directory (including `~/.codex/auth.json`) is restored from a prior snapshot > - Both the host machine and the sandbox independently maintain `auth.json` credentials; on sandbox reuse, these can diverge > - The previous merge code had fail-open edge cases: if host auth was in an unusable state, if the auth JSON object shapes differed between host and sandbox, or if the subscription account identities didn't match, the merge would proceed silently with whatever data was available > - This PR adds fail-closed behavior: if host Codex auth is unusable, if auth parser shapes differ, or if subscription account identities don't match, the merge fails explicitly rather than silently continuing with stale or incorrect credentials > - The benefit is that Codex agents on reused sandboxes now fail fast and loudly when auth is in a broken state, instead of silently running with wrong credentials and producing confusing downstream failures ## Linked Issues or Issue Description No pre-existing public GitHub issue. This is a targeted security hardening fix for the Codex reused-sandbox auth merge path. **Problem:** When a Codex sandbox is reused, the merge logic that reconciles host and sandbox `auth.json` credentials failed open in several cases: - Host `auth.json` present but in an unusable state (missing required keys, empty token material, malformed JSON) → merge would proceed with whatever the sandbox had - Host and sandbox auth payloads had different shapes (e.g., one uses `OPENAI_API_KEY`, the other uses a `tokens` object) → parser-differential case not detected - Subscription account identities (`tokens.account_id`) differed between host and sandbox → stale sandbox identity would be used silently **Fix:** All three cases now fail closed. The merge returns an explicit error rather than proceeding with potentially stale or mismatched credentials. Related PRs: - Refs #9262 — sandbox Codex auth shadow warning (adjacent auth area) - Refs #9259 — auth precedence exports (adjacent auth area) ## What Changed - `packages/adapters/codex-local/src/server/codex-home.ts` — New file with `hasUsableAuthPayload()`, `codexHomeHasUsableAuth()`, and full Codex home setup/teardown. Includes fail-closed auth merge guards: rejects unusable host auth, detects parser shape differentials, and checks subscription account identity match before merging - `packages/adapter-utils/src/workspace-restore-merge.ts` — New file with directory snapshot diffing and restore-merge logic; the merge operation fails closed when auth validation fails - `packages/adapters/codex-local/src/server/codex-home.test.ts` — Unit tests covering auth usability checks, symlink management, and fail-closed merge paths - `packages/adapter-utils/src/workspace-restore-merge.test.ts` — Unit tests for snapshot/restore-merge behavior including fail-closed cases - `packages/adapter-utils/src/sandbox-managed-runtime.ts` — Updated to invoke the fail-closed auth merge during sandbox restore ## Verification Tests run and passing: ```sh corepack pnpm exec vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts packages/adapters/codex-local/src/server/codex-home.test.ts corepack pnpm exec vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts corepack pnpm --filter @paperclipai/adapter-utils typecheck corepack pnpm --filter @paperclipai/adapter-codex-local typecheck git diff --check origin/master HEAD ``` All passed locally before push. ## Risks - **Intentional behavioral change (breaking for previously-silent failures):** Reused sandboxes that previously completed auth merge with unusable host auth, parser-differential auth shapes, or mismatched account identities will now fail with an explicit error. This is the correct behavior — the prior silent-proceed path was the bug. Users affected will see a clear error message rather than a confusing downstream auth failure. - **Auth.json symlink migration:** `ensureSymlink()` detects stale copied `auth.json` files (written by older Paperclip versions) and replaces them with symlinks on first run. This is safe: the target is always under the Paperclip-managed company home, never the user's real `~/.codex`. Directories at the symlink path are left untouched (EISDIR is not silently swallowed). - **Low risk for non-reuse paths:** The fail-closed logic only activates during sandbox restore/reuse. Fresh sandbox allocations are unaffected. ## Model Used - **Provider:** Anthropic - **Model ID:** claude-sonnet-4-6 (Claude Sonnet 4.6) - **Context window:** 200k tokens - **Mode:** Agentic coding with tool use; extended thinking not used - **Role:** Code author (Priya Raman, BackendEngineer) with Harold Kim (Git Expert) handling push and PR operations ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Priya Raman <priya.raman@paperclip.local> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Harold Kim <harold@paperclip.ing> |
||
|
|
931eec3fbf |
feat(mcp) [split 4/8]: wire gateway runtime and Smoke Lab (#9559)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 4/8 and focuses on gateway runtime, Smoke Lab, plugins, and server wiring > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The policy core needs runtime execution, endpoint guards, route registration, heartbeat integration, and adapter MCP injection to become operational. - Proposed solution: Adds the remaining server routes/wiring/consumers, runtime tests, adapter-utils MCP contracts, and Claude/Codex injection implementations required by the server layer. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/03-server-tool-access`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: SecurityEngineer for gateway, endpoint guard, token issuance, and runtime wiring; Greptile on every PR. ## What Changed - Adds the remaining server routes/wiring/consumers, runtime tests, adapter-utils MCP contracts, and Claude/Codex injection implementations required by the server layer. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - Changed server test set — 26 files, 382 tests passed - Affected server adapter tests — 38 tests passed after concrete adapter boundary move - Adapter-utils and Codex focused tests — 76 tests passed ## Risks - Remote endpoint validation, token handling, and runtime supervision are security-sensitive and can fail closed or deny legitimate access if misconfigured. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |