mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
dfcda67650d4fc60b1cca537efca7fc9d52c718c
123
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bd7a13eb9c |
fix(daytona): replace bwrap user-namespace with su privilege drop (#10805)
## Thinking Path
> - Paperclip runs AI agents inside Daytona sandbox environments
> - The Daytona plugin wraps agent commands in bwrap to prevent
accidental writes outside the workspace
> - bwrap previously used `--unshare-user --uid <uid> --gid <gid>` to
run commands as the sandbox user inside the container
> - This creates a uid_map of `<uid> 0 1` — inside-uid maps to
outside-uid 0 (root)
> - Files owned by the sandbox user (outside-uid 1001) appear as
overflow uid 65534 (nobody) from inside the namespace
> - So all writes to the workspace fail with Permission denied, and the
agent cannot run
> - This PR replaces the user-namespace approach with `su -s /bin/sh
<user>`, which gives the correct uid mapping
> - The benefit is that bwrap works correctly on Daytona: writes succeed
and the advisory isolation is preserved
## Linked Issues or Issue Description
No existing issue. Describing the bug inline per the bug report
template.
**What happened?**
Running a Codex agent in a Daytona sandbox failed immediately. The agent
could not create directories inside the workspace:
```
mkdir /home/daytona/paperclip-workspace/.paperclip-runtime/codex/paperclip-bridge/queue ... failed with exit code 1:
bwrap: Can't find source path /home/daytona/paperclip-workspace: Permission denied
```
Root cause: bwrap used `--unshare-user --uid 1001 --gid 1001` (via
sudo/root). This writes uid_map `1001 0 1` — inside-uid 1001 maps to
outside-uid 0. Files owned by outside-uid 1001 (the workspace) appear as
overflow uid 65534 (nobody) from inside the namespace. `--bind-try`
suppresses ENOENT but not EACCES, so the bind exits 0 and the wrapper
proceeds — but every write inside then fails with Permission denied.
The bwrap capability probe (`sudo -n bwrap --unshare-user --uid 0 --gid
0 --ro-bind / / -- true`) did not test the workspace bind or the su
invocation, so it incorrectly reported bwrap as available.
**Expected behavior**
The agent should start and run normally inside the Daytona sandbox.
Directory creation and file writes in the workspace should succeed.
**Steps to reproduce**
1. Configure Paperclip with a Daytona sandbox provider
2. Start a Codex agent task targeting a Daytona sandbox
3. Observe the `adapter_failed` error: `mkdir ... failed with exit code
1: bwrap: Can't find source path /home/daytona/paperclip-workspace:
Permission denied`
**Paperclip version or commit**
master (reproducible on current HEAD before this fix)
**Deployment mode**
Self-hosted server
**Agent adapter(s) involved**
- [x] Codex
**Relevant logs or output**
```
mkdir /home/daytona/paperclip-workspace/.paperclip-runtime/codex/paperclip-bridge/queue \
/home/daytona/paperclip-workspace/.paperclip-runtime/codex/paperclip-bridge/queue/requests \
/home/daytona/paperclip-workspace/.paperclip-runtime/codex/paperclip-bridge/queue/responses \
/home/daytona/paperclip-workspace/.paperclip-runtime/codex/paperclip-bridge/queue/logs \
failed with exit code 1: bwrap: Can't find source path /home/daytona/paperclip-workspace: Permission denied
(adapter_failed)
```
## What Changed
- Replaced `--unshare-user --uid <uid> --gid <gid>` with `su -s /bin/sh
<username>` in `buildBwrapCommand`. bwrap runs as real root (for
bind-mount capability), then `su` drops into the sandbox user. Inside
uid=1001 maps to outside uid=1001, so workspace files are writable.
- Removed `detectSandboxUidGid` (ran `id -u` + `id -g`). Added
`detectSandboxUsername` (runs `id -un`) — the username is what `su`
needs.
- Updated `detectBwrapAvailable` probe to test the actual invocation:
`sudo -n bwrap --ro-bind / / [--bind-try <workspace> <workspace>] -- su
-s /bin/sh '<user>' -c true`. This catches both the su failure mode and
the EACCES-on-workspace-bind case.
- Updated `detectBwrapCapability` to run sequentially (username first,
then probe with that username and remoteCwd).
- Replaced `sandboxUid`/`sandboxGid` in lease metadata with
`sandboxUsername`.
- All three `detectBwrapCapability` call sites now pass `remoteCwd`.
- Updated `BwrapExecPlan` type and `resolveBwrapExecPlan` to use
`username: string` instead of `identity: { uid, gid }`.
- Updated tests TDD-style: rewrote tests to describe the new behavior
first, then implemented to pass them.
## Verification
**SSH verification** (run against a live Daytona sandbox before writing
the fix):
```bash
# Old approach — fails
sudo -n bwrap --unshare-user --uid 1001 --gid 1001 \
--ro-bind / / --bind-try /home/daytona/paperclip-workspace /home/daytona/paperclip-workspace \
-- sh -c "mkdir -p /home/daytona/paperclip-workspace/test"
# → mkdir: cannot create directory: Permission denied
# New approach — works
sudo -n bwrap --ro-bind / / \
--bind-try /home/daytona/paperclip-workspace /home/daytona/paperclip-workspace \
-- su -s /bin/sh daytona -c "mkdir -p /home/daytona/paperclip-workspace/test && echo ok"
# → ok (files owned by uid 1001)
```
**Unit tests:**
```
cd packages/plugins/sandbox-providers/daytona && npx vitest run
# 114 passed, 2 pre-existing failures (macOS tar compatibility in file-sync.ts, unrelated)
```
## Risks
- `su` must be available in the Daytona sandbox image. It is a standard
POSIX tool present in every image tested. If absent,
`detectBwrapAvailable` returns false and execution falls back to the
unwrapped path (same behavior as before).
- The advisory bwrap wrapper was never a security boundary — it is
best-effort. The behavioral change (root → sandbox user inside the
container) is strictly better: files created by the agent now have the
correct ownership.
- `sandboxUid` and `sandboxGid` are removed from lease metadata. Any
external code reading those fields will get `undefined`. They were only
used internally by `resolveBwrapExecPlan`, which now reads
`sandboxUsername`.
## Model Used
Claude Sonnet 4.6 (`claude-sonnet-4-6`) via Claude Code CLI — tool use
enabled, extended context. The model diagnosed the bug via SSH
inspection, designed the fix, and implemented it TDD-style (tests first,
then implementation).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
||
|
|
af6b32d82a |
feat(observability): add provider OTel spans - cache-hit flag, plugin tracer, Daytona pack/transfer spans (#10764)
## Thinking Path > - Paperclip uses pull requests to ship control-plane changes with review and audit. > - This change extends sandbox provider telemetry. > - The first PR added startup and exec spans. > - This PR adds provider spans for cache hits, plugin tracer flow, and Daytona file sync steps. > - The host must keep the trust boundary and reject unsafe span data. > - The result is fuller OTel coverage with bounded labels and safe parentage. ## Linked Issues or Issue Description **Problem or motivation** Sandbox provider telemetry lacks an explicit cache-hit signal, plugin span context, and file-sync pack and transfer spans. The current proxy for a cache hit is fragile. The plugin SDK also needs a safe tracer surface and a per-call parent context. **Proposed solution** Add an explicit `cacheHit` flag from the Daytona sandbox handle lookup. Expose `ctx.tracer` on the plugin context and pass a `traceparent` string with each call. Record provider spans on the host with allowlist clamping and capability checks. Wrap Daytona file sync pack and transfer steps in spans. **Alternatives considered** Keep the old `providerGetMs == 0` proxy. Reject that path because the cache-hit decision belongs at the handle lookup, not in a timing proxy. The host stays the trust boundary for worker span data. It clamps labels, rejects bad parent data, and gates the worker-to-host span RPC by capability. A security review completed before this PR opened. ## What Changed - Added an explicit `cacheHit` flag from the Daytona sandbox handle lookup. - Added `ctx.tracer` on the plugin context and a `traceparent` field in the per-call context. - Added host-side span recording with attribute clamping and capability checks. - Wrapped Daytona file sync pack and transfer steps in spans. - Added tests for the host trust boundary, the plugin tracer no-op path, and the Daytona span paths. ## Verification - `pnpm --filter @paperclipai/plugin-daytona exec vitest run` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run` - `pnpm --filter @paperclipai/server exec vitest run plugin-host-services environment-execution-target instrumentation plugin-worker-manager` - `tsc --noEmit` for server, plugin-sdk, and adapter-utils ## Risks - Span data now crosses the worker boundary, so the host allowlist and capability gate must stay strict. - The `traceparent` path must stay valid and host-owned. - Daytona file sync spans must not add new round trips. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes:` / `Closes:` / `Refs:` or described the issue in the PR body - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53bcf3897f |
feat: sync @-mentioned projects into remote sandboxes (confined sandbox transport, flag ON) (#10564)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The run layer must move project context into the sandbox that executes the agent > - Local sandboxes already stage referenced projects for @-mentions > - Remote confined sandboxes dropped the whole referenced set, so the agent lost needed files and paths > - This pull request keeps the confined sandbox transport aligned with the local behavior for referenced projects > - It does this behind a remote-only flag that defaults on, while SSH keeps the old drop-only path > - The benefit is that remote runs can read the same referenced project context that local runs already provide ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change touches server orchestration, sandbox transport, and observability. **Problem or motivation** A run can @-mention another project. Local targets stage each referenced project and give the agent a path. Remote confined sandboxes dropped the full referenced set, so the agent could not read those project files or paths. **Proposed solution** Enable referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`, which defaults on. Keep the SSH transport out of scope and keep it dropping referenced projects. Repoint each referenced workspace hint at its staged `project-<projectId>` sandbox directory. Publish `PAPERCLIP_WORKSPACES_JSON` on the confined sandbox lane. Count each per-project remote staging failure as a `staging` failure in the requested-vs-synced metrics. **Alternatives considered** Keep the remote path drop-only. That keeps the gap open. Move the change into SSH too. That expands scope beyond the target transport and adds risk. **Roadmap alignment** No matching item in `ROADMAP.md` showed up in this review. **Additional context** The change lands in three commits. The first commit opens the gate for the confined sandbox transport. The second commit repoints the workspace hints and publishes the workspace map. The third commit records per-project staging failure data. ## What Changed - Opened remote referenced-project sync for the confined sandbox transport behind `PAPERCLIP_MULTI_PROJECT_WORKSPACE_SYNC_REMOTE`. - Repointed referenced workspace hints to the staged `project-<projectId>` sandbox directories and published `PAPERCLIP_WORKSPACES_JSON`. - Counted per-project remote staging failures as first-class `staging` failures in the requested-vs-synced observability. ## Verification - The pushed ref `refs/heads/feat/sync-referenced-projects-remote-sandbox` resolves to the authorized submit SHA. - `git log --oneline origin/master..origin/feat/sync-referenced-projects-remote-sandbox` shows exactly the three expected commits. - The handoff reports server typecheck clean, adapter-utils typecheck clean, and the listed unit tests passing. - The handoff also reports no open review comments and no Greptile score yet. ## Risks - The change touches authorization and sandbox path handling, so regressions could block remote runs or expose the wrong project context. - The new flag defaults on, so any bug in the remote path affects normal remote use. - SSH stays out of scope, so the two transport paths must remain distinct. ## Model Used OpenAI GPT-5 via Codex. Tool use enabled. Context window not reported in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c4f62644b0 |
docs(daytona): document operator enablement for the advisory bwrap wrapper (#10560)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Daytona sandbox provider now uses an advisory bwrap wrapper > - Operators need clear host and image setup for bubblewrap, sudo, and user namespaces > - The repo should document that setup, but it should not own provisioning > - This pull request adds the operator guidance to the shared sandbox requirements and the Daytona README > - The benefit is that operators can enable the wrapper with the same steps the code expects ## Linked Issues or Issue Description Refs #10554 and #10541. ## What Changed - Added an advisory bwrap prerequisites section to `packages/plugins/sandbox-providers/SANDBOX-REQUIREMENTS.md`. - Added an operator enablement section to `packages/plugins/sandbox-providers/daytona/README.md`. - Documented the install commands, the sudoers rule, the user namespace setting, and the verification command. - Kept provisioning out of the repo and left it to the image or snapshot layer. ## Verification - `git diff --check origin/master...HEAD` - `gh pr checks 10560` ## Risks - Low risk. This change updates documentation only. - The docs can drift if the host setup changes later. - Provisioning still lives outside the repo. ## Model Used - OpenAI Codex, GPT-5, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8c910b9a40 |
feat(daytona): activate the advisory bwrap wrapper at the execute seam (#10554)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - Daytona runs code in a remote sandbox > - The execute seam decides how agent commands run > - The pure bwrap builder and the capability probes already merged in #10541 > - This pull request activates the advisory bwrap wrapper at the execute seam > - The wrapper uses bwrap when the lease has the needed data and keeps plain execution when wrap is not safe > - The benefit is safer isolation without changing the current runtime contract ## Linked Issues or Issue Description - No public GitHub issue exists for this work. - Related PR: #10541 Problem: The Daytona execute seam can run commands without the advisory bwrap wrapper when the lease does not yet provide the wrapper data. Proposed solution: Activate the advisory bwrap wrapper when the lease reports bwrap support and a known uid or gid. Keep plain execution when wrap is not safe. Alternatives: - Always wrap every command. This can break current behavior and can create root-owned files when the uid or gid is missing. - Reject execution when bwrap is missing. This can fail a lease for a best-effort wrapper and would change the current runtime contract. Roadmap alignment: This work fits the Cloud / Sandbox agents roadmap area. ## What Changed - `executeOneShot` now accepts a bwrap execution plan. - `onEnvironmentExecute` now reads the lease bwrap flags and the writable sync paths. - A helper resolves when the seam should wrap the command. - The writable set now includes the workspace path and the collected read-write sync destinations. - The wrapper re-binds stdin after the fresh `/tmp` mount. - The README explains the advisory wrapper behavior and the writable set model. ## Verification - `pnpm --filter @paperclipai/plugin-daytona exec vitest run plugin.test.ts` - `pnpm --filter @paperclipai/plugin-daytona exec tsc --noEmit` ## Risks - A missing bwrap tool, a missing `sudo -n` rule, or a blocked user namespace keeps plain execution. - The change shifts the execute seam, so command setup needs careful review. - The wrapper is advisory, so it does not add a security boundary by itself. ## Model Used - OpenAI Codex, GPT-5, tool use, code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
70816c18e5 |
feat(daytona): add advisory bwrap command builder and capability probes (#10541)
## Thinking Path > - Paperclip helps people manage AI agents for work > - The Daytona sandbox provider needs a clear advisory wrapper path and best-effort capability checks > - The wrapper must not change the security model or block a lease when the host lacks bubblewrap support > - The lease metadata must carry the capability result so later steps can make a stable choice > - This pull request adds a pure command builder for the advisory wrapper > - This pull request adds non-throwing probes for bubblewrap and sandbox uid or gid data > - The benefit is a safer advisory path with no behavior change in the execution seam ## Linked Issues or Issue Description ### Problem or motivation The Daytona sandbox provider needs a clear advisory wrapper path and best-effort capability checks. The provider must not fail a lease when the host lacks bubblewrap support. The wrapper must stay advisory only. It must not change the security model. ### Proposed solution Add a pure command builder for the advisory wrapper and non-throwing probes for bubblewrap and sandbox uid or gid data. Store the probe result on the lease metadata so later steps can make a stable choice. Keep the execution seam unchanged. ### Alternatives considered Do nothing and keep the current execution seam unchanged. That path gives no signal when a file change is not durable. This pull request adds the signal without changing runtime behavior. ### Roadmap alignment This work fits the Daytona sandbox provider path and keeps the advisory wrapper outside the execution seam. It does not change the current security model. ### Additional context The wrapper is advisory only. It adds no security. The read-only root is a feedback signal. ## What Changed - Added `buildBwrapCommand` as a pure string builder for the advisory wrapper command. - Added `detectBwrapAvailable` and `detectSandboxUidGid` as best-effort probes that never throw. - Stored `bwrapAvailable`, `sandboxUid`, and `sandboxGid` on the lease metadata in the acquire, resume, and probe hooks. - Added a README section that describes the advisory wrapper model. - Kept the execution seam unchanged. ## Verification - `vitest run` for `packages/plugins/sandbox-providers/daytona/src/plugin.test.ts` - The run passed all 106 tests, including the new builder and probe coverage. - The standalone `tsc` run showed only the known baseline noise that already exists on `master`. ## Risks - Risk is low because the execution seam does not change. - The new wrapper stays advisory and does not alter the sandbox security model. - The probe results only add metadata and do not fail the lease on missing host support. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5cffd5c72e |
feat(sandbox): declare and collect the advisory read-write intent (#10521)
## Thinking Path > - Paperclip is the control plane for AI agent companies > - Sandbox providers move files between the host and the sandbox > - The runtime needs advisory data for which sync paths are writable > - The host already knows that intent when it prepares the sync map > - Daytona can record that intent for later use > - This pull request adds the advisory access field and collects writable paths > - The benefit is that later runtime work can use the data without changing current flow ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description below follows the feature-request template fields. ### Problem or motivation An agent runs inside an ephemeral sandbox. Some sandbox paths keep agent changes; other paths do not. Today the runtime has no declared signal for which sync destinations the agent may change and keep. A later feedback wrapper needs this signal to give the agent real-time feedback when a write lands on a non-persistent path. ### Proposed solution Add an optional advisory `access: "rw" | "ro"` field to the sync file-mapping types. The host sets `rw` for the workspace, git-history, and asset destinations, and `ro` for referenced-project trees. An absent value defaults to `ro`. The Daytona provider collects the `rw` target directories into a per-lease writable set for later use. This change adds the metadata and the collection only. No execution path reads the writable set yet, so runtime behavior does not change. ### Alternatives considered Derive a static writable set in provider code. This alternative is weaker: the layer that authors each sync destination already knows the intent, so a per-operation declaration is more accurate and does not hard-code a path list. ### Roadmap alignment This is the first step toward an advisory sandbox feedback wrapper. The wrapper is best-effort and adds no security. The ephemeral sandbox stays the only boundary. ### Additional context The field is advisory metadata. It does not change the file transfer and adds no security. ## What Changed - Added an optional `access` field to the sync file mapping types. - Set the workspace, git history, and asset destinations to writable. - Set referenced project destinations to read only. - Added a writable set store in the Daytona provider. - Recorded the parent directory of each writable sync mapping during sync in. ## Verification - The handoff reports these checks before push. - `pnpm --filter @paperclipai/adapter-utils exec vitest run sandbox-managed-runtime.test.ts` - `pnpm --filter @paperclipai/plugin-sdk exec tsc --noEmit` - `vitest run src/plugin.test.ts` in the Daytona package - `tsc --noEmit` in the Daytona package ## Risks - Low risk. - The new field is advisory. - The writable set store is best effort and in memory. - A cold store falls back to the workspace baseline. ## Model Used OpenAI Codex, GPT-5, tool use, code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in the PR - [x] I have not referenced internal or instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [ ] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d5b9f6c8c9 |
perf(sandbox): coalesce git-workspace stage-sync into one confined syncIn (#10488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cloud and sandbox agents need less round-trip overhead at workspace start > - The workspace start path now uploads the git history and the working tree overlay in two separate sync operations > - That adds extra mkdir, guard, upload, and rename work for the same workspace bytes > - This pull request merges those two uploads into one guarded sync operation > - The benefit is one sync round trip for the workspace bytes with the same guard and login-shell contract ## Linked Issues or Issue Description No public GitHub issue exists for this change. ### Problem A git-backed workspace start uploads the git-history clone and the working-tree overlay as two separate host-to-sandbox sync operations. ### Proposed Solution Merge the two uploads into one sync operation. Keep both tar mappings under the same confinement guard. Run the git-history extract first, then the overlay extract, then optional cleanup. ### Alternatives Considered Keep two sync operations. That keeps the current round-trip cost and duplicates the guard, upload, and rename work. ### Roadmap Alignment This change supports the roadmap work on cloud and sandbox agents by reducing workspace start overhead. ## What Changed - Merge the git-history tar and the working-tree overlay tar into one sync operation. - Run the git-history extract first, then the overlay extract, then optional remove-deleted-paths cleanup. - Keep both temporary tar targets under the runtime directory so the first extract cannot delete the second tar before use. - Keep the symlink-escape guard, login-shell contract, and file-mapping checks on the merged file set. ## Verification - `tsc --noEmit` on `@paperclipai/adapter-utils` - `@paperclipai/adapter-utils` full project test suite: 354 pass / 4 skipped - Orchestrator proof: one merged operation with both tars and two ordered extract commands - Orchestrator proof: the confinement guard covers both tar mappings - Daytona plugin `plugin.test.ts`: 94 pass ## Risks - Any regression in the extract order could change workspace start behavior. - Any regression in the merged guard could block valid uploads or miss an escape attempt. ## Model Used OpenAI GPT-5, tool use enabled, current Codex session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
674d71548a |
perf(sandbox): fold dead sandbox start round trips (bridge dirs + handle-cache seed) (#10485)
## Thinking Path > - Paperclip is the control plane for AI agents. > - Sandbox startup uses bridge directories and a Daytona workspace handle. > - The cold path made repeat directory creation calls and one avoidable handle fetch. > - Those calls add delay but do not change state. > - This pull request folds the bridge directory setup into one exec, removes redundant process-session setup, and seeds the Daytona handle cache at acquire. > - The benefit is fewer deterministic host-to-sandbox round trips and faster cold starts. ## Linked Issues or Issue Description - Problem: Cold sandbox start does extra directory creation work and re-fetches a handle it already has. - Expected result: The startup path should create each directory once and reuse the fresh handle. - Related PRs I found on GitHub: #9280, #9293. ## What Changed - Added `makeDirs` to the bridge queue client and used one `mkdir -p` exec for the callback bridge directories. - Removed the two upfront `mkdir` execs for the process-session bridge stdin and events directories. - Seeded the Daytona sandbox handle cache at acquire so realize can reuse the fresh handle. - Reset the process-scoped cache in the compatibility test so the second sync run sees the expected exec count. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run` - 351 passed, 4 skipped. - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run` - 91 passed. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` - clean. - I checked `ROADMAP.md` for sandbox round-trip work. I found no duplicate planned core work for this change. ## Risks - Low risk. The change removes redundant calls and adds cache seeding. - A wrong cache scope would hide the handle. The seed now checks the lease scope and fails loudly. - The daytona package `tsc --noEmit` still depends on SDK types that are not installed in this isolated workspace. CI covers that path. ## Model Used - OpenAI GPT-5, tool-using, with code execution in the current workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7083c275c8 |
refactor(sandbox): retire the dead noProfile flag from the exec path (#10461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The sandbox exec path starts agent commands and passes runtime options to the server and plugin layers > - This path kept a noProfile flag after the exec wrappers stopped sourcing a login profile > - The flag no longer changed behavior, so it left dead API surface in the protocol and runtime helpers > - This pull request removes that dead flag from the plugin protocol, the server drivers, and the managed-runtime helpers > - It also updates the tests and points the agent runtime README at the sandbox requirements file > - The benefit is a smaller and clearer exec-path contract with no behavior change ## Linked Issues or Issue Description - No public GitHub issue exists. ### What happened? The sandbox exec path kept a `noProfile` field after the exec wrappers stopped sourcing a login profile. ### Expected behavior The plugin protocol, server drivers, and managed-runtime helpers should not expose or forward a dead field. ### Steps to reproduce 1. Run a managed-runtime command through the sandbox exec path. 2. Inspect the protocol payload and runtime helper inputs. 3. Observe that `noProfile` is present even though it no longer changes behavior. ### Paperclip version or commit `60c7da86fc7a6c1dbf37bbcd86e25ecaaff01607` ### Deployment mode Built from source (pnpm dev / pnpm build) ### Additional context This pull request removes the dead field, updates the affected tests, and updates the README note for the sandbox profile path. ## What Changed - Removed noProfile from the plugin protocol and the server exec-path call sites. - Updated the managed-runtime helpers to use the narrower exec-path contract. - Updated the affected tests and added the README pointer to SANDBOX-REQUIREMENTS.md. ## Verification - `git grep -n "noProfile" -- packages/ server/` returns zero matches. - `tsc --noEmit` passed for `@paperclipai/adapter-utils`, `@paperclipai/plugin-sdk`, and `@paperclipai/server`. - `command-managed-runtime.test.ts` passed: 22/22. - `environment-runtime.test.ts` passed: 24/24. ## Risks - Low risk. The flag was already a no-op. - A hidden external caller may still send the removed field. ## Model Used - OpenAI GPT-5, tool-enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d6e235cbcf |
perf(sandbox-providers): drop nvm sourcing from exec wrappers (#10443)
## Thinking Path > - Paperclip keeps agent work on a controlled execution plane. > - Sandbox exec wrappers run on the hot path for agent commands. > - The current change removes the explicit `nvm.sh` load step from those wrappers. > - The sandbox image already restores PATH through profile startup. > - This pull request keeps profile sourcing where the wrapper still needs it and drops only the `nvm.sh` load step. > - The result is a smaller command path with the same node and agent CLI resolution. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Problem: The sandbox exec wrappers spent extra time sourcing `nvm.sh` before each command. The sandbox image already restores PATH in `/etc/profile.d/00-restore-env.sh`, so that explicit `nvm.sh` work was redundant. Proposed solution: Remove the `nvm.sh` source step from all six wrappers. Keep the profile sourcing that the provider still needs for PATH setup. Alternatives considered: Keep the existing shell setup and accept the launch cost. That keeps the current behavior, but it leaves the hot path slower than needed. Roadmap alignment: This change keeps the sandbox command path small and predictable. It does not change the adapter contract or the node resolution rules. ## What Changed - Removed `nvm.sh` sourcing from all six sandbox exec wrappers. - Kept profile sourcing where the provider still needs it for PATH setup. - Switched Modal to a non-login shell because the script now sources profiles itself. - Updated wrapper tests to assert that built commands do not source `nvm.sh`. ## Verification - Local TypeScript typecheck passed in each changed package. - Focused provider tests passed for Daytona, E2B, Modal, exe-dev, Cloudflare bridge, and adapter-utils. - One Daytona test failure is pre-existing and unrelated to this change. ## Risks - This change alters shell startup for sandbox exec paths. - A provider that depends on implicit shell setup may need a follow-up. - The current tests cover command shape, but they do not cover every runtime shell path. ## Model Used OpenAI Codex, GPT-5, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f9034ab3ca |
build(deps-dev): bump @types/node from 22.19.21 to 22.20.1 (#10304)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 22.19.21 to 22.20.1. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
7797995038 |
perf(plugin-daytona): opt-in no-profile fast path for default-PATH execs (#10352)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - The Daytona adapter turns tasks into shell commands and manages execution overhead > - Many short-lived exec calls still pay for login-shell profile sourcing even when the binary already resolves on the sandbox default PATH > - That extra startup work adds latency on the hot path for repeated command execution > - This pull request adds an opt-in fast path that skips profile sourcing only when the caller explicitly requests it and the command does not need shell initialization > - The benefit is lower per-call latency for eligible commands without changing the conservative default behavior for commands that need the profile ## Linked Issues or Issue Description This change does not reference a public GitHub issue. It follows the same Daytona startup-speed work as merged PR #10335 and narrows the execution path for eligible commands while keeping the default login-shell behavior intact. ## What Changed - Added an optional `noProfile` flag to `PluginEnvironmentExecuteParams`. - Refactored Daytona login-shell script assembly so the profile and nvm sourcing block is omitted only on the explicit fast path. - Preserved environment prefixing, `cd`, shell quoting, `NONINTERACTIVE_GIT_ENV`, stdin handling, and `durationMs` behavior on both paths. - Added regression tests for the fast path omission, the preserved execution parameters, and the default profile-sourcing path. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-daytona exec vitest run src/plugin.test.ts` - `pnpm --filter @paperclipai/plugin-sdk tsc --noEmit` - Reverted the guard locally to confirm the two behavior tests fail again, then restored the change. ## Risks - If a caller opts into `noProfile` for a command that depends on shell initialization, the command can fail to resolve its binary. - The API comment and opt-in design keep that risk narrow; the default path remains unchanged. ## Model Used OpenAI GPT-5 (Codex tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ccba45e4d |
feat(sandbox-providers): native providers honor postUploadCommands (Daytona executes; Kubernetes executes) (#10347)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies, so runtime handoffs have to preserve the exact behavior an agent asked for. > - The sandbox-provider layer is where uploaded files and follow-up commands become a real in-sandbox operation. > - The provider-delegable sync-in seam already exists from the prior PR; without this follow-up, native providers can still accept command-bearing uploads and silently drop the commands. > - That is a fail-open gap for native sandbox providers, because the file transfer succeeds while the intended post-upload work never runs. > - This pull request teaches Daytona and Kubernetes to execute `postUploadCommands` in order, inside the sandbox, after the files land. > - The benefit is consistent and safer sync semantics: native providers either run the commands as requested or fail fast instead of pretending the operation completed fully. ## Linked Issues or Issue Description This PR builds on the previously merged provider-delegable sync-in seam and closes the remaining gap for native providers that still dropped `postUploadCommands`. Problem: - A sync operation could include ordered `postUploadCommands`, but a native provider could finish the file upload and skip the commands entirely. - That creates fail-open behavior for command-bearing uploads, especially when the caller relies on the provider to execute the follow-up action in the sandbox. Proposed fix: - Execute `postUploadCommands` inside the sandbox after file placement. - Preserve the provided command order. - Fail fast on the first non-zero exit or timeout. - Keep command execution verbatim and confine any provided `cwd` under the workspace root. Related public PR: - Refs: #10340 ## What Changed - Daytona `performSyncIn` now executes ordered `postUploadCommands` through the existing `executeCommand` seam. - Kubernetes `performSyncIn` now executes ordered `postUploadCommands` through its streaming pod exec path. - Added workspace confinement for provided `cwd` values and defaulted missing `cwd` to the remote root. - Added tests covering the new post-upload command execution behavior in both provider packages. ## Verification - Latest validation recorded on the handoff branch: `@paperclipai/plugin-sdk` and `@paperclipai/plugin-kubernetes` typechecks passed. - Daytona Vitest: `63/63` passing in `plugin.test.ts`. - Kubernetes Vitest: `195/195` passing, including `file-sync.test.ts`. - `git log --oneline origin/master..HEAD` showed a single expected commit on the branch. ## Risks - Command execution semantics are stricter now, so malformed commands or a bad `cwd` will fail the sync instead of being ignored. - The change makes provider behavior more explicit, which can surface previously hidden failures in callers that assumed commands were optional. - Timeout behavior may differ slightly between providers, so the failure mode is intentionally fail-fast. ## Model Used OpenAI Codex, GPT-5-based coding agent with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3d23c3b2c3 |
perf(plugin-daytona): cache the started sandbox handle per lease (#10335)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Daytona sandbox adapter spends time on repeated per-exec sandbox lookups during a lease > - That repeated lookup is pure overhead once the started sandbox handle is already known and trusted for the lease > - The adapter still needs strict isolation and fail-closed behavior because the handle is an authenticated compute object, not an inert value > - This pull request memoizes the started sandbox handle per lease in a process-scoped cache, and now advances freshness only after successful reuse so failed commands cannot suppress stale-handle refreshes > - The benefit is lower provider-get latency on the hot path while keeping resume safety, teardown safety, and observability intact ## Linked Issues or Issue Description This is a Daytona performance fix, not a standalone public GitHub issue. ### Problem / Motivation - Repeated `client.get(sandboxId)` calls on the sandbox hot path re-fetch a handle that is already started and trusted for the current lease. - The extra provider round-trip is pure overhead on repeated exec, sync, resume, and interactive-cancel flows. - Cache freshness also has to be tied to successful reuse, or a failed command can make a stale sandbox look freshly used. ### Proposed Solution - Cache the started `Sandbox` handle in process memory, keyed by a non-secret composite lease scope. - Enforce strict identity checks and eviction on release, destroy, interactive cancel, and resume-sentinel mismatch. - Block teardown cleanup until active lease operations finish so delete/stop cannot race in-flight execute or sync work. - Advance freshness only after successful execute, sync, or resume reuse, so failed operations do not mask an auto-stopped sandbox. - Preserve `getDurationMs` in exec metadata so provider-get latency remains observable. ### Alternatives Considered - Keep fetching the sandbox on every exec path. - Cache only by bare lease id. ### Roadmap Alignment - This change is part of the Daytona performance work and narrows per-call overhead without changing the public adapter contract. ## What Changed - Memoized the started Daytona `Sandbox` handle in a process-scoped cache keyed by a composite lease scope. - Added fail-closed identity checks so cache hits and single-flight populate paths reject mismatched sandbox ids. - Evicted cached handles on release, destroy, interactive cancel, and resume sentinel mismatch. - Blocked release, destroy, and interactive cancel teardown cleanup until active lease operations drain. - Advanced cache freshness only after successful execute, sync, and resume reuse. - Kept `getDurationMs` in exec metadata so provider-get latency remains observable. - Expanded the Daytona plugin tests to cover same-lease reuse, cross-scope isolation, eviction paths, rejected populate handling, concurrent single-flight behavior, cached-resume sentinel revalidation, teardown cancellation safety, and failed-execute freshness handling. ## Verification - `pnpm exec vitest run --config vitest.config.ts` in `packages/plugins/sandbox-providers/daytona` — 81/81 passing, including the teardown-cancel, snapshot-capture, syncIn-cancel, and failed-execute freshness regressions. - `git rev-parse origin/perf/daytona-sandbox-handle-cache` matched the authorized submit SHA `528158f998fa88bb0748300f156b38d86d6589cd` before the fixup commits. - `git log --oneline origin/master..origin/perf/daytona-sandbox-handle-cache` shows the expected focused Daytona changes. - GitHub duplicate/related PR search and ROADMAP review were completed before opening the PR. - Remote CI and Greptile completed successfully after this description was updated; the PR is now ready for board handoff. ## Risks - The cache is process-scoped, so correctness depends on the eviction paths staying complete. - A bug in the scope key or identity checks could leak reuse across the wrong lease boundaries, but the implementation fails closed on id mismatches. - Teardown now waits for active operations to drain, so any missed activity bookkeeping could delay cleanup instead of racing it. - Freshness updates now happen after success, which is safer, but it means any missed success-path call would trigger an extra refresh rather than silently masking staleness. ## Model Used OpenAI GPT-5 via Codex, tool-using coding agent; exact context window not surfaced in the workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
030dd9d15c |
feat(adapter-utils): provider-delegable syncIn seam with ordered post-upload commands (#10340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter/runtime layer has to move files into sandboxes safely and efficiently > - The current sync-in path needs a provider-delegable seam so providers can use their native upload transport when available > - The upload contract also needs ordered post-upload commands so extracted content can be finalized fail-fast after transfer > - The fallback path still has to preserve current behavior when the provider does not expose native sync verbs > - This pull request adds the contract and runtime seam for provider-delegable sync-in, plus the single-stream collapse flag > - The benefit is fewer round trips, a cleaner provider-owned upload path, and a compatible fallback for existing runners ## Linked Issues or Issue Description No public GitHub issue exists for this change. The underlying feature request is described below in the repository's feature-request format. ### Problem or motivation Paperclip needs a sync-in path that lets each provider choose the best available upload transport instead of forcing the harness to orchestrate uploads the same way every time. The runtime also needs a way to describe ordered post-upload commands so providers can finalize extracted content fail-fast after transfer. ### Proposed solution Extend the sync contract with ordered post-upload commands, forward that contract through the plugin and environment runtime layers, and make client syncIn always available. When a provider advertises native sync verbs, the client should delegate to that transport; otherwise it should fall back to the existing tarball/write/extract behavior and then run the post-upload commands in order. ### Alternatives considered Keeping upload orchestration entirely host-side would avoid a contract change, but it would block provider-specific transport optimizations and keep the harness responsible for a path the provider can do more efficiently. A separate post-upload API would add another surface without improving the existing sync flow. ### Roadmap alignment This work aligns with the broader runtime and adapter roadmap because it improves provider integration without changing the external product model. It is an additive contract change that preserves backward compatibility for providers that do not expose native sync verbs. ### Additional context The fallback path still needs to preserve existing observable behavior, including command ordering, cwd confinement, and fail-fast execution. The single-stream progress flag is part of the same transport improvement so smaller writes can collapse to a single round trip when the runner supports it. ## What Changed - Added ordered `postUploadCommands` support to the sync operation contract and SDK mirror. - Plumbed the sync-in contract through the plugin and environment runtime layers. - Implemented a runtime client `syncIn` path that delegates to native provider transport when available, otherwise uses the generic tarball/write/extract fallback. - Preserved fail-fast execution of ordered post-upload commands in the fallback path. - Flipped the sandbox runner's single-stream stdin progress flag to collapse small `writeFile` operations to a single round trip. - Added and updated tests for contract forwarding, fallback behavior, cwd rejection, fail-fast behavior, and single-stream collapse. ## Verification - `pnpm --filter @paperclipai/plugin-sdk exec vitest run protocol.postupload.test.ts` - `pnpm --filter @paperclipai/plugin-sdk exec vitest run environment-sync-negotiation.test.ts` - `pnpm --filter @paperclipai/adapter-utils exec vitest run command-managed-runtime.test.ts` - `pnpm --filter @paperclipai/server exec vitest run environment-execution-target.test.ts` - Local typecheck and targeted suite runs reported in the handoff passed before PR creation. ## Risks - The new fallback path could diverge from the previous inline upload behavior if the tarball/extract contract changes. - Provider-native sync handling may expose provider-specific edge cases if a runner advertises sync verbs but does not fully honor the contract. - The single-stream flag changes transport behavior for small uploads, so regressions would likely show up as round-trip or upload failures. ## Model Used OpenAI Codex (GPT-5, tool-using coding agent). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
197718bc00 |
feat(acpx): per-step round-trip + provider-latency attribution for sandbox startup (#10222)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-backed runs need more precise startup observability so operators can see where time is spent before an adapter is invoked > - Aggregate startup timing hides which boundary is actually slow, especially in remote execution where the bottleneck can move between host round-trips, provider-boundary latency, and handshake phases > - This pull request keeps the existing startup timing channel additive while attributing the latency to the specific startup steps that caused it > - The benefit is better diagnosis of sandbox startup regressions without changing control flow or introducing a schema migration ## Linked Issues or Issue Description Refs: #10204 This PR extends the existing sandbox run-startup timing observability with per-step round-trip and provider-latency attribution for the Daytona startup path. It keeps the event payload additive and free-form, and it leaves the control flow, database schema, and external adapter interfaces unchanged. ## What Changed - Added per-step round-trip counting for the host-to-sandbox execute seam - Added provider-boundary duration accumulation for the Daytona execute and re-fetch steps - Split the ACP handshake timing into `createRuntimeMs` and `ensureSessionMs` while preserving the warm-handle skip - Kept the startup timing payload additive and did not add a schema migration ## Verification - `adapter-utils` acpx-engine and startup-timing suites: pass - Daytona plugin suite: pass with mocked SDK and injected-clock duration assertions - `server` environment-execution-target suite: pass - `tsc --noEmit` for adapter-utils and server: pass - Git validation: fetched `origin/feat/sandbox-start-step-timing-attribution`, confirmed it matches the authorized submit SHA `53c573266e618af05c57a0720aaa9d9e0452de61`, and confirmed the branch contains only the expected commit on top of `origin/master` - Searched GitHub for duplicate or related PRs/issues; found one closely related merged PR and no open duplicate on this branch - Checked `ROADMAP.md`; the broad sandbox-agent roadmap section does not call out this specific startup-timing attribution work as a duplicate ## Risks - Low risk: the change is additive and only enriches existing timing data - Downstream consumers that assume aggregate-only startup timing may need to tolerate the additional per-step fields - The finer spawn/initialize/session split remains a follow-up in the external ACP client because that hook is not available here yet ## Model Used OpenAI GPT-5, tool-using coding agent ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
30ff3d7c58 |
feat(routines): expose activity gate API (#9438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Scheduled routines provide recurring control-plane work without manual intervention > - The new activity gate can suppress scheduled runs when no external work occurred > - The core scheduler and database support landed without a public create/update contract > - Agents, operators, and managed plugins need validated fields plus discoverable semantics to opt in safely > - This pull request exposes the activity gate through routine APIs, revisions, plugin contracts, tests, and skill documentation > - The benefit is backward-compatible control over idle scheduled work without losing activity-triggered follow-up ## Linked Issues or Issue Description - Refs #8534 ## What Changed - Added shared activity-gate policy and scope enums with create/PATCH validation. - Persisted activity-gate fields through routine creation, updates, revision snapshots, pipeline snapshots, and revision restores. - Defaulted legacy revision snapshots during restore and added regression coverage for pre-field snapshots. - Extended managed-plugin routine declarations, production reconciliation, and the SDK test harness to preserve non-default gate settings. - Added end-to-end API coverage for create/PATCH/list/detail round-trips, defaults, and invalid enum rejection. - Documented schedule-only semantics, activity windows, own-run/read-action exclusions, scopes, and an hourly quiet-night watcher example. ## Verification - `pnpm exec vitest run packages/shared/src/validators/routine.test.ts server/src/__tests__/routines-service.test.ts server/src/__tests__/routines-e2e.test.ts` - `pnpm exec vitest run packages/shared/src/validators/plugin.test.ts packages/plugins/sdk/tests/testing-actions.test.ts server/src/__tests__/plugin-managed-routines.test.ts server/src/__tests__/routines-service.test.ts -t 'activity gate|preserves declared activity gate settings|resolves routine agent and project refs'` - `pnpm exec vitest run ui/src/lib/workspace-routines.test.ts ui/src/pages/Routines.test.tsx` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/plugin-sdk typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - GitHub CI: all final-head checks green; Storybook visual regression skipped by path rules. - Greptile: 5/5 with no unresolved review threads. ## Risks - Low risk: defaults remain `always` and `company`, preserving existing routine behavior and old revision snapshots. - Managed plugin manifests can now declare the same validated gate settings as the public routine API; omitted values retain core defaults. - Revision snapshots now include the new fields so policy changes are not lost or treated as no-ops during restore. > For core feature work, checked `ROADMAP.md`: this extends the existing Scheduled Routines roadmap item and does not duplicate a separate planned capability. ## Model Used - OpenAI GPT-5.5 via Codex CLI, with repository tool use and code execution; context-window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7f766526a6 |
feat(sandbox): add task-scoped egress grants (#10155)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Confinement providers protect agent runs with default-deny network policies > - Kubernetes environments currently apply only provider-level, namespace-wide egress allowances > - Tasks that legitimately need GitHub or package registries therefore cannot request narrow access, while network failures do not explain the governing policy or how to request a grant > - This pull request adds issue-scoped egress grants that become workload-owned, run-label-selected policies and carries the effective grant through lease audit metadata > - The benefit is that internet-dependent work can run without enabling broad egress for every concurrent task, and denied requests point operators to the exact grant path ## Linked Issues or Issue Description No public issue exists. Related but distinct: Refs #9944, which adds a provider-wide open-internet posture; this PR keeps provider defaults narrow and adds per-task grants. **Problem / motivation** Kubernetes sandbox egress is configured at the provider/tenant level. A task that needs to clone from GitHub or install from PyPI cannot request those destinations without changing the policy for every run in the tenant namespace. DNS/connectivity failures also surface as generic tool errors with no policy name or remediation path. **Proposed solution** Accept `executionWorkspaceSettings.networkEgress.allowFqdns` and `allowCidrs`, forward the setting through heartbeat environment acquisition, and create a workload-owned NetworkPolicy or CiliumNetworkPolicy selected by `paperclip.io/run-id`. Record the effective grant in lease activity/metadata, expose policy context through `PAPERCLIP_NETWORK_EGRESS_*`, and append the grant path to likely policy-related stderr failures. **Alternatives considered** A provider-wide open-internet switch is broader than required and is already covered by #9944. Mutating the existing namespace policy would leak each task's destinations to other concurrent runs. Standard Kubernetes NetworkPolicy cannot enforce FQDNs exactly, so standard mode uses the existing hardened public-IPv4 TCP 80/443 fallback only for the selected run; Cilium mode remains exact. **Roadmap alignment** This extends the existing cloud/sandbox agent roadmap capability with task-level control-plane policy and does not duplicate a planned roadmap item. ## What Changed - Added validated `networkEgress` grants to issue execution workspace settings and forwarded them through environment lease acquisition. - Added workload-owned, run-label-scoped NetworkPolicy/CiliumNetworkPolicy resources for task FQDN/CIDR grants. - Added lease audit metadata, sandbox policy environment variables, and actionable network-denial stderr guidance. - Added focused parser, manifest, policy creation, and denial-message tests plus Kubernetes provider documentation. ## Verification - `pnpm -C packages/shared exec vitest run src/validators/issue.test.ts` — 27 passed. - `pnpm -C packages/plugins/sandbox-providers/kubernetes test -- --run test/unit/network-policy.test.ts test/unit/cilium-network-policy.test.ts test/unit/scoped-network-egress.test.ts` — 21 passed. - `pnpm -C server exec vitest run src/__tests__/execution-workspace-policy.test.ts` — 15 passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-plugin-environment.test.ts server/src/__tests__/environment-runtime.test.ts` — 26 passed. - `pnpm --dir packages/db build && pnpm --dir packages/shared build && pnpm --dir packages/plugins/sdk build` — passed, including migration safety checks. - `pnpm --dir packages/plugins/sandbox-providers/kubernetes typecheck && pnpm --dir server typecheck` — passed after refreshing the worktree's frozen offline dependencies. - End-to-end cluster validation of the `build-cython-ext` benchmark remains for CI/maintainer Kubernetes infrastructure; the focused tests assert `github.com` and `pypi.org` produce a policy selected only by the granted run. ## Risks - Standard NetworkPolicy cannot express FQDNs, so an FQDN grant allows hardened public IPv4 TCP 80/443 for that run; use Cilium mode for exact hostname enforcement. - The new field is additive and absent by default, so existing runs keep the current provider-level policy. - Workload owner references garbage-collect scoped policies with the Job/Sandbox; a cluster/controller that ignores owner references could temporarily strand a policy that still selects no future run ID. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, tool use and code execution. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
caae2778f0 |
fix: deliver plugin agent session turns and replies (#10137)
## Thinking Path > - Paperclip manages agent execution through heartbeat runs and adapter-specific sessions > - Plugins can open an agent session and send a conversational message through the host service > - The host previously stored that message only in opaque wake payload metadata, so local adapters never saw it in their CLI prompt > - The host also forwarded run log chunks but did not expose the persisted final assistant text as the session reply > - This pull request defines both sides of the session contract in the shared wake renderer and terminal run event > - The benefit is that local adapters receive the actual conversational turn and plugins receive one canonical final reply ## Linked Issues or Issue Description Related context: Refs #629 and Refs #2880 describe adjacent `claude_local` final-text visibility failures. They concern issue comments rather than plugin agent sessions, but exercise the same need for a canonical persisted run summary. Companion consumer change: paperclipai/paperclip-gateway#3. Bug description: - **Observed:** calling the plugin host's `agents.sessions.sendMessage()` with `prompt: "hello"` woke a `claude_local` agent, but the generated CLI prompt omitted `hello`. On completion, the session emitted log chunks and a generic `Run completed` done event, so callers could not reliably recover the assistant reply. - **Expected:** the prompt becomes the user-supplied conversational turn for that agent session, and the successful terminal event carries the run's canonical final user-facing assistant text. - **Reproduction:** create a plugin agent session for a local adapter, call `sendMessage()` with a non-empty prompt, inspect the adapter prompt and terminal session event. - **Affected baseline:** `b517b887a` on `master`, local trusted deployment with plugin host services and `claude_local`; `codex_local` shared the wake-rendering gap because both use the common Paperclip wake prompt renderer. ## What Changed - Added a typed `agentMessage` wake payload rendered by the shared adapter prompt path used by `claude_local`, `codex_local`, and other local adapters. - Labeled session content as user-supplied and explicitly non-authoritative: it cannot expand authorization, permissions, task scope, or company boundaries. - Preserved ordinary heartbeat behavior by omitting the section when no agent-session message exists. - Added canonical `finalText` to terminal heartbeat status events from the already-persisted run summary/result/message. - Defined successful `AgentSessionEvent.message` as the canonical final user-facing reply (or `null`) and forwarded it on the terminal `done` event. - Added host, wake-renderer, normal-heartbeat, and terminal-reply regression coverage. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/heartbeat-agent-session-message.test.ts server/src/__tests__/heartbeat-run-status-payload.test.ts server/src/__tests__/plugin-agent-sessions.test.ts server/src/__tests__/heartbeat-run-summary.test.ts` — 87 passed. - `pnpm -r typecheck` — passed across all 31 workspaces. - `pnpm build` — passed. - `pnpm test:run` — 2,860 passed, 1 skipped, 3 unrelated failures: two existing macOS temp-path alias assertions (`/tmp` vs `/private/tmp`) in workspace branch-containment tests and one reproducible auto-port runtime-service adoption failure. The same three failures reproduce when the two files run alone; none touch this change. - Live Slack verification intentionally remains operator-gated because it requires rebuilding/restarting the host. ## Risks - User-controlled chat text now reaches the model prompt, which is an intentional prompt-injection surface. The renderer labels it as untrusted conversational content, while the existing plugin/session company checks and caller authorization remain unchanged. - `finalText` is added to company-scoped heartbeat status events. It is derived from the same persisted summary/result/message already used for run comments; no raw stdout or secrets are added. - Consumers that ignore the new field remain compatible, and successful runs without usable final text still emit `message: null`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex (GPT-5), agentic reasoning with repository/tool use and code execution; context-window size is not surfaced in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7186dce4b |
fix(plugin-sdk): thread companyId through configChanged + fail-closed cross-tenant guard (LOOA-687) (#10096)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - First-party **plugins** run as isolated workers spawned by the host
`plugin-loader`, reading company-scoped config through a governed
`ctx.config.get(companyId)` channel.
> - The host→worker `configChanged` RPC carries `{ config, companyId }`,
but the SDK dispatch dropped the scope — `onConfigChanged(newConfig)`
was companyId-blind by design — so a **proactive** worker kept a single
worker-global config.
> - #10092 added a startup replay that fans out **every** stored
company's config through `configChanged`. With no deterministic
ordering, a plugin configured for more than one distinct company ends up
running as whichever DB row was delivered last.
> - That is a latent cross-tenant identity/secret confusion bug: one
company's bot token could be applied to another company's traffic.
> - This pull request threads `companyId` through `onConfigChanged` and
adds a fail-closed cross-tenant guard at the SDK layer, so a
single-tenant worker can never silently collapse to a second company's
config.
> - The benefit is that the config-delivery class is fixed at the SDK
boundary — before any genuinely multi-company proactive plugin ships —
without changing today's single-tenant behavior.
## Linked Issues or Issue Description
No public GitHub issue — describing in-PR (hardening / latent security):
**Latent cross-tenant config collapse.** The worker-side `configChanged`
dispatch forwarded only `config` and dropped `companyId`, so a proactive
plugin kept a single worker-global config. #10092's startup replay
delivers every configured company's config sequentially with no `ORDER
BY`, so a plugin with configs for more than one distinct company would
apply a nondeterministic last-write-wins global config (one tenant's
credential applied to another's traffic).
- Builds on and must merge after #10092.
- Not exploitable today: the only proactive consumer (the chat gateway)
has single-tenant config rows, so last-write-wins is a no-op. This is a
hardening pre-condition before any multi-company proactive plugin ships.
## What Changed
- **Thread scope through:** `onConfigChanged(newConfig, context)` with a
new exported `PluginConfigChangeContext { companyId }`. Backward
compatible — the second arg is optional; existing single-arg
implementations are unaffected.
- **Fail-closed cross-tenant guard** (`worker-rpc-host.ts`): a
single-tenant plugin that receives `configChanged` for a second,
distinct company with a *different* config is rejected with the new
`PLUGIN_RPC_ERROR_CODES.CROSS_TENANT_CONFIG` instead of silently
overwriting the applied tenant's config. Idempotent replays of the
*same* config under a different scope row remain allowed.
- **Opt-in `multiCompanyConfig: true`** on the plugin definition for
plugins that genuinely serve multiple companies from one worker (keying
per-company state on `context.companyId`); the guard is bypassed for
those.
- **Deterministic `ORDER BY companyId`** on `registry.listConfigs`, so
the startup replay binds a single-tenant worker to a stable company
across restarts.
- **Loader visibility:** a `CROSS_TENANT_CONFIG` rejection is logged at
`warn` (was best-effort `debug`) so the misconfiguration is surfaced.
- **Regression test**
(`packages/plugins/sdk/tests/worker-rpc-host.test.ts`): two distinct
companies delivered via the startup-replay path fail closed and stay
bound to the first company; an idempotent same-config replay under a
different scope row is allowed; a `multiCompanyConfig` plugin receives
per-company config with the correct `context.companyId`.
## Verification
- SDK `tsc --noEmit`: clean.
- SDK vitest `worker-rpc-host.test.ts`: 7/7 pass (incl. 3 new). The
two-distinct-company case **fails against pre-fix code** and passes
after the fix.
- #10092 embedded-postgres `plugin-config-startup-delivery.test.ts`: 3/3
pass (unaffected by the new `ORDER BY`).
- Full server `tsc --noEmit` against this SDK: clean.
## Risks
- **Low functional risk.** The second `onConfigChanged` arg is optional
and existing implementations are unchanged. Today's single-tenant
gateway keeps working — idempotent same-config replays are explicitly
allowed, so the go-live is preserved.
- **Behavioral shift on misconfig:** a genuinely multi-company plugin
that has NOT opted into `multiCompanyConfig` now fails closed
(`CROSS_TENANT_CONFIG`) rather than silently collapsing to one tenant.
This is the intended safer default; opt in with `multiCompanyConfig:
true` to serve multiple companies from one worker.
- **Not in scope (residual).** Per-company workers/connections for a
genuinely multi-company gateway increase resource use and are tracked
separately (ties into the #10092 fan-out/timeout follow-up). This PR
fixes the class and fails closed; it does not build multi-tenant
connection management.
## Model Used
Claude — Anthropic `claude-opus-4-8` (Opus 4.8), extended thinking, with
tool use / code execution via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
Paperclip ticket id — branch predates this rule; not renaming an open PR
mid-review
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
doc surface; internal SDK/host behavior only
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: anicca <annica@Michaels-Mac-Studio.local>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
5a00040c4c |
feat(plugin-sdk): interactions/approvals respond + attachment read capabilities for the chat gateway (#10066)
Adds the remaining 5 plugin capabilities + 7 worker→host RPC methods (interactions read/respond, approvals read/respond, attachment read) needed by the Slack chat gateway plugin (v0.5.0) to pass manifest capability validation and load. - Security review: PASS (LOOA-642) after the viewer-role privilege-escalation blocker (LOOA-648) was fixed on this branch (requireActiveHumanMember now rejects viewer on impersonation write-paths, matching assertCompanyAccess). - CI: Build, Typecheck, all server suites (3/3 + serialized 4/4), workspaces, e2e shard 2/2, and all security scanners (Snyk/Socket/Superagent/Greptile/security-review) green. - One e2e flake (signoff-policy 'non-participant cannot advance stage') is unrelated: it exercises execution-policy stage advancement (routes/issues.ts, untouched by this PR) and failed on a heartbeat_run_events FK race + 409 checkout conflict. Unblocks LOOA-629 (Slack gateway go-live) and the interview-ask feature. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9edde68373 |
feat(kubernetes): native file-sync lifecycle hooks over pod exec (#10053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - AI agents run in sandboxed execution environments (Kubernetes pods, Daytona workspaces, etc.) and need to sync files between the host and those environments — for workspace setup, asset delivery, and output retrieval > - The existing sync path for Kubernetes uses a base64-over-exec chunk loop: each ~4 MB chunk requires its own `execInPod` round-trip, so large syncs balloon into many exec calls with corresponding overhead > - `execInPod` supports piped stdin/stdout, meaning the full transfer can be done as a single exec that streams a raw `tar` archive over the data channel — one round-trip regardless of file size, with nothing base64-encoded and nothing buffered whole in memory on either side > - PR-1 (#10013, merged) added the `onEnvironmentSyncIn`/`onEnvironmentSyncOut` opt-in hook API to the sandbox provider interface and documented the protocol; PR-2 (#10028, merged) implemented these hooks for the Daytona provider > - This pull request implements the same two lifecycle hooks in the Kubernetes sandbox provider, so workspace/asset file sync streams through one `execInPod` per operation instead of the chunk loop > - The benefit is significantly fewer exec round-trips for large syncs and flat memory use on both host and pod, with security properties preserved: atomic replace, secret-mode enforcement, path confinement, TOCTOU-safe snapshot, and member-confinement on host-assembled archives from sandbox-authored tar output ## Linked Issues or Issue Description This is the third and final PR in a sequential series: - Refs #10013 — PR-1: opt-in sync hook API + provider docs (merged) - Refs #10028 — PR-2: native file-sync lifecycle hooks for Daytona provider (merged) **Feature:** Native single-exec file-sync lifecycle hooks for the Kubernetes sandbox provider. *Motivation:* The existing Kubernetes sync path encodes files as base64 and loops over `execInPod` one chunk at a time (~4 MB per exec). For large workspaces or asset sets this is slow and resource-intensive. The Kubernetes `execInPod` API supports piped stdin/stdout, enabling a raw-`tar` streaming transfer that needs only one exec regardless of file count or size and never buffers the whole payload in memory. *Proposed solution:* Implement `onEnvironmentSyncIn` and `onEnvironmentSyncOut` in the Kubernetes provider using a streaming `execInPod` with a tar pipeline — for syncIn the host builds the archive on disk and streams its raw bytes into the pod's stdin (`head -c <exact-size> | tar -x`, no base64); for syncOut in-pod `tar` writes to the exec's stdout and the host streams those bytes straight to a file. Path confinement, atomic replace, secret-mode enforcement, TOCTOU protection, and a streamed-bytes fail-closed guard are all enforced. ## What Changed - **New `src/file-sync.ts`** in `packages/plugins/sandbox-providers/kubernetes/` — `performSyncIn` and `performSyncOut` over an injected pod-exec closure, keeping transfer logic hermetically unit-testable - **New `execInPodStreaming` in `src/pod-exec.ts`** — a streaming exec primitive that binds a caller-supplied stdin readable and a stdout writable to the exec WebSocket data channel, added alongside the existing `execInPod` (which is unchanged). This lets a transfer stream raw bytes to/from disk instead of buffering the payload as a single string - **Updated `src/plugin.ts`** — registers `onEnvironmentSyncIn`/`onEnvironmentSyncOut`; resolves the `sandbox-cr` pod exactly like `onEnvironmentExecute` and delegates; `job` backend rejects file-sync calls explicitly (out of scope) - **syncIn path:** host builds the tarball to a temp file → streams its raw bytes over exec stdin, bounded in-pod by `head -c <exact-archive-size> | tar -x` (no base64 anywhere) → extract into a `/proc/self/fd`-pinned reserved `0700` staging dir → `chmod`-before-`mv -f` atomic replace per file (directory mappings use `followSymlinks`→`-h`) - **syncOut path:** in-pod validate + realpath-snapshot each source (closes the validation→copy TOCTOU window) → single-exec `tar -c` streamed over exec stdout → host streams that stdout straight to a temp file through a byte-counting transform → member-confined extraction of the sandbox-authored archive - **Security properties:** secret files land at requested mode with no widened window; every interpolated path is shell-quoted and confined lexically plus via in-pod `realpath`; the outbound stream is bounded by a **streamed-bytes disk guard** (`MAX_SYNC_OUTPUT_BYTES`, 8 GiB default, per-call overridable) that fails the transfer closed — writing no target file — if an untrusted pod emits more bytes than allowed. Neither host nor pod buffers the whole payload, so there is no in-memory size cap on the transfer - **No changes** to `execInPod`, `wrapCommandWithEnv`, or `FastUploadInterceptor` (the `environmentExecute` path is untouched) - **No dependency or lockfile changes** - **New tests** in `test/unit/file-sync.test.ts` (atomic-replace, `0600` secret mode, symlink preserve/deref, dir-mapping, exclude, path-confinement rejection, streamed-output guard fail-closed) and `test/unit/pod-exec.test.ts` (streaming stdin/stdout, caller-sink error fail-closed), plus extended `test/unit/plugin.test.ts` ## Follow-up: Legacy Job-Lease Base64 Fallback Fix Addresses the Greptile 4/5 blocking finding ("Handle existing job leases", `server/src/services/environment-runtime.ts`). Job leases provisioned before the `nativeFileSyncUnsupported` metadata flag existed carry `backend: "job"` but no flag, so `supportsSync()` treated them as native-capable and routed their sync to the pod-exec hook — which the job backend rejects (it has no exec channel) instead of using the byte-identical base64 fallback. The fix adds a belt-and-suspenders gate on the persisted `backend === "job"` field alongside the existing `nativeFileSyncUnsupported` flag check, so pre-existing job leases continue syncing via the base64 fallback after deployment. No behaviour change for `sandbox-cr` leases. ## Verification - `pnpm --filter @paperclipai/sandbox-provider-kubernetes test` — 19 files / 182 tests green, including the existing `upload-interceptor` and `pod-exec` suites - `tsc --noEmit` in the kubernetes package — 0 errors - The sync hooks are opt-in; existing `environmentExecute` behaviour is unaffected and tested by the unchanged existing suites ## Risks - **Opt-in only:** `onEnvironmentSyncIn`/`onEnvironmentSyncOut` are registered conditionally; providers that do not register them fall back to the existing chunk loop. No regression risk on the existing path. - **Shell-injection surface:** all path interpolation uses shell-quoting; paths are additionally confined lexically and via in-pod `realpath` before use. - **TOCTOU on syncOut:** the in-pod snapshot validates and records file metadata before the tar call, closing the window between validation and copy. - **Archive member confinement:** host-side reassembly rejects any tar member whose resolved path escapes the target directory, preventing a malicious in-pod tar from writing outside the intended destination. - **Untrusted-output volume:** an over-large outbound stream trips the streamed-bytes disk guard and fails closed (no target written and the temp sink is swept) rather than filling host disk or memory; the guard bounds disk unconditionally and bounds memory insofar as WebSocket write-backpressure holds. ## Model Used Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`) — produced by a Claude-based AI agent using agentic tool use and multi-step code generation. 200K context window, extended reasoning, code execution and verification capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e55d702916 |
feat(plugin-sdk): human-attributed issue comments for chat gateway plugins (#10050)
Adds the `issue.comments.create_human_attributed` capability and `ctx.issues.createComment` `actorUserId` option, with host-side active-human-member verification. LOOA-627. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f215444919 | feat(daytona): native file-sync lifecycle hooks over batch transfer (#10028) | ||
|
|
b247bf7150 |
feat(runtime): opt-in sandbox file-sync lifecycle hooks (API + provider docs) (#10013)
## Thinking Path > - Paperclip is an open-source AI-agent management platform; agents run tasks inside sandboxed environments (Daytona, Kubernetes, E2B, etc.) > - The control-plane ↔ sandbox file-transfer path flows through the `environmentExecute` seam in `protocol.ts` — the only verb available to plugins — which forces a base64-over-exec chunked loop for every file move: workspace files, assets, Codex home sync > - This transport is correct and safe, but it bypasses provider-native bulk/streaming APIs (Daytona `uploadFiles`, K8s `FastUploadInterceptor` / volume mounts), leaving significant throughput on the table for large workspaces > - The right fix is an opt-in seam extension: providers with faster native transfer declare two optional verbs; providers that do not opt in stay on the existing fallback with zero code or behavior change required > - This PR adds the first layer of that extension — two optional verbs (`environmentSyncIn` / `environmentSyncOut`) in the plugin SDK, the runtime plumbing to prefer the native path for the two clean destroy-then-replace cases, and a doc for the contract > - The core correctness invariant is byte-identical fallback: if no provider opts in, execution is exactly what ships today; `assertSyncOperationsConfined` enforces host-side path confinement for providers that do opt in > - No provider advertises the verbs yet → zero production behavior change; future PRs wire up Daytona and K8s providers against this contract ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Description follows the `feature_request` issue template: **Subsystem affected:** packages/plugins — plugin system; packages/adapter-utils — adapter runtime; server/ — EnvironmentRuntimeService **Problem or motivation:** Sandbox file transfers currently always use a base64-over-exec chunked loop regardless of what the underlying provider supports. For workspaces larger than a few MB this becomes the dominant wall-clock cost of every sandbox run, and it bypasses bulk/stream APIs that providers like Daytona already expose natively. **Proposed solution:** Add two optional, opt-in plugin hooks — `onEnvironmentSyncIn` / `onEnvironmentSyncOut` — to the plugin SDK. When a provider defines both hooks and both are advertised via the existing `supportedMethods` negotiation, the runtime prefers the native path for the two clean destroy-then-replace transfer cases; all other cases fall back to the existing byte-identical base64 transport. **Alternatives considered:** An unconditional verb would require every provider to implement or stub the verb. The opt-in / `METHOD_NOT_IMPLEMENTED` pattern (already used by `environmentExecute`) preserves backward compatibility with zero provider changes required. **Roadmap alignment:** Consistent with the ✅ "Cloud / Sandbox agents" and ✅ "Plugin system" milestones; extends the plugin seam rather than adding control-plane-level logic. **Additional context:** Searched open pull requests and issues for duplicate sandbox file-sync / native-transfer work; none found. ## What Changed - **`packages/plugins/sdk`** - `protocol.ts`: two new optional `HostToWorkerMethods` — `environmentSyncIn` / `environmentSyncOut` — plus generic `SyncOperation`, `SyncFileMapping`, and `SyncOutcome` types - `define-plugin.ts`: optional `onEnvironmentSyncIn` / `onEnvironmentSyncOut` fields on `PluginDefinition`; worker advertises each verb only when its hook is defined (else `METHOD_NOT_IMPLEMENTED`, mirroring `environmentExecute`) - `worker-rpc-host.ts`: route new verbs to plugin hooks - `index.ts`: re-export new public types - **`packages/adapter-utils`** - `command-managed-runtime.ts`: expose optional `syncIn` / `syncOut` on `CommandManagedRuntimeRunner` (available only when both verbs are advertised); add `assertSyncOperationsConfined` host-side path-confinement guard - `sandbox-managed-runtime.ts`: `SandboxManagedRuntimeClient` gains optional `syncIn` / `syncOut`; orchestrator prefers native path for default-provision asset inbound and workspace-download-into-fresh-dir outbound; all other paths keep the existing base64 fallback - `sandbox-file-sync.test.ts` (new): 234-line characterization suite — native-opt-in branch, fallback branch, `assertSyncOperationsConfined` escape-path rejection, `followSymlinks` → tar `-h` - `command-managed-runtime.test.ts`: negotiation + native-sync + confinement tests - **`server/src/services/environment-runtime.ts`**: `EnvironmentRuntimeService` delegates to `syncIn` / `syncOut`, gated on advertised support - **`server/src/services/environment-execution-target.ts`**: minor typing fix alongside the new verbs - **`doc/plugins/SANDBOX_FILE_SYNC_HOOKS.md`** (new): documents the full contract — opt-in / no-op guarantee, operation ordering, provider-may-tar, atomicity, `followSymlinks`, secret modes (0600, no window), path confinement, `operationId` opacity, resource bounds, shell-quoting ## Verification ```bash # SDK suite pnpm --filter packages/plugins/sdk test # Adapter-utils suite (includes new sandbox-file-sync characterization tests) pnpm --filter packages/adapter-utils test # Expected: 255 pass / 4 skip # Type-check across affected packages pnpm --filter packages/plugins/sdk typecheck pnpm --filter packages/adapter-utils typecheck # Server changed-file spot check: cd server && npx tsc --noEmit --skipLibCheck 2>&1 | grep -E "environment-(runtime|execution-target)" | head -20 ``` Key behavioral invariant to spot-check: with no provider opting in (the current state), run any sandbox task and confirm file-transfer behavior is byte-for-byte identical to what the pre-PR code produces. The characterization tests assert this at the unit level. ## Risks - **Zero production risk today**: no provider advertises `environmentSyncIn` / `environmentSyncOut`, so the new code paths are unreachable in production; all real traffic stays on the existing base64 fallback - **Path confinement**: `assertSyncOperationsConfined` rejects any `targetPath` that escapes the declared root — this is the primary security boundary for future providers. The test suite covers escape-path rejection - **Atomicity**: the contract delegates atomicity to providers; the doc explicitly calls out that directory-level ops are not guaranteed atomic - **Secret transport**: credential assets (e.g., Codex `auth.json`, directory mappings) continue to use the existing tar path — they do not go through the new verbs in any current provider > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: Anthropic Model: `claude-sonnet-4-6` (Claude Sonnet 4.6) Context window: 200 K tokens Capabilities: extended tool use, multi-file code generation, agentic reasoning via the Paperclip agent framework ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8e856d2518 |
build(deps): bump @codemirror/state from 6.7.0 to 6.7.1 (#9890)
Bumps [@codemirror/state](https://github.com/codemirror/state) from 6.7.0 to 6.7.1. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/state/commits">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
dc7f09be0d |
build(deps-dev): bump vitest from 4.1.8 to 4.1.10 (#9886)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.8 to 4.1.10. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v4.1.10</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li><strong>browser</strong>: Check fs access in builtin commands [backport to v4] - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>OpenCode (claude-opus-4-8)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10680">vitest-dev/vitest#10680</a> <a href="https://github.com/vitest-dev/vitest/commit/5c18dd267"><!-- raw HTML omitted -->(5c18d)<!-- raw HTML omitted --></a></li> <li><strong>vm</strong>: Fix external module resolve error with deps optimizer query for encoded URI [backport to v4] - by <a href="https://github.com/SveLil"><code>@SveLil</code></a> and <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10661">vitest-dev/vitest#10661</a> <a href="https://github.com/vitest-dev/vitest/commit/bae52b511"><!-- raw HTML omitted -->(bae52)<!-- raw HTML omitted --></a></li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v4.1.9...v4.1.10">View changes on GitHub</a></h5> <h2>v4.1.9</h2> <h3>🐞 Bug Fixes</h3> <ul> <li>Fix <code>importOriginal</code> with optimizer and query import [backport to v4] - by <strong>Hiroshi Ogawa</strong>, <strong>David Harris</strong>, <strong>Codex</strong>and <strong>Vladimir</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10546">vitest-dev/vitest#10546</a> <a href="https://github.com/vitest-dev/vitest/commit/a5180190c"><!-- raw HTML omitted -->(a5180)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Wait for orchestrator readiness before resolving browser sessions [backport to v4] - by <strong>Vladimir</strong> and <strong>Séamus O'Connor</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10555">vitest-dev/vitest#10555</a> <a href="https://github.com/vitest-dev/vitest/commit/7fb29651a"><!-- raw HTML omitted -->(7fb29)<!-- raw HTML omitted --></a></li> <li>Wait for iframe tester readiness before preparing [backport to v4] - by <strong>Vladimir</strong> and <strong>Séamus O'Connor</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10497">vitest-dev/vitest#10497</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/10556">vitest-dev/vitest#10556</a> <a href="https://github.com/vitest-dev/vitest/commit/fbc626c40"><!-- raw HTML omitted -->(fbc62)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>mocker</strong>: <ul> <li>Hoist vi.mock() for vite-plus/test imports [backport to v4] - by <strong>Hiroshi Ogawa</strong>, <strong>LongYinan</strong>, <strong>Claude Opus 4.8</strong> and <strong>Vladimir</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10548">vitest-dev/vitest#10548</a> <a href="https://github.com/vitest-dev/vitest/commit/2c9559c02"><!-- raw HTML omitted -->(2c955)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Prevent test run hang on worker crash [backport to v4] - by <strong>Ari Perkkiö</strong> and <strong>Jattioui Ismail</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10543">vitest-dev/vitest#10543</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/10564">vitest-dev/vitest#10564</a> <a href="https://github.com/vitest-dev/vitest/commit/934b0f587"><!-- raw HTML omitted -->(934b0)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5><a href="https://github.com/vitest-dev/vitest/compare/v4.1.8...v4.1.9">View changes on GitHub</a></h5> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/db616d227b6e0cb07a94f5d1bba262ee95db7e46"><code>db616d2</code></a> chore: release v4.1.10 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10718">#10718</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/bae52b5112a6fd8200101b88bf8af9685d077295"><code>bae52b5</code></a> fix(vm): fix external module resolve error with deps optimizer query for enco...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/a7a61e78c7d0718f00173cff6800a91a344457d4"><code>a7a61e7</code></a> chore: release v4.1.9 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10598">#10598</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/934b0f587cb61d8338d83f525295322692a2db40"><code>934b0f5</code></a> fix(pool): prevent test run hang on worker crash (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10543">#10543</a>) [backport to v4] (#...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7fb29651afbae2a9b0cefe6c031a9308f168ac60"><code>7fb2965</code></a> fix(browser): wait for orchestrator readiness before resolving browser sessio...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/a5180190c1be7089e3705e3dd9e84fea118d09d3"><code>a518019</code></a> fix: fix <code>importOriginal</code> with optimizer and query import [backport to v4] (#...</li> <li>See full diff in <a href="https://github.com/vitest-dev/vitest/commits/v4.1.10/packages/vitest">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2683555d13 |
build(deps): bump @codemirror/language from 6.12.3 to 6.12.4 (#9894)
Bumps [@codemirror/language](https://github.com/codemirror/language) from 6.12.3 to 6.12.4. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/language/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1de0a3bb1e |
feat(mcp) [split 2/8]: add governed access contracts (#9557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 2/8 and focuses on database schema and shared governance contracts > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The governed access model needs additive persistence and synchronized shared types before server enforcement can compile. - Proposed solution: Adds migrations 0148–0169, tool-access and Smoke Lab schema, shared types/validators/gallery helpers, and the minimal compile-required contract consumers identified by boundary testing. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/01-demo-servers`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for migrations/validators; Greptile on every PR. ## What Changed - Adds migrations 0148–0169, tool-access and Smoke Lab schema, shared types/validators/gallery helpers, and the minimal compile-required contract consumers identified by boundary testing. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` — passed, including migration numbering and safety checks - `pnpm --filter @paperclipai/db test` — passed - `pnpm --filter @paperclipai/shared test` — passed ## Risks - Migration or contract mistakes could affect every upper layer; all migrations are additive/idempotent and compile consumers are included in this boundary. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cde4e128c |
feat: run ACP sessions in sandbox execution targets (#9390)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters (Claude, Codex, Gemini) default to the ACP engine lane, which needs a live bidirectional stdio session with the agent process > - Sandbox execution targets only exposed one-shot command execution, so every ACP-capable adapter refused remote targets and fell back to the CLI lane with a "supports only the local Paperclip host" warning > - Running agents in sandboxes is a core deployment mode, and losing ACP there means losing streaming updates, structured events, and default-lane parity with local runs > - This pull request adds a provider-agnostic process-session bridge that relays the ACP stdio session into the sandbox over the existing sandbox runner contract, and updates the adapters to use it > - The benefit is that the default ACP lane now behaves the same on the local host and in any sandbox provider, with CLI fallback reserved for targets that genuinely cannot host a bidirectional session ## Linked Issues or Issue Description No existing public issue covers this; inline description following the feature request template: **Problem or motivation** Configuring an ACP-capable adapter (e.g. Claude) with a sandbox environment made every run fall back to the CLI lane with the warning "Claude ACP currently supports only the local Paperclip host, but this run targets a remote environment." The ACP engine only knew how to spawn a local subprocess, while sandbox providers only expose one-shot command execution — so there was no way to hold the bidirectional stdio session ACP requires. **Proposed solution** Add a process-session bridge in `adapter-utils`: a local ACPX-spawnable proxy script connects to a token-authenticated loopback TCP server, which relays JSON-framed stdin/stdout/stderr events to and from a small relay script executed inside the sandbox via the provider's ordinary runner. Claude/Codex/Gemini adapters now treat sandbox targets with a runner as ACP-capable, resolve agent commands against the remote target, and fall back to CLI only when the sandbox exposes no bidirectional path. The sandbox callback bridge injects a run-scoped API endpoint and bridge token so the agent inside the sandbox can reach Paperclip (including work-product handoffs) without ever receiving the host run JWT. **Alternatives considered** A provider-specific lane was prototyped first: Daytona minting SSH access metadata at lease time, converted into an SSH execution target. It was dropped because it only worked for providers able to advertise SSH, added per-provider surface area, and left every other sandbox provider on the CLI fallback. The merged design rides the one-shot runner contract all providers already implement; a regression test pins that sandbox targets stay on the bridge lane even when lease metadata advertises SSH access. **Roadmap alignment** Directly advances the "Cloud / Sandbox agents" roadmap item — agents running in remote and sandboxed environments keep the same control-plane behavior as local ones. No overlap with other planned core work. ## What Changed - `packages/adapter-utils/src/execution-target.ts`: new `startAdapterExecutionTargetProcessSessionBridge()` plus helpers — writes a token-authenticated local proxy script (spawnable by ACPX) and a remote relay script synced into the sandbox, with a loopback TCP server streaming JSON-framed stdio between them; events emitted before the ACP client attaches are buffered so none are lost. - `packages/adapter-utils/src/acpx-engine/execute.ts`: the ACP engine can execute against remote sandbox targets through the bridge instead of requiring a local subprocess, including remote cwd/env shaping. - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: sandbox-scoped API bridging extended to allow work-product handoffs; the sandbox payload env carries a bridge token, never the host run JWT. - `packages/adapters/claude-local`, `codex-local`, `gemini-local` (`src/server/acp.ts`): default-lane selection no longer rejects all remote targets; command resolution is remote-aware (`ensureAdapterExecutionTargetCommandResolvable`, `resolveAdapterExecutionTargetCwd`); the fallback reason is now scoped to sandboxes that expose only one-shot execution. - `server/src/__tests__/environment-execution-target.test.ts`: pins that sandbox targets resolve to the bridge lane, including when lease metadata advertises SSH access. - Non-sandbox remote targets (e.g. SSH) keep the CLI lane: the ACP engine's remote transport is sandbox-only, so default-lane selection falls back for those targets across all three adapters, and tests covering CLI-specific remote behavior pin `engine: "cli"` explicitly. - The bridge authenticates loopback connections before they can own the session or receive buffered output (token required, idle unauthenticated peers dropped), and remote event writes are serialized so the exit event always lands after stdout/stderr have drained. - Daytona plugin: formatting-only residue from the earlier iteration; no functional change. ## Verification - `vitest run` over the touched suites — `packages/adapter-utils/src/acpx-engine/execute.test.ts`, `packages/adapter-utils/src/execution-target-sandbox.test.ts`, `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`, the three adapter `acp.test.ts` files, and `server/src/__tests__/environment-execution-target.test.ts` — 102 tests pass. - End to end: with a Claude agent configured on a Daytona sandbox environment, the primary-model test now selects the default ACP lane (no fallback warning), and the full round trip (wake → sandbox execution → API bridge → comment post) was exercised twice from inside a live sandbox. ## Risks - Behavioral shift: adapters that previously always fell back to CLI on sandbox targets now default to ACP there; `engine=cli` still pins the CLI lane explicitly. - The bridge relays stdio as JSON lines over loopback TCP guarded by a per-session random token; the remote relay runs inside the sandbox under the provider's runner. Providers with slow one-shot execution will see higher session startup latency — the CLI fallback remains for genuinely incapable targets. - No schema or migration changes. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) — extended thinking enabled, agentic tool use via the Claude Agent SDK harness; implementation iterated with local Vitest verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no shipped docs describe the old local-only ACP limitation) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> Co-authored-by: Cody <cody@paperclip.local> |
||
|
|
17dde9d3f2 |
fix(sandbox): keep custom-image snapshots applied to config tests, probes, and saves (#9385)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox environments can capture reusable custom images (provider snapshots) so agents boot with pre-installed tools and CLI logins > - The custom-image runtime fingerprint check included provider secret-ref paths (e.g. the Daytona `apiKey`), while capture-time fingerprinting excluded them, so any config carrying a credential never matched its captured snapshot > - As a result, agent config tests and environment probes silently booted the provider base image instead of the snapshot, test sandboxes were deleted before operators could inspect them, and any environment save orphaned the snapshot without warning > - The UI compounded the confusion by displaying an internal template id that matches nothing in the provider dashboard > - This pull request aligns runtime fingerprints with capture-time exclusions, re-stamps fingerprints on saves that cannot affect the snapshot (warning when they can), archives test/probe sandboxes instead of deleting them, and surfaces the provider snapshot ref in the UI > - The benefit is that custom images actually apply to config tests and probes, survive unrelated config edits, and are debuggable against the provider dashboard ## Linked Issues or Issue Description No public GitHub issue exists for this; describing it in-PR per the bug template. Related: Refs #9329 (saved-environment probe company context — this branch carries an equivalent fix), Refs #8794 (introduced reusable sandbox custom images). **What happened?** With a Daytona environment whose provider config stores the API key as a secret reference and an active captured custom-image snapshot: - Agent config tests and environment probes booted the provider base image (`daytonaio/sandbox:0.8.0`) instead of the captured snapshot, so CLI upgrades/logins baked into the snapshot were missing and the probe reported "login required" and an outdated CLI. - The environment card showed an internal template id (e.g. `b5be03e1-ca5…`) that does not correspond to any snapshot name in the provider dashboard, making the active image impossible to correlate. - Test/probe sandboxes were deleted immediately after the run, so the sandbox a test used could not be inspected afterwards. - Saving the environment config (even fields unrelated to the image) changed the stored fingerprint, silently detaching the snapshot with no warning. **Expected behavior** Config tests and probes boot the captured snapshot when one is active; the UI shows the provider-facing snapshot/template ref; test sandboxes stay inspectable for a short window; unrelated config edits keep the snapshot linked, and edits that genuinely invalidate it produce an explicit warning. **Steps to reproduce** 1. Configure a sandbox environment on Daytona with the API key stored as a company secret reference. 2. Capture a custom image snapshot from the environment page and mark it active (e.g. after installing/logging into a CLI in the setup sandbox). 3. Run the agent config test or an environment probe: the sandbox boots the base image, not the snapshot, and the sandbox is deleted immediately after the test. 4. Save the environment config with an unrelated field change: the snapshot silently stops applying. **Paperclip version or commit** `master` at the merge-base of this branch. **Deployment mode** Self-hosted local instance (macOS, pnpm dev server) with the Daytona sandbox provider plugin. ## What Changed - Runtime custom-image fingerprint checks now exclude provider secret-ref paths, matching capture-time exclusions, so configs carrying credentials match their captured snapshots (`environment-custom-image-runtime.ts`). - Agent config tests and saved-environment probes force fresh, non-reused sandboxes and pass company context so lease-backed probes can resolve company secrets and boot the real snapshot (`environment-probe.ts`, `routes/agents.ts`, `routes/environments.ts`). - Test/probe sandboxes are released by archiving (stop + 60-minute provider-side auto-delete) instead of immediate deletion, so operators can inspect the exact sandbox a test used (Daytona plugin). - On environment PATCH save, changes that cannot affect the captured snapshot re-stamp the template's source fingerprint so the snapshot stays linked; boot-source or provider-identity changes (new manifest field `templateIdentityPaths`) mark the template detached and the save response reports it (`environment-custom-images.ts`, shared plugin types/validators). - The custom-image overview exposes `activeTemplateMatchesConfig`; the environments UI shows the provider snapshot/template ref (internal id moved to a tooltip), warns via toast when a save detaches the snapshot, and shows a persistent "Not in use" warning when the active template no longer matches the saved config (`CompanyEnvironments.tsx`, `api/environments.ts`). ## Verification - `pnpm vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-probe.test.ts server/src/__tests__/environment-routes.test.ts server/src/__tests__/agent-test-environment-routes.test.ts` — server coverage for fingerprint exclusions, re-stamp/detach on save, probe company context, and fresh-sandbox test behavior. - `pnpm vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts` — archive-on-release and snapshot ref handling. - `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx` — snapshot ref display, detach toast, and "Not in use" warning. - Manually verified end-to-end on a live self-hosted instance against real Daytona: config test boots the captured snapshot (CLI login and version persist), the test sandbox remains visible in the provider dashboard as archived, and saving unrelated fields keeps the snapshot applied. ## Risks - Fingerprint exclusion widening: a provider credential rotation alone no longer detaches a captured snapshot; that is the intended behavior (the snapshot content does not depend on the credential), and provider-identity fields (e.g. Daytona `apiUrl`) still detach via `templateIdentityPaths`. - Archived test sandboxes consume provider-side resources for up to their auto-delete window instead of being freed immediately; bounded (60 minutes) and only for test/probe sandboxes. - New optional manifest field `templateIdentityPaths` is backward-compatible; providers that omit it keep current matching behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
953b315dfb |
Shorten skill frontmatter descriptions (#9353)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip agents can load repository and catalog skills, and Codex renders skill names and frontmatter descriptions into startup context. > - Long descriptions consume the fixed skill metadata budget before Codex can use the progressively disclosed skill bodies. > - The repo `.agents/skills` descriptions and a few shipped catalog descriptions had grown into operational documentation instead of short trigger metadata. > - This pull request keeps the strongest trigger language in frontmatter while leaving detailed procedures in each skill body. > - The benefit is lower prompt overhead, more reliable skill triggering, and a regression guard that prevents description drift from returning. ## Linked Issues or Issue Description No public GitHub issue found for this maintenance item. ### Pre-submission checklist - [x] I have searched existing open and closed issues and this is not a duplicate. - [x] I am working against `master`. - [x] I have confirmed the issue originates in Paperclip's shipped skill metadata, not in a local agent adapter or provider. ### What happened? Codex startup renders discovered skill names and frontmatter descriptions into a fixed skill metadata budget. Several repository skill descriptions and one shipped catalog description had grown into long-form operational guidance, which can force Codex to truncate descriptions before the model has enough trigger signal to select the right skill. ### Expected behavior Skill frontmatter descriptions should stay short trigger summaries: one capability sentence plus a “use when” clause. Detailed procedures should stay in the skill body and load only after the skill triggers. ### Steps to reproduce 1. Inspect `.agents/skills/*/SKILL.md` and `packages/skills-catalog/catalog/**/SKILL.md` frontmatter descriptions. 2. Measure folded YAML `description` values. 3. Observe descriptions above the intended short-trigger range, including descriptions above 300 characters. 4. Run the new shipped catalog test to verify future descriptions stay capped. ### Paperclip version or commit Reproduced on `master` at `cc81eefb6047d8eaf57faf785f421c03dc97073c`. ### Deployment mode Local dev / source checkout metadata inspection. This is not database-related. ### Installation method Built from source. ### Agent adapter(s) involved Codex, because Codex startup uses the skill metadata prompt budget. The metadata source itself is core repository/catalog content. ### Database mode Not database-related. ### Access context Not applicable; this is static repository metadata. ### Node.js version `v22.22.2` in the verification environment. ### Operating system Linux container environment. ### Relevant logs or output Final measurement after this PR: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. ### Relevant config None. ### Additional context The shipped catalog manifest was regenerated so the generated package metadata matches the edited catalog `SKILL.md` sources. ### Privacy checklist - [x] I have reviewed all pasted output for PII and redacted where necessary. ## What Changed - Shortened long `.agents/skills/*/SKILL.md` frontmatter descriptions to concise capability plus use-when trigger clauses. - Shortened the over-budget shipped skills catalog descriptions for wireframe, Paperclip capsules, and reflection coach. - Regenerated `packages/skills-catalog/generated/catalog.json` so shipped metadata matches source skill frontmatter. - Added a Vitest regression guard that caps repo skill source descriptions and generated catalog descriptions at 300 characters. ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog test` — 5 files passed, 19 tests passed - `pnpm --filter @paperclipai/skills-catalog typecheck` - Final measurement: 29 source `SKILL.md` files, max description length 215 chars, 5,449 total description chars, estimated 1,363 description tokens at 4 chars/token. Note: the clean PR worktree was created from `origin/master` and contains only this commit, but it does not have `node_modules`; running `pnpm --filter @paperclipai/skills-catalog test` there failed at tool/package resolution (`vitest`, `tsc`, `@paperclipai/shared`). The dependency-equipped workspace passed the commands above before the commit was cherry-picked onto the clean branch. ## Risks Low risk. This changes skill metadata and tests only. The main risk is over-trimming a useful trigger phrase, mitigated by keeping explicit “use when” clauses and leaving detailed guidance in the skill bodies. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent via Paperclip/Codex, with shell and file-edit tool use. Exact API model ID and context window were not exposed in the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3e63a7e3e5 |
Fail fast stalled Daytona Git network commands
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Daytona sandbox provider executes agent commands inside ephemeral Daytona workspaces via `executeCommand` > - Git operations inside these workspaces can block indefinitely when remotes are unreachable or when git prompts for credentials interactively (e.g. via askpass or terminal prompts) > - When git blocks, it consumes the full 900s adapter RPC ceiling, which surfaces as a hard timeout crash on the Paperclip side rather than an actionable error > - This PR adds noninteractive credential defaults and a 120s cap on detected Git network subcommands so stalled operations fail fast with a useful message > - The benefit is that engineers and agents see an actionable error pointing at missing credentials or unreachable remotes instead of an opaque 900s RPC crash ## Linked Issues or Issue Description <!-- No public GitHub issue exists for this internal infrastructure fix. Describing the issue inline. --> **Bug: Daytona sandbox Git network commands stall for up to 900 seconds** **What happened?** `executeCommand` hangs for the full 900 s adapter RPC ceiling when a Daytona workspace git network command (push/fetch/pull/clone) prompts for credentials interactively or the remote is unreachable. The command blocks silently for up to 900 s then crashes with a generic timeout error that names no actionable root cause. **Steps to reproduce** Run any agent handoff that includes a `git push`, `git fetch`, or `git pull` to a remote inside a Daytona workspace where the remote is unreachable or credentials are missing. **Expected behavior** Command fails fast (within ~120 s) with an actionable error naming the unreachable remote or the missing noninteractive credential. **Deployment mode** Daytona sandbox provider (`packages/plugins/sandbox-providers/daytona`). Root cause: Daytona one-shot execution wrappers did not set `GIT_TERMINAL_PROMPT=0`, `GCM_INTERACTIVE=Never`, or disabled askpass helpers, so git blocked waiting for interactive terminal input; no per-operation timeout existed for network-bound git subcommands. ## What Changed - Added `GIT_TERMINAL_PROMPT=0`, `GCM_INTERACTIVE=Never`, `GIT_ASKPASS=echo`, `SSH_ASKPASS=echo`, `SSH_ASKPASS_REQUIRE=force` to all Daytona one-shot execution wrapper invocations so git never blocks waiting for a credential prompt; callers may override via the `env` parameter - Detects Git network subcommands (`push`, `fetch`, `pull`, `ls-remote`, `clone`, `remote update`, `submodule update`) and caps their timeout at 120 s instead of the full 900 s adapter RPC ceiling - Returns an actionable timeout message that names the unreachable remote or the missing noninteractive credential rather than propagating the raw SDK error - Adds two new Vitest tests: one verifying noninteractive credential defaults are injected, one verifying the 120 s network cap and the improved timeout message ## Verification - `corepack pnpm exec vitest run packages/plugins/sandbox-providers/daytona/src/plugin.test.ts --config packages/plugins/sandbox-providers/daytona/vitest.config.ts` — 40 tests pass - `corepack pnpm exec tsc -p tsconfig.json --noEmit` from `packages/plugins/sandbox-providers/daytona` — clean - `corepack pnpm check:no-git-push` — clean > **Known CI note:** The standalone provider package is intentionally excluded from the root workspace. Direct `corepack pnpm test` from the provider directory fails before running tests due to Vitest tsconfig-root resolution in a grafted checkout. The root-config Vitest invocation above is the passing test signal. ## Risks - **Low overall risk.** The new `GIT_TERMINAL_PROMPT=0` / askpass defaults only affect Daytona one-shot execution; they do not touch any shared git config or host environment. - Callers that previously relied on interactive credential prompts inside Daytona (an unlikely pattern for agent workspaces) will now fail fast instead of prompting — this is the intended behavior. - The 120 s network timeout applies only when the command string starts with a recognized git network subcommand, so non-network git operations and all non-git commands are unaffected. - No PII, telemetry schema, crypto, auth flow, or new external endpoint changes. ## Model Used Claude Sonnet 4.6 (`claude-sonnet-4-6`) — extended thinking mode, tool use, code execution. Anthropic Claude running via Paperclip Claude Code adapter. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Harold Kim <harold.kim@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eedc7ddef2 |
Make ACP the default engine for local adapters (#9238)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Adapter packages are the bridge between the control plane and local agent harnesses such as Claude Code, Codex, and Gemini CLI. > - ACP support was concentrated in a separate `acpx_local` adapter, which made ACP feel like a separate agent choice instead of an execution capability of the harness adapters. > - Claude, Codex, and Gemini now have ACP-capable harnesses, so the native adapter should own ACP selection, fallback, config, transcript parsing, and environment diagnostics. > - The standalone ACPX adapter still needs a compatibility path for existing rows, but it should not be offered as an active adapter for new agents. > - This pull request moves the shared ACP runtime into `@paperclipai/acpx-engine`, wires Claude/Codex/Gemini local adapters to prefer ACP when prerequisites are available, and retires `acpx_local` to a tombstone. > - The benefit is one adapter per harness, richer ACP transcripts by default where possible, and a migration path for existing Claude/Codex ACPX agents. ## Linked Issues or Issue Description Closes #5932 — the broken default `acpx_local` Claude path is replaced by native `claude_local` ACP support, existing Claude/Codex ACPX rows migrate to native adapters, and new agents no longer choose the standalone ACPX adapter. Refs #4893 — original merged ACPX local adapter runtime that this PR replaces with native per-harness ACP engines. Refs #6590 — prior ACPX-Claude seamlessness work folded into the new native Claude ACP path. Refs #197 — related open generic ACP/Kiro adapter work; this PR does not close it because Kiro/custom generic ACP remains a separate adapter decision. Refs #7018 — related Kimi-specific `acpx_local` shell failure; this PR retires the built-in standalone adapter but does not add a native Kimi adapter. Refs #8864 — related ACPX prompt/API guidance PR; this PR moves runtime guidance into the shared/native ACP engine path instead of the old standalone adapter. Refs #8881 — related `acpx_local` POSIX shell failure from the old `acpx` pin; this PR updates ACP dependencies but does not claim custom/OMP ACP support as a first-class native adapter. Refs #8964 — related open `acpx_local` stderr cleanup PR; this PR makes the old runtime path obsolete for new agents but keeps it as a non-closing reference. Problem description: - The standalone `acpx_local` adapter duplicates Claude/Codex agent choices that already have first-class local adapters. - ACP should be an execution engine capability of each harness adapter when the underlying harness supports ACP. - Existing `acpx_local` agents should either migrate to native harness adapters or fail with an explicit retirement message instead of silently falling back to the process adapter. ## What Changed - Added `@paperclipai/acpx-engine` as the shared ACP execution, session-codec, CLI formatter, and UI parser package. - Wired `claude_local`, `codex_local`, and `gemini_local` to auto-select ACP by default when prerequisites pass, with `engine=cli` opt-out and `engine=acp` strict mode. - Added ACP config schema/UI fields, environment checks, session-codec preservation, transcript parsing, and adapter capability metadata for the native adapters. - Retired `acpx_local` to a server tombstone, removed its UI/package/runtime image surface, and added a migration for existing Claude/Codex ACPX agents. - Updated package manifests, lockfile, release tooling, docs, Kubernetes sandbox defaults, and tests. ## Verification - `corepack pnpm --filter @paperclipai/acpx-engine typecheck` - `corepack pnpm --filter @paperclipai/adapter-claude-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-codex-local typecheck` - `corepack pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `corepack pnpm --filter @paperclipai/acpx-engine exec vitest run` - `corepack pnpm --filter @paperclipai/adapter-claude-local exec vitest run src/server/acp.test.ts src/server/execute.acp-fallback.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-codex-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts` - `corepack pnpm --filter @paperclipai/adapter-gemini-local exec vitest run src/server/acp.test.ts src/ui/build-config.test.ts src/ui/parse-stdout.test.ts` - `corepack pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && corepack pnpm --filter @paperclipai/server exec tsc --noEmit` - `corepack pnpm --filter @paperclipai/server exec vitest run src/__tests__/adapter-routes.test.ts src/__tests__/adapter-session-codecs.test.ts src/__tests__/adapter-models.test.ts` - `corepack pnpm --filter @paperclipai/ui typecheck` - `corepack pnpm --filter @paperclipai/ui exec vitest run src/adapters/metadata.test.ts src/adapters/adapter-display-registry.test.ts src/components/AgentConfigForm.test.ts src/components/AgentConfigForm.render.test.tsx src/components/transcript/RunTranscriptView.test.tsx` - `node --test scripts/bootstrap-npm-package.test.mjs scripts/release-package-map.test.mjs scripts/verify-release-registry-state.test.mjs` Note: the server typecheck script calls `pnpm` internally; this dev shell exposes pnpm through Corepack only, so I ran the two script steps manually with `corepack pnpm`. ## Risks - Migration changes existing `acpx_local` Claude/Codex agents to native adapter types and clears old ACPX task sessions/runtime state. - Custom ACP commands remain on the retired tombstone and will need a separate future adapter/plugin path. - ACP auto-selection depends on local Node and ACP server command prerequisites; remote and unsupported environments fall back to CLI unless `engine=acp` is explicit. - `@paperclipai/acpx-engine` is a new public package and needs npm trusted-publishing bootstrap before release automation can publish it. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 via Codex coding agent. Exact hosted model build and context-window size are not exposed in this runtime. Tool use included shell execution, repository editing, GitHub CLI operations, and local test/typecheck execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a058f761f3 |
build(deps-dev): bump rollup from 4.61.1 to 4.62.2 (#9070)
[//]: # (dependabot-start) ⚠️ **Dependabot is rebasing this PR** ⚠️ Rebasing might not happen immediately, so don't worry if this takes some time. Note: if you make any changes to this PR yourself, they will take precedence over the rebase. --- [//]: # (dependabot-end) Bumps [rollup](https://github.com/rollup/rollup) from 4.61.1 to 4.62.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/releases">rollup's releases</a>.</em></p> <blockquote> <h2>v4.62.2</h2> <h2>4.62.2</h2> <p><em>2026-06-19</em></p> <h3>Bug Fixes</h3> <ul> <li>Do not add spurious side-effect-free external imports to chunks when using minChunkSize (<a href="https://redirect.github.com/rollup/rollup/issues/6411">#6411</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6411">#6411</a>: Skip side-effect-free external imports when hoisting is disabled (<a href="https://github.com/morgan-coded"><code>@morgan-coded</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6416">#6416</a>: refactor(rust/parser_ast): extract property AstConverter write buffer kind logic to new method (<a href="https://github.com/fabianbernhart"><code>@fabianbernhart</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>v4.62.1</h2> <h2>4.62.1</h2> <p><em>2026-06-19</em></p> <h3>Bug Fixes</h3> <ul> <li>Preserve multipart file extensions when deconflicting output chunks (<a href="https://redirect.github.com/rollup/rollup/issues/6408">#6408</a>)</li> <li>Fix an issue where getLogFilter would match additional logs (<a href="https://redirect.github.com/rollup/rollup/issues/6415">#6415</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6393">#6393</a>: Use import attributes for importing JSON (<a href="https://github.com/selfisekai"><code>@selfisekai</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6408">#6408</a>: fix: insert conflict numbers before first extension in multi-extension filenames (<a href="https://github.com/LeSingh1"><code>@LeSingh1</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6415">#6415</a>: fix: advance value past wildcard prefix before suffix check in getLogFilter (<a href="https://github.com/JSap0914"><code>@JSap0914</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6417">#6417</a>: chore(deps): update msys2/setup-msys2 digest to 66cd2cc (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6418">#6418</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6419">#6419</a>: chore(deps): update dependency eslint-plugin-unicorn to v66 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6420">#6420</a>: chore(deps): lock file maintenance minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>v4.62.0</h2> <h2>4.62.0</h2> <p><em>2026-06-13</em></p> <h3>Features</h3> <ul> <li>Ensure that shared dependencies between manual chunks and entry points receive a serparate chunk (<a href="https://redirect.github.com/rollup/rollup/issues/6374">#6374</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6374">#6374</a>: Extract the static dependencies imported by manual chunks into separate chunks (<a href="https://github.com/TrickyPi"><code>@TrickyPi</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6405">#6405</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6406">#6406</a>: chore(deps): pin dependency concurrently to v9 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6407">#6407</a>: chore(deps): lock file maintenance minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6409">#6409</a>: chore(deps): update minor/patch updates to v6.2.0 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/blob/master/CHANGELOG.md">rollup's changelog</a>.</em></p> <blockquote> <h2>4.62.2</h2> <p><em>2026-06-19</em></p> <h3>Bug Fixes</h3> <ul> <li>Do not add spurious side-effect-free external imports to chunks when using minChunkSize (<a href="https://redirect.github.com/rollup/rollup/issues/6411">#6411</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6411">#6411</a>: Skip side-effect-free external imports when hoisting is disabled (<a href="https://github.com/morgan-coded"><code>@morgan-coded</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6416">#6416</a>: refactor(rust/parser_ast): extract property AstConverter write buffer kind logic to new method (<a href="https://github.com/fabianbernhart"><code>@fabianbernhart</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>4.62.1</h2> <p><em>2026-06-19</em></p> <h3>Bug Fixes</h3> <ul> <li>Preserve multipart file extensions when deconflicting output chunks (<a href="https://redirect.github.com/rollup/rollup/issues/6408">#6408</a>)</li> <li>Fix an issue where getLogFilter would match additional logs (<a href="https://redirect.github.com/rollup/rollup/issues/6415">#6415</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6393">#6393</a>: Use import attributes for importing JSON (<a href="https://github.com/selfisekai"><code>@selfisekai</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6408">#6408</a>: fix: insert conflict numbers before first extension in multi-extension filenames (<a href="https://github.com/LeSingh1"><code>@LeSingh1</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6415">#6415</a>: fix: advance value past wildcard prefix before suffix check in getLogFilter (<a href="https://github.com/JSap0914"><code>@JSap0914</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6417">#6417</a>: chore(deps): update msys2/setup-msys2 digest to 66cd2cc (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6418">#6418</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6419">#6419</a>: chore(deps): update dependency eslint-plugin-unicorn to v66 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6420">#6420</a>: chore(deps): lock file maintenance minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>4.62.0</h2> <p><em>2026-06-13</em></p> <h3>Features</h3> <ul> <li>Ensure that shared dependencies between manual chunks and entry points receive a serparate chunk (<a href="https://redirect.github.com/rollup/rollup/issues/6374">#6374</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6374">#6374</a>: Extract the static dependencies imported by manual chunks into separate chunks (<a href="https://github.com/TrickyPi"><code>@TrickyPi</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6405">#6405</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6406">#6406</a>: chore(deps): pin dependency concurrently to v9 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6407">#6407</a>: chore(deps): lock file maintenance minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6409">#6409</a>: chore(deps): update minor/patch updates to v6.2.0 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6410">#6410</a>: chore(deps): lock file maintenance minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6412">#6412</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6413">#6413</a>: chore(deps): update dependency eslint-plugin-unicorn to v65 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rollup/rollup/commit/8faa18777374582bb813d54ce3623f4acf1f9e0b"><code>8faa187</code></a> 4.62.2</li> <li><a href="https://github.com/rollup/rollup/commit/a38a795c481d589661abf0d5029162ccc4fb79e1"><code>a38a795</code></a> refactor(rust/parser_ast): extract property AstConverter write buffer kind lo...</li> <li><a href="https://github.com/rollup/rollup/commit/6cc5c316bb5c4497923005601c2742132ae117ff"><code>6cc5c31</code></a> Skip side-effect-free external imports when hoisting is disabled (<a href="https://redirect.github.com/rollup/rollup/issues/6411">#6411</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/caacf701b89e5be4a94b3ffdbf70b51e5cfa3a1a"><code>caacf70</code></a> 4.62.1</li> <li><a href="https://github.com/rollup/rollup/commit/d1e8297966f5921c249da31c51f31d8cb92d3010"><code>d1e8297</code></a> Add missing ignore</li> <li><a href="https://github.com/rollup/rollup/commit/1ba1fc23874c888605b7cbbf69e3d39f88536978"><code>1ba1fc2</code></a> fix: insert conflict numbers before first extension in multi-extension filena...</li> <li><a href="https://github.com/rollup/rollup/commit/532bd0ade7e796d6b5990c22fe0ac033464bb8aa"><code>532bd0a</code></a> Use import attributes for importing JSON (<a href="https://redirect.github.com/rollup/rollup/issues/6393">#6393</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/2cd8194aa81b2bafaad049b8de06aa7602eb81ac"><code>2cd8194</code></a> fix: advance value past wildcard prefix before suffix check in getLogFilter (...</li> <li><a href="https://github.com/rollup/rollup/commit/dfac590bd53405167d18c8a6ca6aefc83b854886"><code>dfac590</code></a> fix(deps): update minor/patch updates (<a href="https://redirect.github.com/rollup/rollup/issues/6418">#6418</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/1d6db3d32587e7ff46f7fff8eabf18b85ef8ea50"><code>1d6db3d</code></a> chore(deps): update dependency eslint-plugin-unicorn to v66 (<a href="https://redirect.github.com/rollup/rollup/issues/6419">#6419</a>)</li> <li>Additional commits viewable in <a href="https://github.com/rollup/rollup/compare/v4.61.1...v4.62.2">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
ad961227f5 |
feat(secrets): add user-specific runtime secrets (#8825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs often need provider credentials, API tokens, and other environment-bound secrets. > - Company-level secrets work for shared credentials, but they do not model values that should differ by human operator. > - Without a user-scoped model, a run can dispatch without knowing whether the responsible human has supplied the needed value. > - Paperclip also needs run attribution to make those user-scoped runtime checks deterministic and auditable. > - This pull request adds user-specific secret definitions, per-user values, environment bindings, responsible-user attribution, and runtime resolution gates. > - The benefit is that teams can define the secret once, let each user provide their own value, and block runs before dispatch when required user secrets or active definitions are unavailable. ## Linked Issues or Issue Description Refs #224 Refs #6057 This PR implements user-specific secret support as a core secret-management capability rather than a one-off adapter setting. It is related to existing public work on company secrets UI and runtime secret refs, but is distinct because the value is owned by the responsible user and resolved at run dispatch time. Related PR search before opening found existing secrets work such as #1550, #8256, #8614, #8634, and #8647; none of those add the full user-secret definition/value/runtime gate covered here. ## What Changed - Added user-secret definitions and per-user "My secrets" values, keeping stored values out of access metadata. - Added `user_secret_ref` environment bindings and UI affordances to pick them alongside existing secret refs. - Added responsible-user runtime resolution so user-secret refs resolve against the human responsible for the run. - Added pre-dispatch missing-secret gates so runs fail before adapter dispatch when required user values are absent or definitions are inactive. - Added low-trust allowlist hardening for user-secret runtime access. - Added issue, routine, run, and agent API key responsible-user attribution and fail-closed dispatch behavior when attribution cannot be resolved. - Added denial-copy mapping so responsible-user authorization failures surface as actionable run outcomes instead of opaque setup failures. - Added OpenAPI documentation for the user-secret routes. - Rebases cleanly on current `master`; migrations were renumbered incrementally as `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant` after upstream `0126`/`0127` migrations. - Removed previously committed local design screenshots so the PR contains code/docs/tests only. ## Verification - PASS: PR head `2527febd106bcf3ca264ca0da7fca491084192d6` is based on `paperclipai/paperclip:master`. - PASS: `git diff --check` - PASS: `git diff --name-only public/master...HEAD | rg '^(pnpm-lock\\.yaml|\\.github/workflows/|screenshots/)' || true` produced no files. - PASS: migration journal audit confirmed unique indexes through `130` with tail entries `0126_issue_comment_derived_attribution`, `0127_environment_custom_images_instance_scoped`, `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant`. - PASS: `pnpm --filter @paperclipai/ui typecheck` - PASS: `pnpm --filter @paperclipai/server typecheck` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-active-run-output-watchdog.test.ts src/__tests__/heartbeat-stale-queue-invalidation.test.ts src/__tests__/heartbeat-workspace-finalize-branch.test.ts src/__tests__/issue-monitor-scheduler.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-comment-wake-batching.test.ts src/__tests__/heartbeat-retry-scheduling.test.ts src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts src/__tests__/heartbeat-plugin-environment.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/low-trust-red-team-routes.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` (55 tests) - PASS: `pnpm vitest run server/src/__tests__/secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` (89 tests after final Greptile cleanup fixes) - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-issue-liveness-escalation.test.ts` (17 tests after the final rebase CI fix) - PASS: focused server Vitest batches covering heartbeat recovery, project env, plugin env, routines, low-trust, pipelines, monitors, watchdog, and stale queue paths. - PASS: GitHub checks are green on `2527febd106bcf3ca264ca0da7fca491084192d6`, including Typecheck + Release Registry, Build, General tests, serialized server suites, e2e, Canary Dry Run, verify, security checks, and Greptile Review. - PASS: Greptile Review completed successfully on `2527febd106bcf3ca264ca0da7fca491084192d6` with Confidence Score 5/5, and GraphQL review-thread audit returned zero unresolved non-outdated threads. ## Risks - Runtime behavior now depends on a run having a correct responsible user; missing or incorrect responsibility assignment can block runs before adapter dispatch. - `user_secret_ref` bindings intentionally expose metadata without values, but UI/API callers may need to handle the new binding kind explicitly. - External secret providers and IAM policies are not automatically provisioned by this PR; operators still need to configure provider-side access for non-local vaults. - The PR is broad across db/shared/server/UI/runtime paths, so release validation should include both API and UI secret workflows before merge. - The migration renumbering is intentionally incremental after upstream migrations; the branch migrations use guarded column/table/index/constraint creation so users who tested the older draft numbering should not hit duplicate DDL for the existing objects. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), Codex local adapter with shell/tool use and code execution. Context window and internal reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2f94a66ba1 |
Show live descendant status in inbox rows (#8876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The inbox is where operators quickly scan which issues are active, blocked, or waiting for attention > - A blocked parent can still have active descendant work, but the inbox previously depended on only loaded rows to infer that state > - That made collapsed or partially loaded issue trees look more stuck than they really were > - This pull request carries live descendant summary data through the issue list API and inbox UI > - The benefit is a more accurate blocked-inbox signal, so operators can distinguish truly stalled work from blocked parents that still have live child activity ## Linked Issues or Issue Description No public GitHub issue was found for this exact inbox descendant-status polish. Feature request fields: **Subsystem affected** Cross-cutting: `server/`, `packages/shared`, plugin/MCP API surfaces, and `ui/` inbox rendering. **Problem or motivation** Inbox rows need to show when blocked or collapsed parents still have live descendant work, even when the live child row is not loaded in the current client tree. Without a server-provided descendant summary, a parent can look stalled even though active work continues below it. **Proposed solution** Expose an optional live descendant count on issue list results, request it from inbox views, and use it to render covered blocked status and live-below indicators. Keep the field opt-in so other issue list callers keep their existing payload shape and query cost. **Alternatives considered** Relying only on client-loaded subtree state was ruled out because it misses collapsed or unloaded descendants. Always returning the count was also avoided because most list callers do not need this extra summary. **Roadmap alignment** This is scoped operator-visibility polish for the existing inbox. It does not duplicate a named `ROADMAP.md` milestone. **Additional context** The recursive summary query is guarded against parent cycles, and the UI still falls back to loaded subtree live counts when server summary data is absent or stale. ## What Changed - Added optional `includeLiveDescendantSummary` support to issue list contracts, SDK surfaces, MCP tools, routes, services, and tests. - Added `liveDescendantCount` to issue list results when requested. - Updated inbox and blocked-inbox queries to request live descendant summaries. - Updated inbox row status rendering so blocked parents with live descendants show covered blocker treatment without duplicating the live-below chip. - Hardened live descendant summary traversal against parent cycles and preserved the loaded-subtree fallback path for blocked inbox rows. - Added focused tests for the API parameter, service behavior, helper logic, cycle handling, and inbox UI query/rendering behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/issue-list-assignee-filter-routes.test.ts ui/src/lib/inbox-live-descendants.test.ts ui/src/components/IssueColumns.test.tsx ui/src/components/BlockedInboxView.test.tsx ui/src/pages/Inbox.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - Rebased cleanly onto current upstream `master` before pushing. - Confirmed the branch diff does not include `pnpm-lock.yaml` or `.github/workflows/*` changes. ## Risks Low to moderate risk. The new descendant count is opt-in on list requests, but it adds query work when the inbox asks for it. The recursive traversal now tracks visited ancestors to avoid cycle failures. The UI uses the server count as a supplement to existing loaded-tree state, so stale or absent counts fall back to the prior behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-enabled with local shell and git access. Reasoning mode and context window are managed by the Paperclip/Codex runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f3fa07d59 |
build(deps): bump @pierre/diffs from 1.1.22 to 1.2.11 (#8744)
Bumps @pierre/diffs from 1.1.22 to 1.2.11. [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
29a26cbd65 |
build(deps): bump @codemirror/state from 6.5.4 to 6.7.0 (#8736)
Bumps [@codemirror/state](https://github.com/codemirror/state) from 6.5.4 to 6.7.0. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/codemirror/state/blob/main/CHANGELOG.md">@codemirror/state's changelog</a>.</em></p> <blockquote> <h2>6.6.0 (2026-03-12)</h2> <h3>New features</h3> <p><code>EditorSelection.range</code> now takes an optional <code>assoc</code> argument.</p> <p><code>SelectionRange.extend</code> can now be given a third argument to specify associativity.</p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/state/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
8d9f9fd240 |
Add reusable sandbox custom images (#8794)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A growing part of that work runs in sandboxed environments rather than on the operator's local machine. > - Today sandbox providers can start fresh workspaces and run probes, but they do not have a shared contract for capturing and reusing prepared sandbox state. > - Operators need a way to set up tools, credentials, and project dependencies once, then reuse that prepared image for later agent runs. > - This pull request adds reusable sandbox custom images across the provider contract, server runtime, and board UI. > - It also keeps probes and sandbox copy flows aligned with pre-authenticated/custom-image environments. > - The benefit is faster, more reliable sandbox runs without repeatedly rebuilding the same environment setup. ## Linked Issues or Issue Description No public GitHub issue was found for this change. Inline feature request follows. ### Problem or motivation Sandboxed agents need reusable prepared runtime state so repeated runs do not require manual setup every time. Operators often need system packages, CLIs, SDKs, dependency caches, credentials, and project tooling available before an agent can work productively. ### Proposed solution Add a provider-level custom-image capability, server-side setup/capture lifecycle, Daytona/fake provider support, and board UI controls for creating, testing, selecting, and deleting custom images. ### Alternatives considered Leaving this as provider-specific setup outside Paperclip would keep the control plane blind to image state and would not give agents consistent environment metadata. Re-running setup commands for every lease is simpler, but slower and less reliable for interactive or credentialed setup. ### Roadmap alignment Checked `ROADMAP.md`; this aligns with the Cloud / Sandbox agents roadmap area and does not duplicate any related public issue or PR found by search. Additional context: - Subsystem affected: cross-cutting (`packages/db`, `packages/shared`, `packages/plugins`, `server`, `ui`). - Duplicate search: searched GitHub for `sandbox custom image` and `sandbox template environment`; no related public issues or PRs were found. ## What Changed - Added custom-image shared types, validators, constants, API paths, and database schema/migration. - Added server services/routes for custom-image templates and setup sessions, including runtime cleanup and provider metadata handling. - Extended plugin/sandbox provider capabilities for interactive setup, template capture, and template deletion. - Implemented custom-image support in the fake sandbox provider and Daytona provider. - Updated environment runtime/config handling so active custom images flow into leases, probes, and agent execution. - Added board UI controls and API client support for custom-image setup, capture, selection, status, and error states. - Hardened sandbox copy/probe behavior for insecure clipboard contexts and pre-authenticated sandbox images. - Added targeted coverage across shared validators, DB schema, server routes/services, provider plugins, adapter probes, and UI flows. ## Verification - `pnpm install --frozen-lockfile --ignore-scripts` - `pnpm vitest run packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/test.ts packages/adapters/codex-local/src/server/test.remote.test.ts packages/adapters/codex-local/src/server/test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm vitest run packages/db/src/environment-custom-images-schema.test.ts packages/shared/src/environment-custom-images.test.ts packages/shared/src/validators/plugin.test.ts server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/workspace-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts ui/src/pages/CompanyEnvironments.test.tsx ui/src/pages/CompanySettings.test.tsx` - `pnpm -r typecheck` - `pnpm build` - `rm -rf packages/db/dist && pnpm test:run` - Public-safety scan of the final diff found no internal Paperclip issue links, private instance URLs, or real secret patterns. ## Risks - Adds a database migration and new environment runtime tables, so migration ordering and rollback need care. - Provider implementations may differ in how reliably they can capture/delete images; unsupported providers surface capability-gated UI states. - Custom-image state can contain operator-prepared tooling and credentials inside the provider image, so providers must enforce their own access controls and cleanup semantics. - Broad surface area across shared contracts, server runtime, plugins, adapters, and UI means CI and Greptile review should be watched closely. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (`gpt-5`) via Codex CLI with tool use and code execution. Assisted with branch cleanup, conflict resolution, local verification, and PR preparation. Earlier branch implementation work was assisted by Paperclip-managed Claude/Codex agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed65d08d57 |
[codex] Gate skill mutations with skills:create permission (#8616)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents and board users operate inside a company-scoped control plane where permissions decide which mutating actions they can perform > - Company skills are part of the reusable agent-company setup surface, but skill mutation had been coupled to broader agent-creation authority > - That coupling meant importing or managing skills required a permission that also implies hiring power, which is broader than the operation needs > - Paperclip already has a grant-based permission vocabulary, so skill mutation should be authorized through a dedicated `skills:create` capability while preserving existing default behavior for trusted agents > - This pull request adds the skill creation permission contract, enforces it on company skill mutations, exposes it in agent permission management, and documents the changed CLI/API expectations > - The benefit is a narrower, auditable permission path for skill import/create/update/delete flows without forcing agents to receive broader agent-creation authority ## Linked Issues or Issue Description No public issue is linked. Problem: company skill mutation APIs were effectively tied to broader agent creation authority. This PR splits skill mutation authorization onto the public `skills:create` permission while keeping existing default skill creation behavior for agents unless explicitly disabled. Related public PR found during duplicate search: #5330. That PR uses an older `canManageSkills` shape; this PR implements the `skills:create` grant path instead. ## What Changed - Added `skills:create` to shared permission constants and agent permission types/validators as `canCreateSkills`. - Backfilled default human/member role grants for `skills:create`. - Updated company skill mutation routes to require board/user or agent access to `skills:create`, while preserving legacy/default agent behavior through `canCreateSkills` unless explicitly disabled. - Updated agent permission update handling, UI permission controls, duplicate-agent payloads, plugin SDK fixtures, and agent detail API surfaces for `canCreateSkills`. - Added regression coverage for skill route authorization, permission schema/default behavior, invite grants, omitted permission updates, and duplicate-agent payloads. - Updated CLI and Paperclip skill documentation for the new skill creation permission. ## Verification - `pnpm exec vitest run server/src/__tests__/agent-permissions-service.test.ts server/src/__tests__/agent-permissions-routes.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/invite-join-grants.test.ts ui/src/lib/duplicate-agent-payload.test.ts` — 5 files, 90 tests passed. - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/server typecheck && pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm test:run ...changed files...` was attempted first, but the stable wrapper rejects explicit file arguments; direct Vitest was used for the same targeted files. ## Risks - Moderate authorization risk: this changes the gate for company skill mutations, so the tests cover board grant checks, agent explicit grant checks, legacy default allowance, and explicit denial. - Migration/backfill risk is low: the migration only grants `skills:create` to existing human roles that already need broad management capability. - UI/API compatibility risk is low: `canCreateSkills` remains default-on for full agent permissions, and the update validator preserves omitted values so unrelated permission edits do not re-enable disabled skill creation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with terminal/tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b3209486ca |
fix: default Daytona sandboxes to auto-archive so they leave the disk quota (#8561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run their work inside sandboxes, provisioned through pluggable sandbox-provider plugins; the Daytona plugin is one of them > - Daytona bills storage against an org-wide disk quota, and a *stopped* sandbox still counts against that quota — only an *archived* sandbox is moved to cold storage and stops counting > - The Daytona plugin created sandboxes without supplying any auto-stop/auto-archive/auto-delete intervals, so Daytona fell back to its own defaults (archive after 7 days, never auto-delete) > - When in-product cleanup fails or never runs (crashed runs, failed lease destroys, orphaned probes), stopped sandboxes then sit for a week at a few GiB each until the org storage quota fills and blocks all workers > - This pull request makes the Daytona plugin apply quota-safe defaults on create (auto-stop 15m, auto-archive 60m, auto-delete 7d) so every sandbox eventually leaves the disk quota on its own, even when our own cleanup fails > - The benefit is that the storage quota no longer fills from leaked/idle sandboxes, while operators keep full control to override the intervals per environment ## Linked Issues or Issue Description No public GitHub issue. Describing in-PR (bug report): ### What happened Running many agents in Daytona sandboxes eventually exhausted the org's storage quota, and Daytona returned an out-of-disk error that blocked all sandbox creation. Root cause: a *stopped* Daytona sandbox still consumes disk quota; only an *archived* sandbox is moved to cold storage and frees it. The plugin created sandboxes without specifying auto-stop/auto-archive/auto-delete intervals, so Daytona used its built-in defaults (auto-archive after 7 days, auto-delete disabled). Any sandbox that our own cleanup failed to remove — or that was orphaned by a crashed run — therefore lingered for up to a week, accumulating until the quota filled. ### Expected behavior Idle/leaked sandboxes should leave the storage quota automatically within a short window, without relying solely on in-product cleanup succeeding. ### Steps to reproduce 1. Configure the Daytona sandbox provider with no explicit auto-stop/auto-archive/auto-delete intervals. 2. Run many agent sandboxes over time (or leak some via crashed/failed runs that skip in-product cleanup). 3. Observe stopped-but-not-archived sandboxes accumulating against the org's 30 GiB storage quota until Daytona returns an out-of-disk error and blocks new sandbox creation. ### Deployment mode Self-hosted Paperclip using the Daytona sandbox-provider plugin. ## What Changed - `daytona/src/plugin.ts`: `parseDriverConfig` now defaults `autoStopInterval` → 15 min, `autoArchiveInterval` → 60 min, `autoDeleteInterval` → 7 days when the value is unset. Explicit per-environment values (including `0` and `-1`) are preserved and passed through unchanged. - `daytona/src/manifest.ts`: documents the new defaults and the quota rationale on each field, and sets the manifest `default` so the values surface in the environment configuration screen where operators can override them. - `daytona/src/plugin.test.ts`: adds tests covering the new defaults and that explicit values / disabling sentinels (`0`, `-1`) are respected. ## Verification ``` pnpm --filter @paperclipai/plugin-daytona test ``` - 24 tests pass, including the new default-behavior tests. - Confirmed an explicit `autoArchiveInterval: 0` / `autoDeleteInterval: -1` in config is still forwarded unchanged (defaults only apply when the field is absent). ## Risks Low risk. Behavior change is limited to sandboxes created with no explicit interval config: they now auto-stop/archive/delete on Daytona-managed timers instead of Daytona's longer built-in defaults. Auto-archive is reversible (resuming an archived sandbox restores it). The only persistent action is auto-delete, which is a 7-day backstop for sandboxes nobody resumes, matching prior intent. Any environment that needs long-lived warm sandboxes can override the intervals (including disabling them with `0`/`-1`). ## Model Used Claude Opus (claude-opus-4-8), via the Paperclip agent harness with tool use. Extended reasoning enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ea1e321d86 |
Fix Daytona resource overrides for image-backed sandboxes (#8564)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Remote sandbox providers are part of the execution environment layer that lets agents run outside the local host. > - The Daytona sandbox-provider plugin exposes CPU, memory, disk, and GPU settings so operators can request larger execution sandboxes. > - The plugin was forwarding resource settings through snapshot/default creation, but Daytona rejects resource overrides on that path. > - Image-backed Daytona creation does support resource settings, so Paperclip should only send those values when an image is configured. > - This pull request makes the Daytona contract explicit in validation and runtime acquisition. > - The benefit is that operators get a clear Paperclip error for unsupported snapshot/default resource configs, while image-backed sandboxes still receive the requested allocation. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? Daytona sandbox environments can be configured with CPU/memory/disk/GPU resource settings, but snapshot/default Daytona sandbox creation rejects those resource overrides. Paperclip could pass unsupported resource fields through to Daytona and surface an opaque provider error. ### Expected behavior Paperclip should only send resource fields on Daytona creation paths that support them, and it should reject unsupported resource configurations before creating a sandbox. ### Steps to reproduce 1. Configure a Daytona sandbox environment with resource values such as `cpu: 4` and `memory: 4`. 2. Leave the environment on snapshot/default creation by not setting an image. 3. Acquire a Daytona sandbox lease. 4. Observe Daytona reject the resource override on the snapshot/default path. ### Paperclip version or commit Reproduced against the current Daytona sandbox-provider plugin behavior before this PR. This PR fixes the contract in the plugin code and tests. ### Deployment mode Local authenticated Paperclip instance using the Daytona sandbox provider. Additional scope: - Daytona provider only. - No schema changes. - No non-Daytona provider changes. Related search performed: - `gh search issues "Daytona resources snapshot repo:paperclipai/paperclip" --limit 10` - `gh search prs "Daytona resources snapshot repo:paperclipai/paperclip" --limit 10` ## What Changed - Restrict Daytona `resources` create params to image-backed sandbox creation. - Add validation/runtime guard for resource settings without an image. - Keep image-backed resource metadata and reusable-lease sentinel behavior covered by tests. - Update Daytona plugin tests for image-backed resources and snapshot/default rejection. ## Verification - `pnpm -C packages/plugins/sandbox-providers/daytona test` - `pnpm -C packages/plugins/sandbox-providers/daytona build` - Manual Daytona SDK probe in a local authenticated instance: image-backed creation with `daytonaio/sandbox:0.8.0` returned a 4 CPU / 4 GiB sandbox and deleted cleanly. - Greptile Review: 5/5, no inline comments. ## Risks - Existing Daytona environments that set resource values while relying on snapshot/default creation will now fail validation/acquisition with a clear error instead of calling Daytona and failing there. - Operators who need custom resource sizes should use image-backed creation or a provider-side resource-sized snapshot workflow. - Low migration risk: no database schema changes and no changes to non-Daytona providers. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent. Exact context window is not exposed in this environment. Tool-enabled workflow with shell, GitHub CLI, database inspection, and local test/build execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
71acf83070 |
build(deps): align react-dom with react@19.2.7 (replaces dependabot #8471) (#8557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI, plugin examples, and plugin SDK all share the same React dependency graph. > - Dependabot PR #8471 raised `react` and `@types/react`, but left `react-dom` and `@types/react-dom` on older 19.x ranges. > - That let dependency resolution install incompatible React runtime versions, which makes React refuse to run and breaks UI-oriented test suites. > - Current PR policy keeps `pnpm-lock.yaml` owned by CI for non-dependabot authors, so this PR updates manifests and lets CI regenerate the lockfile artifact. > - This pull request replaces #8471 with a complete React toolchain alignment across the package manifests it touched. > - The benefit is a single React 19.2.7 runtime graph that keeps dependency consumers from deduping onto a stale `react-dom` patch. ## Linked Issues or Issue Description Refs #8471 This PR replaces #8471, which updated `react` and `@types/react` but left `react-dom` and `@types/react-dom` behind. The resulting dependency graph can install mismatched `react` and `react-dom` versions, producing React's incompatible-version runtime error in UI tests. ## What Changed - Aligned `react-dom` package ranges to `^19.2.7` anywhere the React dependency set is present. - Aligned `@types/react-dom` package ranges to the 19.2 line. - Kept the `react@^19.2.7` and `@types/react@^19.2.17` bumps from #8471. - Added root pnpm overrides for `react` and `react-dom` so peer-only consumers resolve to the same 19.2.7 runtime instead of a stale patch. - Left `pnpm-lock.yaml` out of the PR, per the PR workflow policy; CI regenerates and shares the lockfile artifact when manifests change. ## Verification - `cd ui && NODE_ENV=test npx vitest run` passed locally before handoff: 243 files / 1739 tests. - `NODE_ENV=development pnpm install --lockfile-only --ignore-scripts --no-frozen-lockfile` passed locally on the rebased branch. - Confirmed regenerated `pnpm-lock.yaml` resolves `react` and `react-dom` to 19.2.7 and contains no `react@19.2.4` / `react-dom@19.2.4` entries. - GitHub Actions are green on this PR, including `policy`, `Typecheck + Release Registry`, all general test shards, `Build`, `e2e`, and `verify`. - Greptile completed at 5/5 with zero unresolved threads. ## Risks Low risk. This is a dependency-version alignment only, but React dependency changes can expose package-manager resolution issues. The root overrides intentionally constrain the React runtime pair to the same patch version to avoid that class of failure. ## Model Used OpenAI Codex based on GPT-5, using command-line tool execution and GitHub CLI operations in a Paperclip-managed workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
65d45688fe |
build(deps-dev): bump esbuild from 0.28.0 to 0.28.1 (#8470)
[//]: # (dependabot-start) ⚠️ **Dependabot is rebasing this PR** ⚠️ Rebasing might not happen immediately, so don't worry if this takes some time. Note: if you make any changes to this PR yourself, they will take precedence over the rebase. --- [//]: # (dependabot-end) Bumps [esbuild](https://github.com/evanw/esbuild) from 0.28.0 to 0.28.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/evanw/esbuild/releases">esbuild's releases</a>.</em></p> <blockquote> <h2>v0.28.1</h2> <ul> <li> <p>Disallow <code>\</code> in local development server HTTP requests (<a href="https://github.com/evanw/esbuild/security/advisories/GHSA-g7r4-m6w7-qqqr">GHSA-g7r4-m6w7-qqqr</a>)</p> <p>This release fixes a security issue where HTTP requests to esbuild's local development server could traverse outside of the serve directory on Windows using a <code>\</code> backslash character. It happened due to the use of Go's <code>path.Clean()</code> function, which only handles Unix-style <code>/</code> characters. HTTP requests with paths containing <code>\</code> are no longer allowed.</p> <p>Thanks to <a href="https://github.com/dellalibera"><code>@dellalibera</code></a> for reporting this issue.</p> </li> <li> <p>Add integrity checks to the Deno API (<a href="https://github.com/evanw/esbuild/security/advisories/GHSA-gv7w-rqvm-qjhr">GHSA-gv7w-rqvm-qjhr</a>)</p> <p>The previous release of esbuild added integrity checks to esbuild's npm install script. This release also adds integrity checks to esbuild's Deno install script. Now esbuild's Deno API will also fail with an error if the downloaded esbuild binary contains something other than the expected content.</p> <p>Note that esbuild's Deno API installs from <code>registry.npmjs.org</code> by default, but allows the <code>NPM_CONFIG_REGISTRY</code> environment variable to override this with a custom package registry. This change means that the esbuild executable served by <code>NPM_CONFIG_REGISTRY</code> must now match the expected content.</p> <p>Thanks to <a href="https://github.com/sondt99"><code>@sondt99</code></a> for reporting this issue.</p> </li> <li> <p>Avoid inlining <code>using</code> and <code>await using</code> declarations (<a href="https://redirect.github.com/evanw/esbuild/issues/4482">#4482</a>)</p> <p>Previously esbuild's minifier sometimes incorrectly inlined <code>using</code> and <code>await using</code> declarations into subsequent uses of that declaration, which then fails to dispose of the resource correctly. This bug happened because inlining was done for <code>let</code> and <code>const</code> declarations by avoiding doing it for <code>var</code> declarations, which no longer worked when more declaration types were added. Here's an example:</p> <pre lang="js"><code>// Original code { using x = new Resource() x.activate() } <p>// Old output (with --minify)<br /> new Resource().activate();</p> <p>// New output (with --minify)<br /> {using e=new Resource;e.activate()}<br /> </code></pre></p> </li> <li> <p>Fix module evaluation when an error is thrown (<a href="https://redirect.github.com/evanw/esbuild/issues/4461">#4461</a>, <a href="https://redirect.github.com/evanw/esbuild/pull/4467">#4467</a>)</p> <p>If an error is thrown during module evaluation, esbuild previously didn't preserve the state of the module for subsequent module references. This was observable if <code>import()</code> or <code>require()</code> is used to import a module multiple times. The thrown error is supposed to be thrown by every call to <code>import()</code> or <code>require()</code>, not just the first. With this release, esbuild will now throw the same error every time you call <code>import()</code> or <code>require()</code> on a module that throws during its evaluation.</p> </li> <li> <p>Fix some edge cases around the <code>new</code> operator (<a href="https://redirect.github.com/evanw/esbuild/issues/4477">#4477</a>)</p> <p>Previously esbuild incorrectly printed certain edge cases involving complex expressions inside the target of a <code>new</code> expression (specifically an optional chain and/or a tagged template literal). The generated code for the <code>new</code> target was not correctly wrapped with parentheses, and either contained a syntax error or had different semantics. These edge cases have been fixed so that they now correctly wrap the <code>new</code> target in parentheses. Here is an example of some affected code:</p> <pre lang="js"><code>// Original code new (foo()`bar`)() new (foo()?.bar)() <p>// Old output<br /> new foo()<code>bar</code>();<br /> new (foo())?.bar();</p> <p></code></pre></p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/evanw/esbuild/blob/main/CHANGELOG.md">esbuild's changelog</a>.</em></p> <blockquote> <h2>0.28.1</h2> <ul> <li> <p>Disallow <code>\</code> in local development server HTTP requests (<a href="https://github.com/evanw/esbuild/security/advisories/GHSA-g7r4-m6w7-qqqr">GHSA-g7r4-m6w7-qqqr</a>)</p> <p>This release fixes a security issue where HTTP requests to esbuild's local development server could traverse outside of the serve directory on Windows using a <code>\</code> backslash character. It happened due to the use of Go's <code>path.Clean()</code> function, which only handles Unix-style <code>/</code> characters. HTTP requests with paths containing <code>\</code> are no longer allowed.</p> <p>Thanks to <a href="https://github.com/dellalibera"><code>@dellalibera</code></a> for reporting this issue.</p> </li> <li> <p>Add integrity checks to the Deno API (<a href="https://github.com/evanw/esbuild/security/advisories/GHSA-gv7w-rqvm-qjhr">GHSA-gv7w-rqvm-qjhr</a>)</p> <p>The previous release of esbuild added integrity checks to esbuild's npm install script. This release also adds integrity checks to esbuild's Deno install script. Now esbuild's Deno API will also fail with an error if the downloaded esbuild binary contains something other than the expected content.</p> <p>Note that esbuild's Deno API installs from <code>registry.npmjs.org</code> by default, but allows the <code>NPM_CONFIG_REGISTRY</code> environment variable to override this with a custom package registry. This change means that the esbuild executable served by <code>NPM_CONFIG_REGISTRY</code> must now match the expected content.</p> <p>Thanks to <a href="https://github.com/sondt99"><code>@sondt99</code></a> for reporting this issue.</p> </li> <li> <p>Avoid inlining <code>using</code> and <code>await using</code> declarations (<a href="https://redirect.github.com/evanw/esbuild/issues/4482">#4482</a>)</p> <p>Previously esbuild's minifier sometimes incorrectly inlined <code>using</code> and <code>await using</code> declarations into subsequent uses of that declaration, which then fails to dispose of the resource correctly. This bug happened because inlining was done for <code>let</code> and <code>const</code> declarations by avoiding doing it for <code>var</code> declarations, which no longer worked when more declaration types were added. Here's an example:</p> <pre lang="js"><code>// Original code { using x = new Resource() x.activate() } <p>// Old output (with --minify)<br /> new Resource().activate();</p> <p>// New output (with --minify)<br /> {using e=new Resource;e.activate()}<br /> </code></pre></p> </li> <li> <p>Fix module evaluation when an error is thrown (<a href="https://redirect.github.com/evanw/esbuild/issues/4461">#4461</a>, <a href="https://redirect.github.com/evanw/esbuild/pull/4467">#4467</a>)</p> <p>If an error is thrown during module evaluation, esbuild previously didn't preserve the state of the module for subsequent module references. This was observable if <code>import()</code> or <code>require()</code> is used to import a module multiple times. The thrown error is supposed to be thrown by every call to <code>import()</code> or <code>require()</code>, not just the first. With this release, esbuild will now throw the same error every time you call <code>import()</code> or <code>require()</code> on a module that throws during its evaluation.</p> </li> <li> <p>Fix some edge cases around the <code>new</code> operator (<a href="https://redirect.github.com/evanw/esbuild/issues/4477">#4477</a>)</p> <p>Previously esbuild incorrectly printed certain edge cases involving complex expressions inside the target of a <code>new</code> expression (specifically an optional chain and/or a tagged template literal). The generated code for the <code>new</code> target was not correctly wrapped with parentheses, and either contained a syntax error or had different semantics. These edge cases have been fixed so that they now correctly wrap the <code>new</code> target in parentheses. Here is an example of some affected code:</p> <pre lang="js"><code>// Original code new (foo()`bar`)() new (foo()?.bar)() <p>// Old output<br /> new foo()<code>bar</code>();<br /> new (foo())?.bar();<br /> </code></pre></p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/evanw/esbuild/commit/bb9db84c02433fbe37b3509f53f9f3e3cc48725e"><code>bb9db84</code></a> publish 0.28.1 to npm</li> <li><a href="https://github.com/evanw/esbuild/commit/9ff053e53b8eeb990f59355dbea365277ac45ee2"><code>9ff053e</code></a> security: add integrity checks to the Deno API</li> <li><a href="https://github.com/evanw/esbuild/commit/0a9bf2135b67c7e28989a5ba19f0f000805a5ab5"><code>0a9bf21</code></a> enforce non-negative size in gzip parser</li> <li><a href="https://github.com/evanw/esbuild/commit/e2a1a7132058ee067fe736eac15f695861b8654e"><code>e2a1a71</code></a> security: forbid <code>\\</code> in local dev server requests</li> <li><a href="https://github.com/evanw/esbuild/commit/83a2cbfc35809f4fd5152da59572d7bed7739d78"><code>83a2cbf</code></a> fix <a href="https://redirect.github.com/evanw/esbuild/issues/4482">#4482</a>: don't inline <code>using</code> declarations</li> <li><a href="https://github.com/evanw/esbuild/commit/308ad745d824c77bc607603451b257d0f2fd9a38"><code>308ad74</code></a> fix <a href="https://redirect.github.com/evanw/esbuild/issues/4471">#4471</a>: renaming of nested <code>var</code> declarations</li> <li><a href="https://github.com/evanw/esbuild/commit/f013f5f99a015bce92ec48d49181d4ad3177b29b"><code>f013f5f</code></a> fix some typos</li> <li><a href="https://github.com/evanw/esbuild/commit/aafd6e48b1088336a5f5a17e930be7e840d43d8c"><code>aafd6e4</code></a> chore: fix some minor issues in comments (<a href="https://redirect.github.com/evanw/esbuild/issues/4462">#4462</a>)</li> <li><a href="https://github.com/evanw/esbuild/commit/15300c30b5e22f7cfcbed850c246d35095658386"><code>15300c3</code></a> follow up: cjs evaluation fixes</li> <li><a href="https://github.com/evanw/esbuild/commit/1bda0c31d7697c0af44b3ab39b81e599e559a395"><code>1bda0c3</code></a> fix <a href="https://redirect.github.com/evanw/esbuild/issues/4461">#4461</a>, fix <a href="https://redirect.github.com/evanw/esbuild/issues/4467">#4467</a>: esm evaluation fixes</li> <li>Additional commits viewable in <a href="https://github.com/evanw/esbuild/compare/v0.28.0...v0.28.1">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d5088c3987 |
build(deps): bump @codemirror/view from 6.39.15 to 6.43.1 (#8468)
Bumps [@codemirror/view](https://github.com/codemirror/view) from 6.39.15 to 6.43.1. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/codemirror/view/blob/main/CHANGELOG.md">@codemirror/view's changelog</a>.</em></p> <blockquote> <h2>6.41.0 (2026-04-01)</h2> <h3>Bug fixes</h3> <p>Fix an issue where <code>EditorView.posAtCoords</code> could incorrectly return a position near a higher element on the line, in mixed-font-size lines.</p> <p>Expand the workaround for the Webkit bug that causes nonexistent selections to stay visible to be active on non-Safari Webkit browsers.</p> <h3>New features</h3> <p>The new <code>EditorView.cursorScrollMargin</code> facet can now be used to configure the extra space used when scrolling the cursor into view.</p> <h2>6.40.0 (2026-03-12)</h2> <h3>Bug fixes</h3> <p>Fix a bug that caused Shift-Enter/Backspace/Delete on iOS to lose the shift modifier when delivered to key event handlers.</p> <p>Fix an issue where <code>EditorView.moveVertically</code> could move to the wrong place in wrapped lines with a large line height.</p> <p>Make sure the selection head associativity is properly set for mouse selections made with shift held down.</p> <h3>New features</h3> <p><code>WidgetType.updateDOM</code> is now called with the previous widget value as third argument.</p> <h2>6.39.17 (2026-03-10)</h2> <h3>Bug fixes</h3> <p>Improve touch tap-selection on line wrapping boundaries.</p> <p>Make <code>drawSelection</code> draw our own selection handles on iOS.</p> <p>Fix an issue where <code>posAtCoords</code>, when querying line wrapping points, got confused by extra empty client rectangles produced by Safari.</p> <h2>6.39.16 (2026-03-02)</h2> <h3>Bug fixes</h3> <p>Perform scroll stabilization on the document or wrapping scrollable elements, when the user scrolls the editor.</p> <p>Fix an issue where changing decorations right before a composition could end up corrupting the visible DOM.</p> <p>Fix an issue where some types of text input over a selection would be read as happening in wrong position.</p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/view/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2dbaf4a7fa |
External object references across issue surfaces (#8512)
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI agents, issues, approvals, comments, and work products. > - The involved subsystem is issue context: markdown links, issue properties, related work, lists, filters, inbox/sidebar status, and plugin-provided external context. > - The gap is that URLs to external systems currently remain mostly plain links, so humans and agents must manually open them to understand status, identity, and liveness. > - This matters because external work objects such as GitHub issues and pull requests are part of the operational state of a Paperclip company. > - The implementation keeps core provider-neutral: shared contracts, storage, sync, routes, and UI surfaces live in core while providers can contribute detection and status resolution. > - This pull request adds the external object reference foundation, GitHub provider support, issue-surface rendering, filters, sidebar/list/inbox signals, and test/story coverage. > - The benefit is that linked external work becomes inspectable Paperclip context without hardcoding every provider directly into the UI. ## Linked Issues or Issue Description No public GitHub issue exists for this work. Feature request: - Problem: URLs in Paperclip issues, comments, documents, and related surfaces do not expose provider status or object identity inline. - Proposed behavior: detect supported external object URLs, persist normalized references, refresh provider status, and render concise status-aware links across issue surfaces. - Users affected: board users, agents, and maintainers who triage issues containing external work links. - Acceptance: external object references are company-scoped, provider-extensible, visible in key issue surfaces, filterable where relevant, and covered by focused shared/server/UI tests. Related PR search: - No open duplicate PRs found for `external object references`. - Closed related prior attempt: #4556. ## What Changed - Added shared external-object contracts, validators, status/liveness helpers, and plugin protocol declarations. - Added database schema and additive migrations for external objects, source mentions, and display metadata. - Added server services/routes for detecting, syncing, summarizing, refreshing, and resolving external objects across issues, documents, comments, projects, and plugins. - Added a GitHub external-object provider plus plugin SDK authoring docs. - Wired UI presentation across markdown links, comments, issue chat, documents, properties, related work, issue rows, filters, inbox/sidebar badges, and Storybook stories. - Rebasing cleanup: moved the branch onto current `master`, repaired stale worktree provision config, hardened environment-sensitive tests/mocks, and removed committed screenshot artifacts from the PR branch to keep the reviewable file set below tool limits. ## Verification - `pnpm exec vitest run packages/shared/src/external-objects.test.ts server/src/__tests__/external-object-routes.test.ts server/src/__tests__/external-objects-service.test.ts ui/src/components/ExternalObjectPill.test.tsx ui/src/lib/external-objects.test.ts` passed after rebasing: 5 files, 56 tests. - Historical branch verification before this PR creation included `pnpm test:run`, `pnpm -r typecheck`, and `pnpm build`; this PR body does not claim those were rerun after the final rebase. ## Risks - Medium: this adds a new cross-surface sync path on issue/document/comment writes. The implementation uses safe sync wrappers so external-object failures warn instead of blocking core mutations. - Medium: the migrations introduce new tables and indexes. They are additive and company-scoped. - Medium: provider-specific URL parsing can miss or misclassify edge cases. Shared canonicalization tests and provider tests cover current GitHub shapes. - Low: UI badge/filter behavior could add visual noise for object-heavy issues; component tests and Storybook stories cover the intended surfaces. > Roadmap checked: `ROADMAP.md` references the plugin system as the current extension path and does not list a duplicate core feature. Related long-range docs discuss external references, work products, preview URLs, and plugin extension points; this PR implements the scoped external-object reference foundation. ## Model Used OpenAI Codex, GPT-5 coding-agent runtime, with shell and GitHub CLI tool use. Reasoning mode: medium. Exact deployed runtime model ID and context window were not exposed in the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd38c150b0 |
feat: reuse Daytona sandbox leases (#8513)
Add opt-in reusable Daytona sandbox lease support, including retryable pending cleanup handling.\n\nPR: https://github.com/paperclipai/paperclip/pull/8513 |
||
|
|
a937b89a47 |
fix(daytona): valid memory input in env config form + size presets (#8389)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments are configured through a JSON-schema-driven form (`JsonSchemaForm`), and each sandbox provider (here, Daytona) supplies a manifest describing its fields > - In the Daytona environment config form, the "memory" field rendered as a free-text input; on save the value round-tripped to `0`, producing a server-side "number must be greater than zero" error even when the user typed a valid number like `2` > - Rather than patch the free-text input, memory is now a true dropdown of the sandbox sizes Daytona actually supports, so an invalid value can no longer be typed or coerced > - The dropdown leads with a blank "None" row that is selected by default (meaning "not configured — use Daytona's defaults"); `0` is never offered because it is not a valid configuration > - The benefit is that operators pick a valid memory size from a constrained list, it saves correctly as an integer, and leaving it unset cleanly omits the field instead of submitting `0` ## Linked Issues or Issue Description No public GitHub issue exists; describing the bug inline per the bug report template. ### What happened? In the Daytona environment configuration form, the memory field is a free-text input labelled "gigabytes of RAM". Entering a value such as `2` and saving coerced the value to `0`, and the server rejected the save with "the number must be greater than zero". There was no way to enter a valid memory size through the form. ### Expected behavior Selecting a valid memory size (e.g. `2`) keeps that value and saves cleanly as an integer. Leaving memory unset is valid and submits no value (Daytona defaults apply). `0` is never selectable. ### Steps to reproduce 1. Open the environment configuration form for a Daytona sandbox provider. 2. In the memory field, type `2`. 3. Click Save. 4. Observe the value becomes `0` and the form errors with "the number must be greater than zero". ## What Changed - Daytona manifest: `memory` is now an `enum` of the supported sandbox sizes `[1, 2, 4, 8]` (GiB). It stays optional, so "not configured" remains valid. `0` is not in the list. - `JsonSchemaForm` `EnumField`: optional enums now render a leading blank **None** row that is selected by default when no value is set, letting the user express "not configured" and clear a previous selection (Radix `Select` forbids an empty-string item value, so the unset state maps to a sentinel that translates back to `undefined`). - `JsonSchemaForm` `EnumField`: when every enum option is numeric, the selected value is coerced back to a number on change so the payload keeps the schema's integer type (a stringified `"2"` would otherwise fail server-side integer validation — this is what fixes the original bug for the dropdown path). - Tests: added `EnumField` coverage in `JsonSchemaForm.test.tsx` (blank row present, no `0`, blank selected by default, numeric coercion, blank → unset) and Daytona manifest coverage in the plugin test (memory enum is `[1,2,4,8]`, excludes `0`, stays optional). ## Verification - `cd ui && npx vitest run src/components/JsonSchemaForm.test.tsx src/pages/CompanyEnvironments.test.tsx` → 16/16 passing (14 + 2). - `vitest run` in `packages/plugins/sandbox-providers/daytona` → 19/19 passing (incl. 3 new manifest tests). - `cd ui && npx tsc --noEmit` → clean. - Behavior: the memory field now renders as a dropdown showing **None / 1 / 2 / 4 / 8**, with **None** selected by default. Picking `2` stores the integer `2`; picking **None** clears the field so it is omitted from the payload (no `0`, no "must be greater than zero" error). ## Risks - Low risk. The blank-row + numeric-coercion changes live in `EnumField`. The blank row is only added for **optional** enums (required enums are unaffected); numeric coercion only triggers when every option is numeric, so existing string enums (`egressMode`, `backend`, `sessionStrategy`) are unchanged. The manifest change is scoped to the Daytona provider. ## Model Used Claude (Anthropic), Opus-class model, used via the Paperclip agent workflow (author + reviewer agents) with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |