mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
bfc19e2ebdaba4c07f630d65ef7376912f2fcfa7
1051
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f4802b1bbc |
feat(runtime-exposure): least-privilege Tailscale HTTPS broker, shared contract, and persisted exposure state (#11524)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services for a project's execution workspaces, so an agent's branch can be previewed while it works > - Those services only listen on plain loopback HTTP. A person on another device, or on a phone, cannot open the preview > - A Tailscale HTTPS mapping solves this, but `tailscale serve` needs host privileges that the Paperclip server process must not hold > - This pull request adds the foundation only: a separate least-privilege host broker, the shared exposure contract, and the database columns that hold exposure state > - Nothing calls the broker yet, so there is no behavior change. The benefit is that the privileged surface is small, reviewable, and isolated before any lifecycle code depends on it ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, the shared type and validator package, and the database schema. **Problem or motivation** A managed runtime service binds to loopback only. There is no supported way to reach that preview from another device. Adding HTTPS directly to the server would mean the server process runs `tailscale serve`, which needs privileges far wider than the task requires. A compromised or buggy server could then map any port to the tailnet. **Proposed solution** Split the privileged work into a separate broker process with a narrow protocol, and define one shared contract that the server, the UI, the runtime, and the broker all read. Land this foundation first, with no caller, so the privileged code can be reviewed on its own. **Alternatives considered** - Call `tailscale serve` from the server process. This was rejected because it gives the server unrestricted mapping authority. - Use `sudo` for single `tailscale` commands. This was rejected because the argument list is the only guard, and it is easy to widen by accident. - Use a generic reverse proxy. This was rejected because it does not remove the need for a privileged Tailscale mapping step. **Roadmap alignment** This supports the existing managed workspace runtime capability. It adds no new product surface on its own. **Additional context** The broker is the security boundary of the feature, so it is deliberately the first slice. Three later pull requests build on it: the server exposure lifecycle, the runtime lease and recovery integration, and the leased-port mediator. ## What Changed - Add the `@paperclipai/tailscale-https-broker` workspace package. The broker listens on a unix socket, authorizes each peer with `SO_PEERCRED`, and answers a small request protocol. - Restrict what the broker will map. It accepts only same-number HTTPS-to-loopback pairs inside the Paperclip port range, refuses protected ports, and confirms that the loopback port belongs to a Paperclip-owned listener. - Parse every request with a strict JSON reader that rejects duplicate keys, prototype keys, and unknown fields. - Write an append-only audit record for each broker decision. - Add the shared exposure contract in `@paperclipai/shared`: the `RuntimeExposureConfig`, `RuntimeExposureState`, and `RuntimeExposureStatus` types, their zod validators, the app and HMR port rules, and the loopback-bind helpers. - Persist exposure state on `workspace_runtime_services` with the new `exposure` column, plus the server-private `exposure_handle` and `backend_url` columns that are never serialized to API clients. - Add the `execution_workspace_runtime_leases` table that the later lease slice uses. - Extend the runtime read-model test fixture for the three new columns. ## Verification Focused checks, all run on this branch: - `pnpm --filter @paperclipai/tailscale-https-broker test` — 12 files, 82 tests pass. This covers peer credentials, port policy, protected ports, the serve config writer, the strict JSON reader, argv parsing, and the socket server. - `pnpm --filter @paperclipai/tailscale-https-broker typecheck` — clean. - `npx vitest run --root packages/shared src/runtime-exposure src/validators/runtime-exposure.test.ts` — 3 files, 40 tests pass. - `pnpm --filter @paperclipai/db typecheck` — runs `check:migrations` first. Migration numbering and migration safety both pass. - `pnpm --filter @paperclipai/shared typecheck` — clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `npx vitest run --root server src/services/workspace-runtime-read-model.test.ts` — 3 tests pass. - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, which is exactly the count on `master` before this branch. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. To confirm the exposure state is inert, start a managed runtime service as usual. The new columns stay null and the service behaves as it does today. ## Risks - Migration risk is low. Both migrations only add a table and three nullable columns. No column is backfilled and no existing column changes. The migration safety check passes. - Behavior risk is low. No code path calls the broker in this pull request, and the shared exposure fields are optional. - The broker is privileged, so it is the real risk surface. It is mitigated by peer-credential authorization, a fixed port range, a protected-port deny list, same-number pair enforcement, listener-ownership checks, strict JSON parsing, and an audit trail. Reviewers should read `packages/tailscale-https-broker/src/authorization.ts` and `src/port-policy.ts` closely. - The broker requires a `tailscale` version floor, which its README records. An older host CLI makes the broker refuse to start rather than map incorrectly. - `pnpm-lock.yaml` changes because a new workspace package is added. The diff is the new importer block, plus one duplicate `tinyexec` entry that pnpm removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
10d0555189 |
fix(interactions): authorize resolvers consistently (#11376)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue interactions give agents and people a structured decision record. > - Resolver routes used different authorization rules. > - Some routes blocked valid agents, including task watchdogs with normal issue access. > - The API did not show who could resolve a pending interaction. > - This pull request gives every interaction kind one resolver policy evaluator. > - The benefit is a clear decision path with consistent governance and company isolation. ## Linked Issues or Issue Description Fixes: #8087 Refs: #7403 Related PR: #11082 proposes board-only confirmation rules. This change keeps human-only review as an explicit policy. **What happened?** Agents could create issue interactions. Some resolver routes still required board access. This left valid agent confirmations pending. Task watchdogs could see the same problem without board identity. **Expected behavior** Every interaction kind must use one resolver policy contract. The contract must support `anyone`, `not_creator`, and `human_only`. It must also apply all normal governance controls. **Steps to reproduce** 1. Create a `request_confirmation` interaction as an agent. 2. Resolve it with another authorized agent. 3. Observe the board-only denial. **Paperclip version or commit** The problem exists on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Add canonical policies for `anyone`, `not_creator`, and `human_only`. - Use one server evaluator for every interaction kind. - Apply named addressees, company limits, review rules, and task watchdog scope. - Charge cross-issue resolutions to the existing per-run action limit. - Return the effective resolver audience in attention and interaction data. - Show the audience, governance choices, and denial reasons in the board UI. - Add telemetry, API documents, product documents, and regression fixtures. - Add migration provenance for safe legacy behavior. - Make migration `0218` safe for complete replays and partial prior runs. ## Product Rules - An interaction records a response. It does not grant authority for the next action. - `anyone` lets any authorized issue participant respond. - `not_creator` requires a responder other than the interaction creator. - `human_only` requires an authorized person. - A named addressee, company policy, or governed action can narrow the audience. - These controls cannot widen the audience. - A task watchdog uses the same rules as an ordinary agent. - A task watchdog does not receive board authority. - An agent resolution on another issue uses the shared cross-issue action limit. - Legacy pending interactions keep their earlier restrictions. - The UI shows the effective audience and a permanent denial reason. ## Verification - `pnpm --filter @paperclipai/db check:migrations` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run packages/db/src/issue-thread-interaction-resolver-policy-migration.test.ts` - The focused PostgreSQL test applies migration `0218` twice. - The test also completes a partial prior run and preserves existing provenance. - The latest GitHub head has 29 successful checks. - The opt-in Storybook visual check skipped as expected. - Greptile reports 5/5 with no open comments. ## Risks - New interaction writes use `anyone` by default. - Callers must select `not_creator` or `human_only` when they need stricter review. - Legacy pending interactions keep the old creator and human restrictions. - Migration `0218` fills only missing provenance fields during recovery. - Cross-issue resolutions can reach the existing action limit. - The shared evaluator affects every interaction kind. - Route, service, database, shared contract, and UI tests cover these rules. > This work matches the Agent Reviews and Approvals direction in `ROADMAP.md`. It does not duplicate a planned item. ## Model Used OpenAI Codex, GPT-5. The runtime does not expose the exact deployment ID or context window. The agent used reasoning, repository tools, shell commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue with the required labels - [x] I have not referenced internal Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open comments - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
92047cac46 |
test(adapter-utils): make the next sandbox flake diagnosable (#11483)
`execution-target-sandbox` has failed twice in CI and not once in several hundred local runs. This does not fix it. It makes the next occurrence carry its own evidence, because a third unreproducible failure would teach nothing. The observed signature was an empty stdout with exit code 0 - the child exited cleanly having produced nothing, which is what a lost stdin frame looks like from the test's side. Three mechanisms were checked and ruled out rather than assumed: the helper resolving on `exit` rather than `close` (a 200-iteration probe produced no truncations, and the failure was empty rather than partial); the wrapper reporting exit before stdout drains (it already listens on `close`); and frame writes racing (the stream wrapper's `writeEvent` is synchronous and sequence-numbered). Two candidates remain and the runtime tree separates them. A stdin queue frame still present means the host wrote it and the wrapper never consumed it; a drained queue with no output means it was consumed and the reply was lost on the way back. The report prints that tree, both proxy streams, the exit code, and the elapsed time - the last because the bridge and proxy run on 5s budgets that are generous locally and tight on a runner sharing a box with 19 other lanes. Timeouts are deliberately unchanged. Raising them would probably make the symptom go away, which is the reason not to do it blind. The first revision capped the tree walk one level above the queue frames, so "the queue is empty" and "the walk never looked" printed identically - the distinction the report exists to make. Caught in review. Verifying that the reporter printed something was not enough; it had to print the thing that discriminates, which is now checked by planting a frame and forcing the assertion. adapter-utils typecheck clean; 44 pass, stable across repeated runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9e9f744f58 |
Show blocker links in the task chat (#11456)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip helps operators supervise agent work through tasks and task threads. > - The redesigned task thread shows the current work and its state. > - A blocked task did not show the dependency that prevented progress. > - Operators had to leave the thread to find the direct and final blockers. > - This pull request adds compact blocker links at the top and bottom of the task thread. > - The benefit is that operators can identify and open the relevant tasks without adding a large notice to the thread. ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread did not show which task directly blocked the current task or which task ultimately blocked its dependency chain. **Subsystem affected** `server/`, `packages/shared/`, and `ui/` task-blocker presentation. **Current behavior** A blocked task can open in the redesigned thread without a visible dependency link at the top or bottom of the conversation. **Proposed behavior** Show one compact amber row for the direct blocker. Show a second row for the selected final blocker when one exists. Render the rows at both ends of the thread. **Reason and benefit** Operators can see the reason for the blocked state and open the relevant task from the conversation. The compact rows preserve thread density. **Breaking changes** None. The new blocker-attention fields are optional. Existing clients remain compatible. ## What Changed - Added a compact task-chat component for direct and selected final blocker links. - Added the blocker rows to the top and bottom of populated and empty task threads. - Added link-ready blocker-attention details so an intermediate selected task stays on its correct direct chain. - Included blocker-link changes in the thread content key so pinned threads follow a newly added bottom row. - Added component, scrolling, server contract, and Storybook coverage for the new states. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx server/src/__tests__/issue-blocker-attention.test.ts` (38 tests passed) - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` ## Risks - Low risk. The rows only render while the task status is `blocked` and an unresolved blocker is available. - Long titles are truncated to keep each blocker on one line. The full task label remains available in the link title. - Older server payloads keep the original leaf-selection behavior because the new sampled details are optional. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The run used tool-enabled reasoning and code execution. The context-window size was not exposed to the run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cd501499a2 |
test: add ACPX run lifecycle characterization baselines (#11461)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The adapter runtime starts, turns, settles, and composes ACPX runs > - Recent lifecycle corrections changed several order and cleanup rules > - Those rules need regression coverage before the planned engine refactor > - This pull request adds characterization suites for the corrected behavior > - The benefit is a clear test baseline for the next refactor ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX adapter runtime and server heartbeat lifecycle need stable regression coverage for their current corrected behavior. **Subsystem affected** Cross-cutting (multiple of the above): `packages/adapter-utils` and `server` test suites. **Current behavior** The runtime has corrected rules for startup, turns, settlement, composed results, and heartbeat terminalization. The repository lacks a single characterization baseline for these rules. **Proposed behavior** Keep the current lifecycle rules pinned by five test suites. Let the later engine refactor change behavior only when it updates these tests with a clear reason. **Reason and benefit** The suites expose order, cleanup, transport, timeout, retry, result, and lease-release changes during the refactor. They also record one known latent defect as current behavior. **Breaking changes** None. This pull request adds tests only. ## What Changed - Add startup characterization coverage for commands, launch values, session fingerprints, sync order, bridge overlap, and cleanup paths. - Add turn characterization coverage for inputs, events, transports, timeout and cancel behavior, retry rules, errors, and usage. - Add settlement characterization coverage for teardown, adapter sync-back, workspace restore order, native sync, and error policy. - Add composed-run characterization coverage for result forms, finalization sets, and host-lane warm save and warm hit behavior. - Add server coverage that checks run terminalization before environment lease release. ## Verification - Run `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts packages/adapter-utils/src/acpx-engine/execute.test.ts`. - Run `npx vitest run server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts`. - The adapter-utils run passes 178 tests, and the server run passes 4 tests. - Check `pnpm --filter @paperclipai/adapter-utils typecheck`. - Check `pnpm --filter @paperclipai/server typecheck`. ## Risks Low risk. The change adds test files and does not change production code. One known cold ensure-session cleanup defect remains pinned as current behavior. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e52b8a343f |
fix: ACP run lifecycle corrections — failure settlement, workspace sync-back, lease cleanup (#11454)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run ACP sessions and manage runtime, workspace, and lease resources. > - Several failure paths left runtime bridges, staged workspaces, or environment leases active after an error. > - These leaks reduce run reliability and can leave later runs without clean resources. > - This pull request closes the failure paths, applies one teardown policy, and adds regression tests. > - The benefit is consistent failure settlement and safer reuse of agent workspaces and leases. ## Linked Issues or Issue Description **What happened?** ACP runs could leave runtime bridges, staged workspaces, or environment leases active after failures. Claude and Gemini ACP runs did not restore the sandbox workspace on teardown. Lease release stopped when one lease returned an error. **Expected behavior** Each ACP failure must return an error result and settle its resources. Teardown must run each step, release leases independently, and restore the host workspace when the sandbox ends. Pending cleanup leases must receive bounded retry attempts. **Steps to reproduce** 1. Run an ACP session that fails after runtime creation or during turn preparation. 2. Run an ACP session that fails during a warm hit or staged runtime handoff. 3. Run lease cleanup with more than one lease when the first release returns an error. 4. Inspect the result phase, teardown calls, workspace state, and lease metadata. 5. Run the regression suites listed in the Verification section. ## What Changed - Settle every ACP failure after runtime creation with an error result and one sandbox.startup span closure. - Close the ACP runtime and remove warm entries after every pre-turn failure. - Run all teardown steps, record teardown errors, release staging leases in finally, and prevent duplicate teardown. - Dispose staged runtimes after seam failures and remove borrowed staged entries with identity guards. - Add fail-open workspace sync-back teardown for Claude and Gemini ACP adapters. - Isolate lease release errors and add bounded retry sweeps for stranded pending_cleanup leases. - Atomically claim pending_cleanup retries and clamp attempt readers to keep the five-attempt bound. - Default absent provider reusableLeases values to false and align the fake provider with its runtime declaration. - Add regression tests for engine, adapter, server, and shared environment behavior. ## Verification - [x] `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 124 tests passed. - [x] `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/claude-local/src/server/acp.test.ts packages/adapters/gemini-local/src/server/acp.test.ts` — 61 tests passed. - [x] `npx vitest run server/src/__tests__/environment-runtime.test.ts server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts server/src/__tests__/reusable-leases-default.test.ts server/src/__tests__/environment-routes.test.ts packages/shared/src/environment-support.test.ts` — passed. - [x] All listed suites ran from the repository root. - [x] GitHub CI completed successfully for `cfc349c9f232711433897915112a1c52c0e462ca`. - [x] Greptile completed with a 5/5 confidence score and no blocking finding. ## Risks The engine changes affect failure settlement and teardown order across ACP runs. The server changes add retry state to existing lease metadata without a schema migration. The adapter changes restore workspaces after sandbox execution. Regression tests cover the changed paths. GitHub CI and Greptile passed for the current head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This change fixes runtime reliability and does not duplicate a roadmap feature. ## Model Used OpenAI GPT-5 Codex. The model used tool-based repository inspection, GitHub operations, and code review support. The runtime does not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (for example, `docs/...` or `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bc9f70f54c |
fix(plugin-daytona): bound the sandbox liveness calls with a per-call timeout (#11408)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox provider plugins run agent work in remote execution environments > - The Daytona sandbox liveness read can stay pending when the connection stops responding > - A pending read blocks the plugin until a broad host-to-worker limit expires > - This pull request adds bounded deadlines to Daytona liveness calls and clears stale handles > - The benefit is a fast and clear error when a Daytona connection stops responding ## Linked Issues or Issue Description Refs #11341 **What happened?** The Daytona sandbox liveness read had no per-call timeout. A silent connection failure left the read pending until the broad host-to-worker RPC limit expired. **Expected behavior** The plugin should stop a liveness call within a defined limit and report a clear timeout error. **Steps to reproduce** 1. Create a Daytona sandbox handle. 2. Make the cached handle freshness read never resolve. 3. Run the next sandbox operation. 4. Observe that the operation waits for the outer RPC limit without a liveness timeout. **Paperclip version or commit** `master` before this change. **Deployment mode** Any deployment mode that uses the Daytona sandbox provider. ## What Changed - Add `withLivenessTimeout` with timer cleanup and `SandboxLivenessTimeoutError`. - Bound `refreshData` with configurable `livenessTimeoutMs`, which defaults to 30000 milliseconds. - Bound sandbox start and recovery calls with the SDK timeout plus a 5000 millisecond margin. - Reject `livenessTimeoutMs` values above 86400000 milliseconds and document the setting. - Evict a cached handle after a failed freshness refresh so the next operation fetches a new handle. - Add a test for a never-resolving freshness refresh and the cached-handle eviction. ## Verification - Run the Daytona plugin test suite with its package Vitest configuration. - Confirm that 150 of 150 tests pass. - Confirm that the new test reports a bounded timeout and a fresh handle on the next operation. - Confirm that GitHub Actions reports green status checks after the pull request starts. ## Risks This change adds an early timeout only to Daytona liveness calls. A value of 0 or less disables the extra bound. The default leaves normal SDK calls within their expected time limit. The main risk is a timeout value that is too short for a slow but healthy connection. ## Model Used OpenAI Codex, GPT-5. The model used tool calls and code execution. The model supplied the PR handoff and did not author the code in this pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdb9a4880d |
fix(security): route paperclipai CLI guidance through safe npx form (CWE-78) (#11400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip provides CLI commands and guidance for operators and agents > - The `pnpm paperclipai` script can pass argument values through a shell > - Shell re-parsing can execute command substitutions inside quoted values > - This pull request routes guidance through inert-argv `npx paperclipai` commands and adds regression coverage > - The benefit is safer operator guidance across documentation and runtime hints ## Linked Issues or Issue Description This pull request fixes a command-injection-class defect in Paperclip CLI guidance. **What happened?** The `pnpm paperclipai <sub> --flag "$VALUE"` form can re-parse argument values through a shell. A command substitution inside a quoted value can execute on the host. **Expected behavior** Paperclip guidance must pass CLI values as inert argument values. Host-derived values must not appear in copyable commands. **Steps to reproduce** 1. Run a Paperclip guidance command that uses the `pnpm paperclipai` script. 2. Provide a quoted value that contains a command substitution. 3. Observe that the shell can evaluate the substitution before the CLI starts. 4. Compare the result with the `npx paperclipai` form. **Paperclip version or commit** `5670984b75d109950c968542a0111ebb6967f4da` **Deployment mode** All deployment modes that show or use the affected CLI guidance. **Installation method** Built from source and installed CLI guidance. **Agent adapter(s) involved** Not adapter-specific (core bug). **Database mode** Not database-related. **Access context** Both. **Additional context** The earlier merged PR [#11343](https://github.com/paperclipai/paperclip/pull/11343) used the unsafe `pnpm exec paperclipai` form. This fresh PR replaces that guidance with the safe `npx paperclipai` form. ## What Changed - Standardize documentation and runtime hints on `npx paperclipai`. - Remove the broken `pnpm exec paperclipai` guidance. - Use a static `<host>` placeholder in private-hostname guidance. - Add regression tests for unsafe forms, continued lines, static hosts, and offline guidance. ## Verification - `git diff --check origin/master...origin/fix/paperclipai-cli-npx-safe-invocation` passes. - The branch adds `server/src/__tests__/cli-invocation-safety.test.ts` and updates private-hostname tests. - CI must run the new tests, typecheck, lint, and build checks. - Local Vitest execution was not available because this worktree has no installed Vitest binary. ## Risks - The change affects operator and agent documentation text. - The runtime hints now show `<host>` instead of a request-derived host value. - No database schema or migration changes exist. - CI will detect any missed unsafe invocation or type error. ## Model Used OpenAI GPT-5, exact model ID `gpt-5`, with tool use and code-review assistance. The model used repository inspection, Git operations, and PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] CI ran the test suites and they pass; local test execution was unavailable in this worktree - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed8075b535 |
fix(adapter-utils): order stdin file writes in the sandbox process-session bridge (#11406)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters can run through a sandbox process-session bridge > - The bridge writes streamed standard input to files before the remote process reads them > - Concurrent file writes can make a later chunk visible before an earlier chunk > - The remote process can then parse a tail fragment and wait forever for the missing head > - This pull request serializes host writes and makes an unexpected file gap a loud error > - The benefit is ordered input with a bounded failure path for sandbox ACP sessions ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description below follows `.github/ISSUE_TEMPLATE/bug_report.yml`. **What happened?** A sandbox ACP process-session bridge could stall after its handshake when a run sent a large prompt. The host sent one un-awaited file write for each standard input chunk. A small later chunk could finish before a large earlier chunk. The remote poller then sent the tail bytes first. The agent parser raised an error on the tail fragment, and the head bytes stayed buffered without a newline. **Expected behavior** The bridge must expose standard input files in sequence. The remote poller must report a clear error when an earlier file remains missing beyond the retry budget. **Steps to reproduce** 1. Start an ACP session through a sandbox process-session bridge. 2. Send a prompt that produces multiple standard input file chunks. 3. Delay finalization of an earlier chunk while a later chunk completes. 4. Observe that the remote parser can receive the later chunk first and the session can stop without a clear error. **Paperclip version or commit** The change targets the current `master` branch at the submitted commit. **Deployment mode** The bug affects sandbox execution. **Agent adapter(s) involved** The failure affects the ACP process-session bridge. **Database mode** Not database-related. **Additional context** The fix keeps the existing per-file atomic write behavior. It adds ordering at the host write boundary and a bounded ordering check in the shared wrapper poll tail. ## What Changed - Add a per-session promise chain for host standard input file writes. - Keep a failed write from blocking later chain entries. - Track the next expected sequence number in the shared wrapper poll tail. - Hold later files while an earlier file is missing within the existing retry budget. - Emit a loud error and advance after the retry budget expires. - Add regression tests for host ordering, gap holding, and the loud error path. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` — 9 tests passed. - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` — 43 tests passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - With the source fix reverted, the 3 new tests fail and the 6 original tests pass. - CI must pass on the pull request before merge. ## Risks Low risk. - The host now serializes writes for each session, which can reduce write parallelism. - A failed write still emits one error and destroys the socket, as before. - The wrapper can emit a loud error after the existing retry budget when a file gap persists. - The change does not alter the atomic per-file write behavior. ## Model Used OpenAI GPT-5, model ID `gpt-5`, with tool use and code execution. The context window and internal reasoning details are not disclosed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f26f2a450 |
fix(adapter-utils): add per-iteration timeout and watchdog to the sandbox callback bridge poll loop (#11341)
## Thinking Path > - Paperclip runs AI agents through adapters and sandboxed execution paths > - The sandbox callback bridge carries file requests between the host and a sandbox > - The poll loop waited forever when a sandbox call stopped responding > - A permanent wait stranded queued requests and hid the run failure > - This pull request adds bounded timeouts, abort handling, recovery backstops, and trace reporting > - The benefit is prompt request failure, safe mutation outcomes, run-level error reporting, and trace visibility ## Linked Issues or Issue Description **What happened?** The sandbox callback bridge could wait forever when a client call stopped responding without a rejection. **Expected behavior** The bridge should fail queued requests and report a run-level error when the sandbox channel stops responding. **Steps to reproduce** 1. Start a sandbox callback bridge. 2. Queue a request. 3. Make the sandbox call stop responding. 4. Observe that the request does not receive a failure response. **Paperclip version or commit** Commit `edc4f71b460c600f97cf44cb486d5cac72ca2db9`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Custom or external sandbox callback bridge. **Database mode** Not database-related. ## What Changed - Add a per-iteration timeout for `listJsonFiles` and `processRequestFile`. - Add a watchdog that fails pending requests when the loop makes no progress. - Abort a hung handler and use a non-retryable 504 backstop when its outcome can be indeterminate. - Retry recovery writes and keep queued requests when a recovery write fails. - Forward the indeterminate-outcome header through the execution target. - Record worker failures through the `sandbox.callbackBridge.workerFailed` trace span. - Add tests for timeout, watchdog, recovery, mutation safety, header forwarding, and fast-request behavior. ## Verification - Run `pnpm exec vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts`. - Confirm that the PR test, typecheck, build, end-to-end, serialized test, and security checks pass. - Confirm that the current PR head is `edc4f71b460c600f97cf44cb486d5cac72ca2db9`. - Confirm that the PR changes four files: the callback bridge, its tests, the execution target, and its tests. ## Risks The default timeout can fail a slow but valid sandbox call. The defaults remain configurable, and the iteration timeout stays below the sandbox response deadline. A mutation that may have committed returns a non-retryable 504 outcome so the caller does not apply it twice. ## Model Used OpenAI Codex, GPT-5 current runtime, with extended reasoning and tool use. The exact context window is not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a53cc8819b |
fix(claude-local): pipe print prompt via stdin (#9500)
Fixes #2444. Refs #4947. The `claude_local` adapter launched Claude Code as `claude --print - --output-format stream-json --verbose`. Paperclip writes the rendered task prompt to Claude's stdin, but current Claude Code releases can treat the stale `-` positional marker as the prompt itself, so Claude received the literal string `"-"` instead of the issue body. The customer's task ran against no content at all. The fix keeps `--print` mode and stdin delivery, and removes the stale `-`. Adds regression coverage on both sides of the delivery path: a `claude_local` assertion that `--print` is present, `"-"` is absent and the prompt still reaches stdin, and an adapter-utils case proving the sandbox run-log command wrapper preserves stdin while streaming logs. Authored by @elJayAdvisor, whose commit is included unchanged with their authorship. The branch had gone stale and was showing CONFLICTING; the conflict was in `execution-target-sandbox.test.ts`, where their new test was added at the same point as master's `creates the process session directories only in the launch exec` case and git interleaved the two into one hunk. Resolved by taking master's file and re-inserting their test whole, after checking every helper it needs still exists there. Verified: the bug was still live on master at `execute.ts:838`; the regression test genuinely catches it — restoring the stale `-` fails `expect(captured.argv).not.toContain("-")`; `@paperclipai/adapter-claude-local` and `@paperclipai/adapter-utils` typecheck clean; 67 pass across the two test files. All CI gates green; Greptile 5/5. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eabecc6f77 |
feat(annotations): include issue document annotations in agent review context (#11332)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Reviewers annotate plans and issue documents with inline comments, and assigned agents act on that feedback > - The server already builds a bounded review context from open plan annotations and includes it in agent wake payloads > - Non-plan issue documents did not get the same treatment: their open annotation threads never reached the agent, and the properties pane did not surface their annotations > - This pull request extends the review-context path and the properties-pane UI to issue documents, at parity with plans > - The benefit is that agent feedback on any issue document reaches the assigned agent, not only feedback on the plan ## Linked Issues or Issue Description **What existing behavior does this improve?** The review-context pipeline that delivers inline annotation feedback to assigned agents, and the properties pane that surfaces those annotations to reviewers. **Subsystem affected** The server review-context path (`server/src/services/plan-review-context.ts`, wake payload assembly in `server/src/services/heartbeat.ts`, `server/src/routes/issues.ts`), shared wake-payload types (`packages/shared`, `packages/adapter-utils`), and the issue properties pane (`ui/src/components/issue-properties/`). **Current behavior** A reviewer can annotate any issue document, not only the plan. The agent wake payload includes open annotation threads for the plan document only. Feedback left on other issue documents is invisible to the assigned agent. In the properties pane, the Artifacts tab also gives no way to see or open a document's annotations. **Proposed behavior** Add `buildDocumentReviewContext` beside the existing plan builder. It collects open annotation threads for all non-plan issue documents, applies the same thread, comment, and character budgets across documents, and reports truncation. Include the result as a new `documentReviewContext` field in agent wake payloads and in the issue wake-context route. Keep the plan context on its legacy builder and field so plan-only wakes stay byte-for-byte compatible. Render the new context in the adapter wake-payload text, and surface annotation counts and the annotation panel for documents in the properties pane's Plans and Artifacts tabs. **Reason and benefit** The floating annotation popover and persistent highlight UI landed earlier; this change completes the loop so agent feedback on any issue document reaches the assigned agent, not only feedback on the plan. **Breaking changes** None. The wake payload gains a new optional `documentReviewContext` field; the existing plan context field and its legacy builder are unchanged, so plan-only wakes stay byte-for-byte compatible. ## What Changed - Add `buildDocumentReviewContext` in `server/src/services/plan-review-context.ts`: bounded review context (shared thread/comment/character budgets, per-document legacy limits) over all non-plan issue documents - Include `documentReviewContext` in agent wake payloads (`server/src/services/heartbeat.ts`) and in the issue wake-context response (`server/src/routes/issues.ts`) - Add shared `DocumentReviewContext` / `DocumentReviewContextDocument` types in `packages/shared` - Normalize and render the new context in adapter wake-payload text (`packages/adapter-utils/src/server-utils.ts`), with tests - Show a `DocumentAnnotationsCountChip` and the annotation panel for documents in the properties pane Plans and Artifacts tabs, with tests - Extend server document-annotations service tests to cover the new context builder ## Verification - Run `npx vitest run packages/adapter-utils/src/server-utils.test.ts server/src/__tests__/document-annotations-service.test.ts` from the repo root — 104 tests pass - Run `TZ=UTC npx vitest run ui/src/components/issue-properties/IssuePropertiesDocumentAnnotations.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/IssueDocumentAnnotations.test.tsx ui/src/components/DocumentAnnotationPopover.test.tsx` from the repo root — 75 tests pass (one pre-existing monitor-row case asserts UTC timestamps, so use `TZ=UTC` locally; CI runs in UTC) - `pnpm run typecheck` in `server/` passes - Manual: annotate a non-plan issue document, then wake the assigned agent with a comment — the wake payload lists the open document annotation threads; the Artifacts tab shows the annotation count chip and opens the panel ## Risks - The wake payload gains a new optional `documentReviewContext` field; consumers that ignore unknown fields are unaffected, and the plan context field is unchanged - The context is new input to agent wakes; shared budgets (same limits as the plan context) bound token cost across all documents - Low UI risk: the properties-pane changes reuse the existing annotation components > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), with extended thinking and agentic tool use (Claude Code harness) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b1fd42ac1 |
test(grok-local): isolate billing env in usage cost test (#11285)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters report run output, token use, and cost data. > - The Grok local adapter now reports real token use and cost data. > - Its new billing test must prove the no-key path and the API-key path. > - The no-key assertion used the caller environment without isolation. > - This made the test fail when `XAI_API_KEY` was already set. > - This pull request isolates that environment state in the test. > - The benefit is stable coverage for the cost gate from #10433. ## Linked Issues or Issue Description Refs #10433 **What happened?** The Grok local usage and cost test asserted subscription billing while it still used the ambient process environment. If `XAI_API_KEY` was set before the test ran, the adapter selected API billing instead. The subscription assertion could then fail on a developer machine or a CI runner with provider credentials. **Expected behavior** The test should prove the subscription path with no `XAI_API_KEY`. It should also prove the API billing path with a test key. **Steps to reproduce** 1. Start from `master` after #10433. 2. Set `XAI_API_KEY` in the shell environment. 3. Run `vitest` for `packages/adapters/grok-local/src/server/execute.test.ts`. 4. Observe that the subscription half can take the API billing branch without test isolation. **Paperclip version or commit** `master` after #10433. **Deployment mode** Built from source. ## What Changed - Isolated `XAI_API_KEY` with save, delete, set, and restore logic around both billing assertions. - Gave the subscription and API billing checks separate run ids and temp roots. ## Verification - `XAI_API_KEY=ambient-test-key corepack pnpm exec vitest run packages/adapters/grok-local/src/server/execute.test.ts` - `corepack pnpm --filter @paperclipai/adapter-grok-local typecheck` ## Risks Low risk. This changes test setup only. It does not change Grok local adapter runtime behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex local coding agent. The agent used shell tools, GitHub CLI, and local test execution. The context window size was not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
44694328a3 |
fix(issues): make DELETE /api/issues/:id succeed for issues with dependents (#11331)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server provides issue APIs and the database stores issue child rows > - The issue delete endpoint removes the parent issue before dependent rows > - Several issue foreign keys had no delete policy, so PostgreSQL returned a foreign-key error > - This pull request adds safe cascade and set-null policies and a clear conflict response > - The benefit is reliable issue deletion with a useful error when a restricted audit row still blocks deletion ## Linked Issues or Issue Description Fixes #7728 Fixes #4660 Fixes #7991 Fixes #4627 Fixes #5086 **What happened?** `DELETE /api/issues/:id` returned HTTP 500 when dependent comments, thread interactions, read states, inbox archives, feedback votes, or ledger rows referenced the issue. The database raised SQLSTATE 23503 because several foreign keys had no delete policy. **Expected behavior** The endpoint must remove dependent rows that have no meaning without the issue. It must keep ledger rows with a null issue reference. It must return HTTP 409 when a restricted decision audit row still references the issue. **Steps to reproduce** 1. Create an issue. 2. Add a comment or thread interaction that references the issue. 3. Send `DELETE /api/issues/:id`. 4. Observe the HTTP 500 response. **Paperclip version or commit** Commit `1f8f456f8340823fe2bd891ae8933d942f190b7b`. **Deployment mode** Local dev with embedded PGlite or external PostgreSQL. ## What Changed - Add `CASCADE` to five issue child foreign keys. - Add `SET NULL` to the finance and cost event issue foreign keys. - Keep decision audit references restricted. - Map SQLSTATE 23503 from the issue delete service to HTTP 409. - Add migration 0217 for the seven changed tables. - Add regression tests for cascade deletion and restricted decision references. ## Verification - Run `pnpm --filter @paperclipai/db typecheck`. - Run `pnpm --filter @paperclipai/server typecheck`. - Run `npx vitest run src/__tests__/issue-remove-cascade.test.ts` from `server/`. - The regression test applies migration 0217 to a fresh embedded PostgreSQL database. ## Risks - Migration 0217 changes only seven foreign keys that reference `issues.id`. - Cascade deletion removes child rows that cannot exist without the parent issue. - Set-null preserves finance and cost ledger rows. - Decision audit rows remain protected, so the endpoint can return HTTP 409. ## Model Used Codex, based on GPT-5, with tool use and code-review support. The implementation author used an AI coding agent. This PR handoff uses the same model family to validate the commit and manage the pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3cd596725d |
build(deps): bump @modelcontextprotocol/sdk from 1.29.0 to 1.30.0 (#10729)
Bumps [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk) from 1.29.0 to 1.30.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@modelcontextprotocol/sdk's releases</a>.</em></p> <blockquote> <h2>1.30.0</h2> <h2>What's Changed</h2> <ul> <li>fix(server): prioritize zod issues and format them by <a href="https://github.com/mozmo15"><code>@mozmo15</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1503">modelcontextprotocol/typescript-sdk#1503</a></li> <li>chore(ci): switch publish to OIDC trusted publishing by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1839">modelcontextprotocol/typescript-sdk#1839</a></li> <li>Add end-to-end test suite by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2167">modelcontextprotocol/typescript-sdk#2167</a></li> <li>v1 stdio buffer limit by <a href="https://github.com/KKonstantinov"><code>@KKonstantinov</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2239">modelcontextprotocol/typescript-sdk#2239</a></li> <li>fix: support Zod 3.25 method literals by <a href="https://github.com/mattzcarey"><code>@mattzcarey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2368">modelcontextprotocol/typescript-sdk#2368</a></li> <li>Validate Content-Type by parsed media type instead of substring match (v1.x) by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2444">modelcontextprotocol/typescript-sdk#2444</a></li> <li>fix: send SSE keep-alive comment frames from Streamable HTTP server transport (v1.x) by <a href="https://github.com/mattzcarey"><code>@mattzcarey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2538">modelcontextprotocol/typescript-sdk#2538</a></li> <li>fix(deps): widen <code>@hono/node-server</code> past GHSA-frvp-7c67-39w9 by <a href="https://github.com/arimu1"><code>@arimu1</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2549">modelcontextprotocol/typescript-sdk#2549</a></li> <li>Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x) by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2547">modelcontextprotocol/typescript-sdk#2547</a></li> <li>chore: bump version to 1.30.0 by <a href="https://github.com/felixweinberger"><code>@felixweinberger</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2563">modelcontextprotocol/typescript-sdk#2563</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/mozmo15"><code>@mozmo15</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1503">modelcontextprotocol/typescript-sdk#1503</a></li> <li><a href="https://github.com/arimu1"><code>@arimu1</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2549">modelcontextprotocol/typescript-sdk#2549</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/2d889f2b329e46680ec9bdd565de4616c497825a"><code>2d889f2</code></a> chore: bump version to 1.30.0 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2563">#2563</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/e3f3daa12cc2603919939b72136ce9d9e800b868"><code>e3f3daa</code></a> Fix SSE keep-alive timer lifecycle in Streamable HTTP server transport (v1.x)...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/bb5a718cbf90796bacbf62218b359196d210426b"><code>bb5a718</code></a> fix(deps): widen <code>@hono/node-server</code> past GHSA-frvp-7c67-39w9 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2549">#2549</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/1dad2634ce5799fb386283d14291d1b4935a9a52"><code>1dad263</code></a> fix: send SSE keep-alive comment frames from Streamable HTTP server transport...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/69749aa5081ddfe675d36da8d96c7e27d83742b8"><code>69749aa</code></a> Validate Content-Type by parsed media type instead of substring match (v1.x) ...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/369513df7b0e9d8a979c86f68ba1930e0d5f27f0"><code>369513d</code></a> fix: support Zod 3.25 method literals (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2368">#2368</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/e7ee57c2f33b8290a78a3cefa27ab635fe67fbff"><code>e7ee57c</code></a> v1 stdio buffer limit (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2239">#2239</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/c36e1ef5bb3b07b0c23fc6d28d4a6b56ebdd9512"><code>c36e1ef</code></a> Add end-to-end test suite (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2167">#2167</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/bf1e022bd219f678b3865093d58595c6c8a67f1a"><code>bf1e022</code></a> chore(ci): switch publish to OIDC trusted publishing (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1839">#1839</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/9edbab7a09f31a288a27df3220edbebff45dbb6c"><code>9edbab7</code></a> fix(server): prioritize zod issues and format them (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1503">#1503</a>)</li> <li>See full diff in <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/v1.29.0...1.30.0">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new releaser for <code>@modelcontextprotocol/sdk</code> since your current version.</p> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
817225415c |
build(deps-dev): bump rollup from 4.62.2 to 4.62.4 (#11319)
Bumps [rollup](https://github.com/rollup/rollup) from 4.62.2 to 4.62.4. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/releases">rollup's releases</a>.</em></p> <blockquote> <h2>v4.62.4</h2> <h2>4.62.4</h2> <p><em>2026-08-01</em></p> <h3>Bug Fixes</h3> <ul> <li>Resolve a regression when using Rollup on older Linux distributions (<a href="https://redirect.github.com/rollup/rollup/issues/6467">#6467</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6463">#6463</a>: docs: add llms.txt documentation index for LLMs and agents (<a href="https://github.com/abyworkings-coder"><code>@abyworkings-coder</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6464">#6464</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6465">#6465</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6466">#6466</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6467">#6467</a>: ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>v4.62.3</h2> <h2>4.62.3</h2> <p><em>2026-07-26</em></p> <h3>Bug Fixes</h3> <ul> <li>Sanitize illegal characters preserved modules input base (<a href="https://redirect.github.com/rollup/rollup/issues/6439">#6439</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6421">#6421</a>: docs: update x_google_ignoreList link to canonical URL (<a href="https://github.com/DucMinhNe"><code>@DucMinhNe</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6422">#6422</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6423">#6423</a>: chore(deps): update actions/checkout action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6424">#6424</a>: chore(deps): update dependency eslint-plugin-unicorn to v68 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6425">#6425</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6426">#6426</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6432">#6432</a>: fix: make isLegal idempotent by not using a global-flag regex (<a href="https://github.com/spokodev"><code>@spokodev</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6433">#6433</a>: docs: clarify sideEffects and moduleSideEffects (<a href="https://github.com/ishaanlabs-gg"><code>@ishaanlabs-gg</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6434">#6434</a>: chore(deps): update dtolnay/rust-toolchain digest to 4be7066 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6435">#6435</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6436">#6436</a>: chore(deps): update actions/cache action to v6 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6438">#6438</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6439">#6439</a>: Sanitize input base before computing preserved module chunk names (<a href="https://github.com/MahinAnowar"><code>@MahinAnowar</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6443">#6443</a>: chore(deps): update dependency eslint-plugin-unicorn to v71 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6444">#6444</a>: fix(deps): update rust crate swc_compiler_base to v60 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6446">#6446</a>: chore(deps): update dtolnay/rust-toolchain digest to 4cda84d (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6447">#6447</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6448">#6448</a>: chore(deps): update actions/setup-node action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6449">#6449</a>: chore(deps): update dependency eslint-plugin-unicorn to v72 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6450">#6450</a>: chore(deps): update dependency pinia to v4 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6451">#6451</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6455">#6455</a>: docs: fix broken commonjs namedExports link in troubleshooting (<a href="https://github.com/Hashim1999164"><code>@Hashim1999164</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/rollup/rollup/blob/master/CHANGELOG.md">rollup's changelog</a>.</em></p> <blockquote> <h2>4.62.4</h2> <p><em>2026-08-01</em></p> <h3>Bug Fixes</h3> <ul> <li>Resolve a regression when using Rollup on older Linux distributions (<a href="https://redirect.github.com/rollup/rollup/issues/6467">#6467</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6463">#6463</a>: docs: add llms.txt documentation index for LLMs and agents (<a href="https://github.com/abyworkings-coder"><code>@abyworkings-coder</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6464">#6464</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6465">#6465</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6466">#6466</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6467">#6467</a>: ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <h2>4.62.3</h2> <p><em>2026-07-26</em></p> <h3>Bug Fixes</h3> <ul> <li>Sanitize illegal characters preserved modules input base (<a href="https://redirect.github.com/rollup/rollup/issues/6439">#6439</a>)</li> </ul> <h3>Pull Requests</h3> <ul> <li><a href="https://redirect.github.com/rollup/rollup/pull/6421">#6421</a>: docs: update x_google_ignoreList link to canonical URL (<a href="https://github.com/DucMinhNe"><code>@DucMinhNe</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6422">#6422</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6423">#6423</a>: chore(deps): update actions/checkout action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6424">#6424</a>: chore(deps): update dependency eslint-plugin-unicorn to v68 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6425">#6425</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6426">#6426</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6432">#6432</a>: fix: make isLegal idempotent by not using a global-flag regex (<a href="https://github.com/spokodev"><code>@spokodev</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6433">#6433</a>: docs: clarify sideEffects and moduleSideEffects (<a href="https://github.com/ishaanlabs-gg"><code>@ishaanlabs-gg</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6434">#6434</a>: chore(deps): update dtolnay/rust-toolchain digest to 4be7066 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6435">#6435</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6436">#6436</a>: chore(deps): update actions/cache action to v6 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6438">#6438</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6439">#6439</a>: Sanitize input base before computing preserved module chunk names (<a href="https://github.com/MahinAnowar"><code>@MahinAnowar</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6443">#6443</a>: chore(deps): update dependency eslint-plugin-unicorn to v71 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6444">#6444</a>: fix(deps): update rust crate swc_compiler_base to v60 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6446">#6446</a>: chore(deps): update dtolnay/rust-toolchain digest to 4cda84d (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6447">#6447</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6448">#6448</a>: chore(deps): update actions/setup-node action to v7 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6449">#6449</a>: chore(deps): update dependency eslint-plugin-unicorn to v72 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6450">#6450</a>: chore(deps): update dependency pinia to v4 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6451">#6451</a>: chore(deps): lock file maintenance (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6455">#6455</a>: docs: fix broken commonjs namedExports link in troubleshooting (<a href="https://github.com/Hashim1999164"><code>@Hashim1999164</code></a>, <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6456">#6456</a>: fix(deps): update minor/patch updates (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot])</li> <li><a href="https://redirect.github.com/rollup/rollup/pull/6457">#6457</a>: chore(deps): update dependency magic-string to v1 (<a href="https://github.com/renovate"><code>@renovate</code></a>[bot], <a href="https://github.com/lukastaegert"><code>@lukastaegert</code></a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rollup/rollup/commit/ddc4ffab628944e45dbb8d66d58aae818015440f"><code>ddc4ffa</code></a> 4.62.4</li> <li><a href="https://github.com/rollup/rollup/commit/86d171076b2855c2036cac1f17eb94eea2b9cd7a"><code>86d1710</code></a> Update audit resolve</li> <li><a href="https://github.com/rollup/rollup/commit/7beedfa94963a32c14116994d1bae48d33673245"><code>7beedfa</code></a> ci: fix linux-gnu glibc regression and enforce glibc ≤ 2.28 compatibility (<a href="https://redirect.github.com/rollup/rollup/issues/6">#6</a>...</li> <li><a href="https://github.com/rollup/rollup/commit/9c2c58d55632d25a56cc72f4ae9d43d66ef54040"><code>9c2c58d</code></a> docs: add llms.txt documentation index for LLMs and agents (<a href="https://redirect.github.com/rollup/rollup/issues/6463">#6463</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/dc692883d8c7575692613b76a0d0d45afcf1a4ef"><code>dc69288</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/rollup/rollup/issues/6466">#6466</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/5ee08215eaafc3ea495a76fb317ec89aa15c0690"><code>5ee0821</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/rollup/rollup/issues/6465">#6465</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/4501389a63d588ce05f141b6452359e7a510a496"><code>4501389</code></a> fix(deps): update minor/patch updates (<a href="https://redirect.github.com/rollup/rollup/issues/6464">#6464</a>)</li> <li><a href="https://github.com/rollup/rollup/commit/a80a1974c584bfa8b694fb5d1a20f3fa75ebaf0a"><code>a80a197</code></a> 4.62.3</li> <li><a href="https://github.com/rollup/rollup/commit/e87e19b31e87a1dd6a749a6c11afe4cb9183cf8d"><code>e87e19b</code></a> Update audit resolve</li> <li><a href="https://github.com/rollup/rollup/commit/72f98e99228ad90a8e13d79709c009a827a4ef34"><code>72f98e9</code></a> Fix build:docs after rollup update (<a href="https://redirect.github.com/rollup/rollup/issues/6460">#6460</a>)</li> <li>Additional commits viewable in <a href="https://github.com/rollup/rollup/compare/v4.62.2...v4.62.4">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
3040db3343 |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.63.0 to 0.66.0 (#11314)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.63.0 to 0.66.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.66.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a> (2026-08-07)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b">7f27c47</a>)</li> <li>expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c">8b31dea</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b">f8fd3ab</a>)</li> </ul> <h2>v0.65.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a> (2026-08-05)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9">b965dd2</a>)</li> <li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15b4eb46f329566837eae58f2ee4b05e3e81bf64">15b4eb4</a>)</li> <li><strong>deps:</strong> Bump <code>@hono/node-server</code> from 1.19.17 to 2.1.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f9123f3e18560b580398aabf49e2190f69746976">f9123f3</a>)</li> <li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/098843842895dcb450746bd064dbb1509e3049d1">0988438</a>)</li> <li><strong>steering:</strong> settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d">a84b810</a>)</li> </ul> <h2>v0.64.2</h2> <h2>Bug Fixes</h2> <ul> <li>restore the single-tool representation for ExitPlanMode (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>) (4302a4b)</li> </ul> <h2>v0.64.1</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0936ec281ec730714c605e3da732069ff47d8969">0936ec2</a>)</li> </ul> <h2>v0.64.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a> (2026-07-30)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8e099e844254c3e91508c79a02e3e7dc2239fcbb">8e099e8</a>)</li> <li><strong>deps:</strong> Bump the minor group with 7 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3f609219592e63b947539f79c696b3cedb421060">3f60921</a>)</li> </ul> <h3>Bug Fixes</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.65.0...v0.66.0">0.66.0</a> (2026-08-07)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b">7f27c47</a>)</li> <li>expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c">8b31dea</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b">f8fd3ab</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.2...v0.65.0">0.65.0</a> (2026-08-05)</h2> <h3>Features</h3> <ul> <li><strong>deps-dev:</strong> Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9">b965dd2</a>)</li> <li><strong>deps-dev:</strong> Bump tinyexec from 1.2.4 to 1.3.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/959">#959</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/15b4eb46f329566837eae58f2ee4b05e3e81bf64">15b4eb4</a>)</li> <li><strong>deps:</strong> Bump <code>@hono/node-server</code> from 1.19.17 to 2.1.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/956">#956</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f9123f3e18560b580398aabf49e2190f69746976">f9123f3</a>)</li> <li><strong>deps:</strong> Bump fast-uri from 3.1.4 to 3.1.5 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/952">#952</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/098843842895dcb450746bd064dbb1509e3049d1">0988438</a>)</li> <li><strong>steering:</strong> settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d">a84b810</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.1...v0.64.2">0.64.2</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>restore the single-tool representation for ExitPlanMode (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/942">#942</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/4302a4b0b6df821b164cbe4857f26cf5b44b532c">4302a4b</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.64.0...v0.64.1">0.64.1</a> (2026-08-02)</h2> <h3>Bug Fixes</h3> <ul> <li>release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/939">#939</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0936ec281ec730714c605e3da732069ff47d8969">0936ec2</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.64.0">0.64.0</a> (2026-07-30)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Bump actions/checkout from 7.0.0 to 7.0.1 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/925">#925</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8e099e844254c3e91508c79a02e3e7dc2239fcbb">8e099e8</a>)</li> <li><strong>deps:</strong> Bump the minor group with 7 updates (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/928">#928</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3f609219592e63b947539f79c696b3cedb421060">3f60921</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li><strong>steering:</strong> add opt-in host-owned fallback (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/919">#919</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/43af4ec29ea5396c2614813af05967bfb0b1bac8">43af4ec</a>), closes <a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/903">#903</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/6b405138fc82be947964612fac04e56654827b66"><code>6b40513</code></a> chore(main): release 0.66.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/961">#961</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8aaf608b4e5a9bead4dd3a4abb060c916281dca6"><code>8aaf608</code></a> ci: fix release flow (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/971">#971</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/f8fd3ab8224420f8ced570e974cde09612939d6b"><code>f8fd3ab</code></a> fix: publish and replace Claude goals reliably (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/967">#967</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/133337ffe5f1137305fa18befafa06f483db32b1"><code>133337f</code></a> ci: simplify the release flow and make it agent-friendly (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/965">#965</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8b31dea11bed54f86c41217759159c415611346c"><code>8b31dea</code></a> feat: expose provider-neutral ACP goal extension (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/964">#964</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/bba912728f7acde6a415b70bf4e6d3b6be99947d"><code>bba9127</code></a> ci: validate PR titles against release-please conventions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/962">#962</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/7f27c47c5c7c49e65014e9f7dc55cba17352d33b"><code>7f27c47</code></a> feat(deps-dev): Bump globals from 17.8.0 to 17.9.0 in the minor group (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/960">#960</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/6d608cb399001329b5f485d750e1114ce7293439"><code>6d608cb</code></a> chore(main): release 0.65.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/957">#957</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/a84b81080a4127edf40bc448fc8bf2b15503304d"><code>a84b810</code></a> feat(steering): settle a steered turn at idle, not at the interrupt (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/958">#958</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/b965dd21917e822b56f5012c3572902f26c065c9"><code>b965dd2</code></a> feat(deps-dev): Bump nanoid from 3.3.16 to 3.3.17 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/951">#951</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.63.0...v0.66.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d68cf32ae9 |
build(deps): bump @codemirror/view from 6.43.1 to 6.43.8 (#11321)
Bumps [@codemirror/view](https://github.com/codemirror/view) from 6.43.1 to 6.43.8. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/codemirror/view/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f0e6c0f549 |
feat(server): receive and apply the Paperclip Cloud onboarding seed (#11098)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud provisions a dedicated tenant stack for each
customer. During signup it asks for a mission, a name and role for the
first agent, and a first task.
> - Cloud pushes those answers into the new stack at activation, as
`POST /api/companies/:companyId/onboarding-seed`.
> - No route served that path. The tenant answered 404, so Cloud
recorded the push as unacknowledged and retried on every portfolio
fetch.
> - The failure was soft. The answers stayed durable in Cloud and the
stack still activated. But the stack opened on the empty first-run
wizard, and it asked the customer again for what they had already given.
> - This pull request adds the receiving endpoint. It validates the
seed, applies it, and acknowledges it.
> - The benefit is that a seeded stack opens with the mission, the agent
and the first task already in place.
## Linked Issues or Issue Description
No public GitHub issue covers this. The problem is described in-PR,
following the feature template.
**Subsystem affected**
server/ — Express REST API and orchestration services. Also
`packages/db` (one new table) and `packages/shared` (one new validator).
**Problem or motivation**
Paperclip Cloud collects onboarding answers at signup and pushes them to
the tenant stack at activation. The tenant had no route for that
request. It answered 404. Cloud treats a non-2xx as "not yet applied",
so it kept the answers and retried, but the stack itself stayed
unseeded. A customer who had already named their mission, their first
agent and their first task arrived at an empty first-run wizard that
asked for all three again.
**Proposed solution**
Serve `POST /api/companies/:companyId/onboarding-seed`. Validate the
body, apply it to the company, then acknowledge it.
The seed is customer free text, so it is bounded and validated in
`packages/shared` and read from the JSON body only. It is never read
from an `x-paperclip-cloud-*` header. That header set is the trusted
identity envelope: every member is derived server-side from the host
plus verified domain records, and that is exactly what makes it
trustworthy. Mixing user content into it would remove the property. A
test plants a mission on a cloud header and asserts that the body value
wins.
Application reuses the shapes the first-run wizard already produces, so
a seeded stack and a manually onboarded one look the same afterwards:
- The mission becomes the company-level goal. A multi-line mission
splits into a title and a description, as the wizard does.
- The agent becomes the company's first hire. Its free-text role ("Chief
of Staff") lands on `title`. The structural `role` stays `ceo`, which is
what the org chart and the default-instructions lookup read.
- The first task becomes an issue in the Onboarding project, assigned to
that agent.
Cloud retries until it gets a 2xx, and it reads any 2xx as "the tenant
holds this content". So the endpoint is idempotent per `revision`. A new
`company_onboarding_seeds` table records the applied revision together
with the goal, the agent and the issue it produced. A replay of a
revision that already matches is a successful no-op. A later revision —
the customer edited their answers — updates those three rows in place
instead of creating a second agent and a second task. The record is
written last, after every other write has landed, so a partial
application cannot present itself as acknowledged.
Everything is applied before the 200 is sent. This is an ordering
guarantee, not eventual consistency. The tests read the database
immediately after the response, with no waiting and no polling, so a
lazy receiver fails them on a fast machine as well as a slow one. That
matters because the redirect into the tenant dashboard is gated on this
acknowledgement.
**Alternatives considered**
Store the seed and let the tenant UI apply it on first load. Rejected:
the dashboard redirect is gated on the acknowledgement, so a background
apply would let the dashboard open before the agent and the task exist.
The whole point is that it must not.
Reuse `POST /companies/:companyId/agents` and `POST
/companies/:companyId/issues` over HTTP from Cloud. Rejected: it needs
three round trips with no shared idempotency key, and it moves the "did
all of it land?" decision to the caller.
**Roadmap alignment**
This completes an existing Cloud-to-tenant contract. It does not add a
new user-facing surface.
## What Changed
- Add `POST /api/companies/:companyId/onboarding-seed` in
`server/src/routes/onboarding-seed.ts`. It authenticates exactly as
`POST /api/companies/:companyId/logo` does, through
`assertCompanyAccess`.
- Add `server/src/services/onboarding-seed.ts`. It applies the mission,
the agent and the first task, and records the applied revision last.
- Add the `company_onboarding_seeds` table: schema, migration `0216`,
and journal entry. It holds the applied revision and the ids of the
goal, agent and issue the seed produced.
- Add `applyOnboardingSeedSchema` in `packages/shared`. It bounds
mission to 2000, agent name to 80, agent role to 120, task title to 200,
and task details to 2000 — the same limits Cloud enforces before it
sends.
- Mount the router in `server/src/app.ts` and register the path in the
OpenAPI document.
- Add `server/src/__tests__/onboarding-seed-route.test.ts` with 13
tests.
- The seeded agent is created on `claude_local`. This mirrors the
teams-catalog default for agents created server-side, where no human
runs an environment test first. `PAPERCLIP_ONBOARDING_SEED_ADAPTER_TYPE`
overrides it.
## Verification
```sh
pnpm typecheck # whole workspace, passes
npx vitest run \
server/src/__tests__/onboarding-seed-route.test.ts \
server/src/__tests__/openapi-routes.test.ts # 15 passed
```
The suite runs against embedded Postgres with migrations applied, so
migration `0216` is exercised by every test.
The route tests cover:
- the happy path — mission, agent and task all applied, read immediately
after the 200
- replay of the same revision — no second agent, no second task, no
second goal, no second project
- a later revision — the goal, agent and task are updated in place
- a multi-line mission splitting into a goal title and description
- a revision-only seed
- the activity log entry written once, and not again on a replay
- a caller without access to the company — 403, and nothing written
- a body with no revision — 400
- each field bound past its limit — 400
- a mission planted on an `x-paperclip-cloud-*` header — ignored, body
wins
- an existing Onboarding project — reused, not duplicated
Not verified here: the full Cloud-to-tenant walk against a live stack.
That needs a deployed Cloud and a provisioned tenant together, which is
separate staging work.
## Risks
Migration `0216` creates one new table. It adds no column to an existing
table, rewrites nothing, and backfills nothing, so it is safe to apply
online. The migration safety check passes.
The endpoint writes to a company. Access is enforced by
`assertCompanyAccess`, the same gate the company logo write uses, and a
test covers the denial.
Behavioral note for stacks that already hold data. If a company already
has a non-built-in `ceo` agent, a first seed updates that agent's name
and title rather than creating a second lead. Likewise a seed adopts an
existing company-level goal rather than adding a parallel one. This is
deliberate: the seed is the customer's own stated answer from signup,
and two competing missions or two leads would be worse than one updated
in place. In the intended case — a stack that Cloud has just activated —
none of these exist yet.
The seeded agent is created on `claude_local` with an empty adapter
config. It is idle and needs the usual credential setup before it runs.
Seeding it does not start it.
## Update — rebased onto master + review hardening
Master moved on after this PR was cut, so it was **rebased onto
`master`** and
the seed migration was **renumbered from `0212` to `0216`** (the merged
#11101
took `0212_onboarding_first_task_unique`); the drizzle journal was
re-stitched
and `check:migrations` passes.
Two things landed on top of the original receiver:
- **Mission-only walk contract (PAP-67 r17.4).** The tenant now owns the
first
agent and the first task via #11101's server-owned onboarding path,
which
stamps `ONBOARDING_FIRST_TASK_ORIGIN_KIND` and races safely on the
partial
unique index `issues_onboarding_first_task_uq`. A comment in the apply
path
documents why this receiver leaves the first task to that path on the
cloud
walk, and a paperclip-cloud `node:test`
(`src/onboarding/walk-seed.test.ts`)
asserts the walk's seed carries no `agent`/`firstTask`. The receiver
retains
the agent/first-task code for its documented body contract, kept inert
on the
cloud path by the mission-only seed.
- **Three Greptile P1 fixes** (`95622fa37`): concurrent application is
now
serialized under a per-company `pg_advisory_xact_lock` (no duplicate
goal/agent/project/task on overlapping pushes); a revised first task
carries
its resolved `assigneeAgentId`/`goalId`; and the
`company.onboarding_seed_applied`
audit write is best-effort so a logging failure can't leave the entry
permanently absent. Two new regression tests cover the first two.
## Model Used
Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking,
with tool use and code execution. Used for the original codebase
investigation, the implementation, and the tests. The rebase, migration
renumber, mission-only contract, and the three P1 fixes were done with
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use
and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1e07d5b9aa |
fix(db): give the last two embedded-Postgres migration tests a timeout (#11313)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `@paperclipai/db` package owns the database schema and its migrations > - Some migration tests start an embedded Postgres server and replay a migration against it > - An embedded Postgres server needs 7 to 12 seconds to start on a CI runner > - Vitest stops a test after 5 seconds unless the test sets its own timeout > - Two of these tests do not set a timeout, so they fail on CI before they assert anything > - This pull request gives both tests a 30 second timeout > - The benefit is that unrelated pull requests stop failing on a test they did not change ## Linked Issues or Issue Description No public issue exists for this. The problem follows. **What happened?** The test `packages/db/src/company-secret-proposals-migration.test.ts` fails on CI. The error is `Test timed out in 5000ms`. The test never reaches its assertions. The suite reports `1 failed | 104 passed`. The failure is not caused by the branch under test. It appeared on three different branches in a few hours: | Run | Head | Failing jobs | | --- | --- | --- | | 31630781317 | `95622fa3` | `General tests (workspaces-b)`, `verify`, `e2e shard (2/3)`, `e2e` | | 31652020976 | `feba90c9` | `General tests (workspaces-b)`, `verify`, `e2e shard (3/3)`, `e2e` | | 31651467721 | `f2115207` | `General tests (workspaces-a (1/2))`, `verify` | The `verify` job reads the result of the general tests. One timeout therefore turns into two red checks. A reviewer sees two failures and reads them as a regression. **Expected behavior** The test starts an embedded Postgres server, replays the migration, and asserts the schema. It must pass on a normal CI runner. **Steps to reproduce** 1. Open any pull request against `master`. 2. Wait for the job `General tests (workspaces-b)`. 3. Read the failure. The test times out after 5000 ms. The failure needs a slow runner. A fast development machine starts embedded Postgres in less than 5 seconds, so the test passes there. **Paperclip version or commit** `master` at `a09d7dcc0`. **Deployment mode** CI only. GitHub Actions, `ubuntu24` runner image. ## What Changed - `packages/db/src/company-secret-proposals-migration.test.ts` — the test now uses a 30 second timeout. The migration suites in this package already use 20 to 60 seconds. 30 seconds is the most common value. - `packages/db/src/status-card-migrations.test.ts` — the same change. This test has the same defect. It does not fail yet because it replays fewer statements. A fix to only one test moves the problem instead of removing it. - Both tests get a comment. The comment tells the next author why the 5 second default is too short. These two tests were the only embedded-Postgres migration tests in the package without a timeout. ## Verification - Run `pnpm vitest run src/company-secret-proposals-migration.test.ts src/status-card-migrations.test.ts` in `packages/db`. Both tests pass. - These suites skip themselves when the Postgres binaries are absent. A pass alone therefore proves nothing. Run the command with `--reporter=verbose`. The output contains Postgres `NOTICE` messages, for example `relation "status_cards" already exists, skipping`. These messages prove the tests ran real SQL. - Run the same command with `--testTimeout=1`. Both tests still pass. This proves the per-test timeout overrides the global timeout. Before this change, the same command fails immediately. - All CI jobs on this pull request pass. The job `General tests (workspaces-b)` passes. This job failed on the three runs listed above. Not done: no attempt to reproduce the timeout on a development machine. A fast machine starts embedded Postgres in less than 5 seconds, so the failure does not occur there. ## Risks Low risk. The change adds two timeout arguments to tests. It changes no source code, no schema, and no dependency. A longer timeout cannot hide a regression here. The tests assert the same conditions as before. A migration that truly hangs now fails after 30 seconds. Before, it failed after 5 seconds with a message that pointed at the wrong cause. The `e2e` failures on the runs above have a different cause. The spec `mcp-user-stories.spec.ts › US-9` fails with `502 — fetch failed` and `fetch failed: bad port`. These errors come from MCP tool-connection health checks. The failures hit different shards on different runs. This pull request does not change that behavior. `e2e shard (2/3)` passes here, which supports the view that those failures are unstable infrastructure. To revert, remove the two timeout arguments. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell command execution for the local test runs, and the GitHub CLI to read the failing CI logs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b7b8fbf688 |
fix(adapter-utils): let explicit PAPERCLIP_API_URL override the derived runtime URL in run env (#10339)
## Thinking Path
> - Paperclip orchestrates AI agents for zero-human companies
> - Every agent run gets a run-scoped bridge into the Paperclip API
through the injected `PAPERCLIP_API_URL` / `PAPERCLIP_API_KEY` env vars,
built by `buildPaperclipEnv` in
`packages/adapter-utils/src/server-utils.ts`
> - `buildPaperclipEnv` resolves that URL as `PAPERCLIP_RUNTIME_API_URL
?? PAPERCLIP_API_URL ?? http://<listen-host>:<port>`, and the server
always exports `PAPERCLIP_RUNTIME_API_URL` derived from
`authPublicBaseUrl` at boot
> - When `authPublicBaseUrl` points at an address that is not reachable
from inside the runtime container (e.g. a VPN/tailnet-only address used
to keep the web UI off the public internet), every local run receives a
dead API URL (`curl` exit 7) and agents only survive by hand-rolling a
localhost fallback
> - An operator-set `PAPERCLIP_API_URL` is the documented escape hatch —
`docs/deploy/environment-variables.md` states the server "preserves the
value" when set externally and that the run-level var "inherits the
server-level value" — but the run env builder inverts the precedence, so
the override never actually reaches runs
> - This pull request swaps the precedence in `buildPaperclipEnv` so an
explicit `PAPERCLIP_API_URL` wins over the derived runtime URL, aligning
the behavior with the documented contract
> - The benefit is that operators with split-horizon topologies (public
auth URL != container-reachable URL) can point agent runs at a reachable
endpoint with one env var, with zero behavior change for deployments
that do not set it
## Underlying Issue
No pre-existing public issue covers this, so per CONTRIBUTING ("Link
Issues or Describe Them In-PR") here are the `bug_report.yml` fields
inline:
- **What happened:** with `PAPERCLIP_AUTH_PUBLIC_BASE_URL` on a
tailnet-only address and `PAPERCLIP_API_URL=http://localhost:3100`
explicitly set in the server environment, every agent run still received
`PAPERCLIP_API_URL=http://100.x.y.z:3100` (the derived,
container-unreachable URL); `curl` from inside the run exits 7 and
agents can only reach the API by hand-rolling a localhost fallback
- **Expected behavior:** the run env inherits the operator-configured
`PAPERCLIP_API_URL`, as documented in
`docs/deploy/environment-variables.md` ("preserves the value", run-level
var "inherits the server-level value")
- **Steps to reproduce:** (1) set `PAPERCLIP_AUTH_PUBLIC_BASE_URL` to an
address not reachable from inside the server container, (2) set
`PAPERCLIP_API_URL=http://localhost:3100` in the server env, (3) trigger
any agent run and inspect the spawned process env: it carries the
derived URL, not the override
- **Version/commit:** reproduced on the `91e58acb` image (2026-07-19);
the precedence is unchanged on current `master` (`a3b293e`)
- **Deployment mode:** single-host Docker Compose, local adapters
(`claude_local`/`codex_local`), web UI exposed via VPN/tailnet only
## Related PRs (dedup search)
Several in-flight PRs touch the same pain point (runs receiving an
unreachable injected API URL) — linked for reviewer context; none of
them honors the documented explicit override, and the older ones appear
stale:
- #9916 — reworks `PAPERCLIP_RUNTIME_API_URL` derivation and port
preservation (server side); complementary, does not change run-env
precedence
- #8130 — honors a pre-set `PAPERCLIP_RUNTIME_API_URL` (server side); a
complementary escape hatch via the runtime var instead of the documented
`PAPERCLIP_API_URL` override
- #8025 — heuristic: prefer loopback when the runtime bind is loopback
(no activity since Jun 12)
- #5692 — heuristic loopback-safe URL inside `buildPaperclipEnv` (no
activity since May 14)
- #4877 — broader same-host injection rework across 10 files (no
activity since May 2)
- #4794 — always forces loopback for spawned agents (no activity since
Apr 30; would break split-horizon setups where a reachable non-loopback
URL is intended)
This PR intentionally takes the Path-1 route from CONTRIBUTING: the
smallest possible change (swap two lines so the documented operator
override wins) plus regression tests, rather than a new heuristic.
## What Changed
- `packages/adapter-utils/src/server-utils.ts`: `buildPaperclipEnv` now
resolves the injected URL as `PAPERCLIP_API_URL ??
PAPERCLIP_RUNTIME_API_URL ?? http://<listen-host>:<port>` (explicit
override first), with a short comment explaining why
- `packages/adapter-utils/src/server-utils.test.ts`: three new tests
covering the override precedence, the derived-URL fallback, and the
listen-host default (including the `0.0.0.0` to `localhost` mapping)
- `server/src/__tests__/paperclip-env.test.ts`: updated the expectation
that encoded the old runtime-URL-first precedence and added the
symmetric fallback case (runtime URL used when no explicit override is
set)
- No docs changes needed: `docs/deploy/environment-variables.md` already
describes the fixed behavior
## Verification
- `vitest run` on the new `buildPaperclipEnv` tests in
`packages/adapter-utils`: 3/3 pass
- `vitest run` on `server/src/__tests__/paperclip-env.test.ts` after the
expectation update: 5/5 pass (the first CI run correctly flagged the one
test that encoded the old precedence)
- Reproduced and verified on a production deployment (single-host
Docker, `PAPERCLIP_AUTH_PUBLIC_BASE_URL` on a tailnet-only address):
- Before: freshly spawned runs received
`PAPERCLIP_API_URL=http://100.x.y.z:3100` (verified in the spawned
process `/proc/<pid>/environ`); `curl` to it from inside the container
exits 7
- After (with `PAPERCLIP_API_URL=http://localhost:3100` in the compose
environment): a fresh run received `http://localhost:3100`, and `curl
$PAPERCLIP_API_URL/api/agents/me` with the run-scoped key returned HTTP
200; the run finished `succeeded` with usage telemetry recorded
## Risks
- Low. Behavior changes only for deployments that explicitly set
`PAPERCLIP_API_URL`; when unset (the default),
`PAPERCLIP_RUNTIME_API_URL` is used exactly as before
- The sandbox callback bridge (`execution-target.ts`) is intentionally
untouched: remote sandboxes genuinely need the publicly reachable URL,
and its `input.hostApiUrl || PAPERCLIP_RUNTIME_API_URL || ...` chain
still provides it
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended
thinking + agentic tool use via Claude Code, operating over SSH against
the affected deployment
## Checklist
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Sergio-LPA <204395363+Sergio-LPA@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
4660562fde |
fix(opencode-local): make the model-availability probe non-fatal (#10294)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents run through adapters; the `opencode-local` adapter shells out to the OpenCode CLI and, before each run, does a pre-flight `opencode models` **availability probe** to fail fast on a misconfigured `provider/model`. > - That probe was written to **throw on any probe failure** — a timeout, a non-zero exit, or a transient `Unexpected error` from the CLI — which aborts the whole heartbeat run. > - In practice the CLI probe fails transiently (provider hiccup, cold cache, momentary CLI error). When that happens *after* the agent has already done its work, the run dies before its terminal disposition is written, so the platform reopens the issue and re-runs it — a spurious crash/re-run loop that affects every agent on the OpenCode adapter. > - This PR makes the probe **non-fatal when it cannot run**: it warns and proceeds with the configured model, letting the real invocation be authoritative. > - It deliberately **keeps** the genuine guard: when the probe *succeeds* and the configured model is absent from a non-empty list, it still throws (this is what catches misconfigured slugs). > - The benefit is that a best-effort pre-flight check can no longer take down an otherwise-healthy run, while the useful misconfiguration guard is retained. ## Linked Issues or Issue Description No public GitHub issue exists; describing inline (bug). **What happened:** an OpenCode-adapter agent run terminated at the adapter level with `` `opencode models` failed: Unexpected error ``. The failure landed after the agent had produced its work, so the terminal-status update never applied and the run was reopened and re-executed. **Expected:** a transient failure of the `opencode models` availability *probe* should not abort the run — the probe is a best-effort pre-flight guard, not a gate. **Actual:** the probe threw on timeout / non-zero exit / empty output, aborting the run and discarding the completed work + disposition. **Scope:** both the local (`models.ts`) and remote/SSH (`execute.ts`) probe paths; affects any agent on the `opencode_local` adapter. Related PRs (context / prior art): - Refs #5119 — added the remote execution-target model-probe validation this PR softens. - Refs #3291 — closed prior attempt to make the `opencode_local` model probe non-blocking (at agent-create time; different entry point). - Refs #8014 — related open work raising the probe timeout (20s → 60s); complementary, not overlapping. ## What Changed - `models.ts` (`ensureOpenCodeModelConfiguredAndAvailable`): if discovery throws (probe can't run) or returns an empty list, **warn and proceed** with the configured model instead of throwing. The "model present in a non-empty list" check is unchanged and still throws when the configured model is genuinely absent. - `execute.ts` (`ensureRemoteOpenCodeModelConfiguredAndAvailable`): remote probe **timeout / non-zero exit / empty output** now warn and return (proceed) instead of throwing. The remote model-absent guard still throws. - `models.test.ts`: the local "discovery cannot run" case now asserts the probe **proceeds** with the configured model (was: asserts it rejects). - `execute.test.ts`: added remote regression tests — non-zero exit, timeout, and empty output all proceed; a successful probe missing the configured model still rejects. ## Verification ```bash pnpm --filter @paperclipai/adapter-opencode-local typecheck # clean # opencode-local server suite (default 5s per-test timeout is too tight for the # heavy SSH tests on some machines; use a realistic timeout): node node_modules/.pnpm/vitest@*/node_modules/vitest/vitest.mjs run \ packages/adapters/opencode-local/src/server/models.test.ts \ packages/adapters/opencode-local/src/server/execute.test.ts \ packages/adapters/opencode-local/src/server/execute.remote.test.ts \ --testTimeout=45000 ``` Result: typecheck clean; all opencode-local server tests pass, including the new remote fail-open tests and the retained "model unavailable on the remote target" guard test. ## Risks - **Fail-open behavior (intentional).** When the probe can't run, a genuinely misconfigured model is no longer caught at pre-flight — it surfaces at the real invocation instead. This is the accepted tradeoff: the probe is best-effort, and the real invocation is authoritative. The high-value guard (probe succeeds + model absent from a non-empty list) is retained, so the common misconfiguration — a bad `provider/model` slug — is still caught. - No API, schema, or migration changes. Behavior change is confined to the two probe helpers. Low risk overall. ## Model Used Anthropic **Claude Opus 4.8** (`claude-opus-4-8`), used via Claude Code with agentic tool use (repo search, file editing, shell/code execution) and extended reasoning. Used to diagnose the crash, implement the fix, and write the tests; the change was reviewed before submission. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change (`fix/opencode-model-probe-non-fatal`) and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — N/A (internal adapter behavior; no user-facing docs affected) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green (functional gates: tests/build/e2e/typecheck/security). Review/Greptile gate re-running after this update. - [ ] Greptile is 5/5 with no open P2s — re-triggered after addressing both P2s (remote test coverage + this template-complete description) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2f1c0e011e |
fix(hermes): surface silent nonzero exit failures (#10107)
## Thinking Path - Followed a silent nonzero Hermes exit from child-process result parsing through heartbeat run, runtime, task-session, and agent finalization. - Found two gaps: the adapter could return `errorMessage: null` for a numeric nonzero exit, and heartbeat later reused the nullable adapter field instead of its normalized fallback. - Kept timeout, signal-cancellation, and specific parsed diagnostics authoritative. ## Linked Issue(s) / Bug Report Related to #9751 (stderr classification) and #9519 (exit-zero finalization), but this is a separate failure mode. Reproduction: run Hermes with a child result equivalent to `exitCode: 1`, `timedOut: false`, and no parsed diagnostic. The heartbeat row derives `Adapter failed`, while runtime/task-session/agent finalization can persist null diagnostics. ## What Changed - Give silent numeric nonzero Hermes exits a stable fallback such as `Hermes exited with code 1`. - Preserve specific parsed errors and timeout/signal semantics. - Reuse the normalized persisted run error for recovered runtime state, task-session `lastError`, and agent `errorReason`. - Add adapter-level and embedded-Postgres regressions. ## Verification - Hermes adapter `execute.onspawn.test.ts` — 7 passed. - Focused heartbeat normalized-error regression — 1 passed (91 skipped). - `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check origin/master...HEAD` — passed. Independent review also ran the full recovery file: the changed regression passed; one unrelated pre-existing timing-sensitive test timed out. ## Risks / Rollout Notes Low risk. Fallback text is used only when a numeric nonzero exit has no better diagnostic. Existing timeout, signal, and parsed-error precedence remains unchanged. ## Model Used OpenAI Codex `gpt-5.6-sol` with repository inspection, test execution, and independent read-only review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (not applicable: internal diagnostics only) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: cucurigoo <cucurigoo@users.noreply.github.com> |
||
|
|
8a5c0615f9 |
fix(adapter-utils): forward sandbox callback bridge traffic to the local listen origin (#10017)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can execute in remote sandboxes, where a callback bridge relays in-sandbox Paperclip API calls back to the host server process > - The bridge worker resolves its forward target from PAPERCLIP_RUNTIME_API_URL / PAPERCLIP_API_URL, which now prefer a configured public base URL and therefore mean "the origin browsers and external agents use" > - The bridge worker runs inside the same process that serves the API, so forwarding through the public origin routes an in-process loopback hop through the network edge > - On a deployment whose public origin sits behind a session-gated edge proxy, every forwarded agent API call is rejected at the edge, so agents in sandboxes cannot read their identity, comment, or hire > - This pull request resolves the bridge forward target from the explicit hostApiUrl override or the local listen host and port only, never the public URL exports > - The benefit is that sandbox agent API calls keep working regardless of how the public base URL is configured or gated ## Linked Issues or Issue Description No existing issue. Describing in-PR following the bug report template: **What happened?** On a cloud deployment with a session-gated public edge, setting a public base URL (PAPERCLIP_PUBLIC_URL) caused every in-sandbox agent API call through the sandbox callback bridge to fail with `403 text/plain "Access denied"` from the edge proxy. With PAPERCLIP_BRIDGE_DEBUG enabled, the bridge logs show the forward target is the public origin, and every proxied request (for example `GET /api/agents/me`) returns the edge proxy's 403 instead of reaching the API. **Expected behavior** The bridge worker runs in the same server process that serves the API, so forwarded calls should target the local listen origin and succeed regardless of how the public origin is configured or gated. **Steps to reproduce** 1. Run the server with a public base URL configured, fronted by a proxy that requires a browser session on API routes. 2. Start a sandbox-executed agent run (any adapter using the sandbox callback bridge). 3. Observe every in-sandbox call to the Paperclip API fail with the proxy's 403; with PAPERCLIP_BRIDGE_DEBUG the forward URL is the public origin. **Paperclip version or commit** Current `master`. **Deployment mode** Self-hosted server behind a reverse proxy. **Agent adapter(s) involved** All sandbox-executed adapters (the bridge is adapter-agnostic). ## What Changed - `packages/adapter-utils/src/execution-target.ts`: `startAdapterExecutionTargetPaperclipBridge` now resolves its forward target as `input.hostApiUrl?.trim() || resolveDefaultPaperclipApiUrl()`. It no longer consults `PAPERCLIP_RUNTIME_API_URL` / `PAPERCLIP_API_URL`, which now describe the public origin for browsers and external agents, exactly the wrong target for an in-process loopback hop. `resolveDefaultPaperclipApiUrl()` builds `http://<PAPERCLIP_LISTEN_HOST>:<PAPERCLIP_LISTEN_PORT>` (exported by server boot before any run executes) and maps wildcard listen hosts to the loopback address of the same family (`0.0.0.0` to `127.0.0.1`, `::` to `[::1]`), so the forward target always matches the address family the server is bound to. `input.hostApiUrl` remains the explicit override seam. A comment documents the reasoning. - `packages/adapter-utils/src/execution-target-sandbox.test.ts`: two new tests. One sets both public URL env vars to an unreachable public https origin and asserts the bridge forwards to the local listen origin (fails before this fix with a 502 because the worker targets the public origin). One asserts an explicit `hostApiUrl` input still overrides everything. - The acpx-engine bridge start (`packages/adapter-utils/src/acpx-engine/execute.ts`) passes no `hostApiUrl` and goes through the same resolution site, so it is covered by the same fix. The sandbox-facing env builder in `server-utils.ts` is intentionally untouched; the bridge env overrides `PAPERCLIP_API_URL` inside the sandbox separately. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` (28 tests pass; the new local-origin test fails without the fix) - `pnpm --filter @paperclipai/adapter-utils typecheck` (clean) - Full adapter-utils suite run; the only failures are pre-existing environment-dependent tests (bubblewrap and shallow-clone tests on macOS) identical on a clean `master` checkout ## Risks - Low risk. Deployments where the bridge previously worked did so precisely because the forward target already resolved to the local origin (no public URL configured, so the chain fell through to the same `resolveDefaultPaperclipApiUrl()` result). The only behavioral shift is for deployments with a public URL configured, where forwarding through the edge was either wasteful (an unnecessary network round trip) or broken (session-gated edge). The explicit `hostApiUrl` override seam is preserved for callers that need a nonlocal target. ## Model Used - Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1f7959bc69 |
fix(codex-local): skip benign stderr warnings when deriving the fallback run error (#10003)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents execute through adapters; the codex_local adapter runs the Codex CLI and reports each run's outcome, including an error message when the CLI exits nonzero > - When no error can be parsed from the CLI's JSONL output, `toResult` in `packages/adapters/codex-local/src/server/execute.ts` falls back to the first non-empty stderr line as the run error > - The adapter itself passes the approvals-bypass flag, so the CLI's first stderr line is always the benign startup warning "YOLO mode is enabled. All tool calls will be automatically approved." > - Failed runs therefore record that warning as their error, hiding the real cause (for example an OpenAI API error further down in stderr) and making failures hard to diagnose from the run record > - This pull request derives the fallback error from the first meaningful stderr line, skipping a conservative set of known benign lines, and keeps the existing behavior when every line is benign > - The benefit is that failed Codex runs surface the actual failure reason instead of a harmless startup warning, without ever producing an emptier message than before ## Linked Issues or Issue Description No public issue exists for the codex_local case. The same bug class was fixed for gemini-local in Refs #5099 and Refs #3476; this PR applies the equivalent fix to codex_local. **What happened?** On a multi-tenant cloud deployment of Paperclip, several codex_local runs failed and their run records showed `error_code=adapter_failed` with the error text "YOLO mode is enabled. All tool calls will be automatically approved." That is a benign Codex CLI startup warning, printed on every run because the adapter passes the approvals-bypass flag itself. The real failure (an OpenAI API error printed later in stderr) was never surfaced. **Expected behavior** When the Codex CLI exits nonzero and no error was parsed from its JSONL output, the run error should be the first stderr line that actually explains the failure, not a startup warning the adapter itself provoked. **Steps to reproduce** 1. Configure a codex_local agent and make the underlying Codex CLI invocation fail after startup (for example, configure a model id the active credentials cannot use). 2. Run the agent so the CLI exits nonzero with no parsed JSONL error. 3. Inspect the run's error message: it shows the YOLO approvals warning (the first stderr line) instead of the real error printed further down in stderr. ## What Changed - Added `firstMeaningfulStderrLine` next to `firstNonEmptyLine` in `packages/adapters/codex-local/src/server/execute.ts`, with a conservative benign-line predicate covering the YOLO approvals warning and `[paperclip] ...` diagnostic lines the adapter injected (for example ACP fallback notes). - Used it only in the `toResult` fallback error derivation. If every stderr line is benign, the existing chain still applies (first non-empty line, then `Codex exited with code N`), so the message never gets emptier than today. Logging is unchanged. - Added `packages/adapters/codex-local/src/server/execute.stderr-error.test.ts`: four end-to-end cases through `execute()` with a mocked CLI process, plus unit coverage for the new helper. Tests were written first and confirmed failing before the fix. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/execute.stderr-error.test.ts` (7 tests pass; 5 failed before the fix as expected) - `pnpm exec vitest run packages/adapters/codex-local` (21 files, 188 tests pass) - `pnpm run typecheck` in `packages/adapters/codex-local` (clean) ## Risks Low risk. Only the derived fallback `errorMessage` changes, and only when a benign line would otherwise have been picked; parsed JSONL errors, logging, retry/quota/auth classification inputs, and the empty-stderr exit-code fallback are untouched. The benign-line list is deliberately conservative (exact prefixes) so real errors are never skipped. ## Model Used Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5521d768b2 |
fix(claude-local): avoid root-only skip permissions failure (#9463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Claude local is one of the adapter paths that lets operators run Claude Code through a local Paperclip runtime. > - Claude Code rejects `--dangerously-skip-permissions` when the process is running as root or through sudo. > - Local/self-hosted Paperclip deployments may run inside root-owned Docker/runtime processes, so the Claude local adapter can fail before it reaches the actual runtime/auth condition. > - Paperclip already uses a curated `--allowedTools` list instead of `--dangerously-skip-permissions` for remote Claude targets. > - This pull request applies the same safer permission strategy to local root processes while preserving existing local non-root and remote behavior. > - The benefit is clearer, safer Claude local diagnostics/execution in containerized setups without widening permissions beyond the existing explicit tool allowlist. ## Linked Issues or Issue Description No directly matching public issue or PR found. Bug description: - **Problem:** `claude_local` can fail its local probe/execution path when Paperclip runs from a root-owned local container/runtime because Claude Code refuses `--dangerously-skip-permissions` under root/sudo. - **Actual behavior:** The adapter may fail immediately with Claude's root/sudo guard before validating the real Claude runtime/auth state. - **Expected behavior:** Local root processes should use the same explicit allowlist strategy Paperclip already uses for remote targets, while local non-root behavior remains unchanged. - **Environment:** Local/self-hosted Docker or container-style Paperclip runtime where the app process UID is `0`. Related but different: #4926 covers MCP config propagation for the Claude local adapter, not the root/sudo permission flag behavior fixed here. ## What Changed - Added root-aware permission argument selection for the Claude local adapter. - Preserved current local non-root behavior: `--dangerously-skip-permissions` is still used when allowed. - Preserved current remote behavior: remote targets continue using explicit `--allowedTools`. - Changed local root behavior to use the explicit `--allowedTools` list instead of `--dangerously-skip-permissions`. - Threaded process UID awareness through Claude local probe and execution paths. - Added unit coverage for skip-disabled, remote, local non-root, local root, and UID-unavailable behavior. ## Verification ```sh ./node_modules/.bin/vitest run --config g15-vitest-claude-local.config.mjs \ packages/adapters/claude-local/src/server/permissions.test.ts ``` Result: ```text 1 file passed 8 tests passed ``` ```sh pnpm --filter @paperclipai/adapter-claude-local typecheck ``` Result: ```text @paperclipai/adapter-claude-local typecheck passed ``` Additional local smoke: - Ran a disposable root-container Claude adapter diagnostic against this patch. - The diagnostic no longer fails with Claude's root/sudo `--dangerously-skip-permissions` error. - It proceeds to the actual environment-specific Claude auth/runtime result. - No credentials, tokens, hostnames, private paths, or internal Paperclip issue references are included in this PR. Public duplicate checks performed: ```sh gh pr list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools' gh issue list --repo paperclipai/paperclip --state open --search 'claude local root permissions dangerously skip permissions allowedTools' ``` ## Risks Low-to-medium risk adapter behavior change: - Local root Claude runs will now use explicit `--allowedTools` rather than broad skip-permissions behavior. - That is intentionally safer, but an environment depending on broader implicit tool access under root may now need the adapter allowlist to include any required tools. - Local non-root behavior is unchanged. - Remote behavior is unchanged. - No database migrations, API contract changes, or UI changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex `gpt-5.5` via Hermes Agent, with shell/file/tool use for repository inspection, patching, local verification, and GitHub CLI operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: LeeJ <elJayAdvisor@users.noreply.github.com> |
||
|
|
04bf7a6ab5 |
feat(observability): instrument stage.sync host steps and home the agent process span (#11301)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - It runs each agent in a remote sandbox and emits OpenTelemetry spans for the sandbox bring-up and the run. > - A real trace showed two gaps. `stage.sync` had about 3 seconds of unattributed host work before its `pack` span. The persistent agent process showed a `sandbox.exec` span that outlived its parent by about 50 seconds. > - The gaps hide real cost and make the trace read as a sequencing bug, so an operator cannot see where startup time goes. > - This pull request wraps the two pre-`pack` host steps in their own spans. It also homes the long-lived process in a run-scoped `sandbox.agentProcess` span. > - The benefit is that startup time is fully attributed and the process reads as a resource that overlaps the turn, not a child that outlives its parent. ## Linked Issues or Issue Description No public issue exists. This is an enhancement to existing telemetry. It is described inline below, following `.github/ISSUE_TEMPLATE/enhancement.yml`. Prior related work: the merged PR #10999 added the run-time wrapper spans and the telemetry data-contract section this PR extends. **What existing behavior does this improve?** The sandbox bring-up and run OpenTelemetry trace. It closes two attribution gaps in that trace. **Subsystem affected** Observability for sandbox execution. The code lives in `packages/adapter-utils`. The span contract lives in `packages/shared/src/telemetry`. **Current behavior** `stage.sync` opens a `pack` span, but the git enumeration and the baseline content-hash walk that run before `pack` have no span, so about 3 seconds read as a gap. On the streamed process-session path the agent process launches fire-and-forget inside the ~2.3 second `bridge.process-session` bring-up step, so its `sandbox.exec` span parents to that step and then runs about 50 seconds. The child dangles past its parent and overlaps `agent.turn`. **Proposed behavior** Wrap the two pre-`pack` host operations in `snapshot.git` and `snapshot.baseline` spans under `stage.sync`. Wrap the streamed launch in a run-scoped `sandbox.agentProcess` span that parents to the live run root (`task.run` at launch). **Reason and benefit** Startup time is fully attributed. The long-lived process reads as a resource that overlaps the sibling `agent.turn`, not a mis-parented child. **Breaking changes** None. The spans are opt-in and export only when an OTLP endpoint is configured. The span seam is a no-op when no runner is injected. No first-party telemetry event changes. ## What Changed - `sandbox-managed-runtime.ts`: add `snapshot.git` and `snapshot.baseline` spans around the git enumeration and the baseline content-hash walk, nested under `stage.sync`, through a shared `runStepSpan` helper that `pack` now also uses. - `execution-target.ts`: wrap the fire-and-forget streamed launch in a run-rooted `sandbox.agentProcess` span, so it parents to the live run root and holds the inner `sandbox.exec`. The `.then`/`.catch` chain became try/catch inside the span callback, with identical frame-ingestion behavior. - `packages/shared/src/telemetry/README.md`: update the span table and the parenting prose. Add `snapshot.git`, `snapshot.baseline`, `pack`, and `sandbox.agentProcess`, and document the intended `sandbox.agentProcess` / `agent.turn` overlap. - Tests: update the executor span-tree test (`childNames` and parent assertions), update the `sandbox-managed-runtime` span-set and nesting tests, and add two `execution-target-sandbox` tests (the launch opens `sandbox.agentProcess`; it parents to the run root, not the bring-up step). ## Verification - Run `npx vitest run` on the three affected test files. Result: 174 tests pass. This includes the updated executor span-tree test and the new `sandbox.agentProcess` open and parenting tests. - Run `tsc --noEmit` in `packages/adapter-utils`. Result: no errors in the changed source or test files. - The full 37-test streamed process-session suite passes unchanged. This confirms the try/catch restructure preserves frame delivery and exit/error behavior. - Pre-existing and unrelated to this PR (present on `master`): `tsc` errors in `execute.ts` / `execute.test.ts` / `remote-spawn-smoke.test.ts` (`onAgentStderr` / `spawnCwd`), and a `check:forbidden-tokens` failure from internal `PAP-###` ids in `ui/src/components/IssueRecoveryActionCard.test.tsx`. This PR does not touch those files, and its own diff is token-clean. ## Risks Low. The change adds instrumentation on the opt-in span path and does not change control flow on the default path. The one production restructure is the streamed launch, which stays fire-and-forget, so bring-up does not block on it. Only the streamed path gains `sandbox.agentProcess`; the legacy poll path launches the process detached and has no host-side long-lived span to home. ## Model Used Anthropic Claude Opus 4.8 (`claude-opus-4-8`), about 200K-token context, agentic tool use through Claude Code. The trace was reviewed through the Honeycomb MCP. The code was written and tested with the model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e31951a17d |
feat: Claude agent setup-token login in a sandbox (#11286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Claude agents that run in a remote sandbox need a safe in-product login path > - The existing host login route cannot open a pseudo-terminal inside that sandbox > - The login flow must protect the browser code, the login URL, and the OAuth token at every step > - This pull request adds the parser, the runner, a Daytona pseudo-terminal transport, and a guarded, owner-bound session route behind an injectable transport > - The route stays inert in the default build and fails closed until a sandbox provider binds the live transport > - The benefit is a company-scoped setup-token flow with one-time secret delivery, redaction, and fail-closed transport checks, ready for a later staged production rollout ## Linked Issues or Issue Description **Agent or provider** Claude Code setup-token login for sandbox agents. **Why this adapter is useful** Sandbox agents need a supported way to sign in without host credentials. An authorized owner completes the browser step and receives the token one time. **How the agent is invoked** When a sandbox provider binds the injectable transport, the server starts `claude setup-token` through a sandbox pseudo-terminal, sends the browser code to the matched prompt, and returns the token through the guarded session route. The default build does not bind the transport. In that state the start route fails closed with a fixed no-secret `503`. It does not start a process and it does not hold a sandbox lease. **Additional context** The transport is injectable, so each sandbox provider binds its own pseudo-terminal. This pull request adds the Daytona transport but does not bind it in the production server. A production wiring needs a lease manager, a live pseudo-terminal factory, a durable token store, and its own security review. The route keeps secrets out of logs, activity details, errors, telemetry, and non-owner responses. ## What Changed - Add strict parsers for the setup-token URL, the prompt, and the success token. - Add a login runner that drives the `claude setup-token` command through a pseudo-terminal. - Add the Daytona pseudo-terminal transport and the sandbox plugin wiring. - Add a company-scoped, owner-bound login session service with rate limits, a reaper, cleanup, and one-time token delivery. - Add the guarded session routes at `/agents/:id/setup-token-login-sessions/*` behind an injectable transport. The routes become the live login path only when a provider binds the transport. - Keep the start route fail-closed in the default build. It returns a fixed no-secret `503` and it does not bind `setupTokenLogin`. - Keep the existing host route `POST /agents/:id/claude-login` in place. This pull request does not replace it. - Keep confidential responses behind a fail-closed TLS transport guard with `Cache-Control: no-store`, and extend redaction for the new fields. - Export the parser and the runner from the Claude local server entry, and document the new session routes in the OpenAPI spec. ## Verification - `pnpm --filter @paperclipai/server exec vitest run setup-token-route setup-token-session` - `pnpm --filter @paperclipai/adapter-claude-local exec vitest run` - `pnpm --filter @paperclipai/server run typecheck` - Confirm that the pull request checks pass on GitHub. ## Risks - Low user-facing risk on merge. The default build does not bind the transport, so the production start route stays fail-closed with a `503`. The merge does not change the production login behavior. - When a provider later binds the transport, the flow starts a live sandbox process and holds a short-lived in-memory secret. Cleanup must stop the child before it releases the sandbox lease. - The transport guard fails closed when the deployment does not provide a trusted TLS path. A wrong proxy allowlist can block a valid request. - The production wiring is out of scope. It needs a lease manager, a live pseudo-terminal factory, a durable token store, and its own security review before the server binds `setupTokenLogin`. ## Model Used Anthropic Claude Opus 4.8 assisted the implementation. It used extended reasoning, code execution, repository tool use, and a 200,000-token context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (the OpenAPI spec covers the new session routes; no user-facing documentation needs changes) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c57c0f7498 |
fix(sandbox-providers): accept bsdtar listings in the syncOut tarball confinement check (#11289)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox providers (Daytona, Kubernetes) sync run results back to the host with a sandbox-authored tarball > - Before extraction, a confinement check parses the host `tar -tvf` listing and fails closed on unparseable lines > - The parser only understands the GNU tar listing dialect; macOS ships bsdtar, whose ls-style listing never matches > - Every sandbox syncOut on a macOS host therefore aborts with "refusing tarball with an unparseable entry listing", and the run fails at copy-back > - This pull request teaches the parser both dialects while keeping the fail-closed and traversal guarantees > - The benefit is that Daytona and Kubernetes sandbox runs work on macOS hosts, with no behavior change on Linux ## Linked Issues or Issue Description No existing issue. Bug description: **What happened?** On a macOS host, every Daytona sandbox run fails at syncOut. The adapter reports: `Daytona syncOut refusing tarball with an unparseable entry listing: -rw-r--r-- 0 daytona daytona 7560 Aug 11 21:43 AGENTS.md`. The Kubernetes provider has the same parser and fails the same way. **Expected behavior** The confinement check accepts a well-formed listing from the host tar, whichever dialect the host tar emits. It still rejects members that escape the extraction directory, and it still fails closed on lines it cannot parse. **Steps to reproduce** 1. Run Paperclip on macOS (system tar is bsdtar). 2. Configure an agent with the Daytona sandbox provider. 3. Trigger any run that syncs files back from the sandbox. 4. The run fails at syncOut with the unparseable-entry-listing error, because bsdtar prints `<perms> <links> <user> <group> <size> <Mon> <day> <time|year> <name>` while the parser expects the GNU `<perms> <owner>/<group> <size> <date> <time> <name>` shape. **Operating system** macOS (bsdtar 3.5.3). Linux hosts with GNU tar are unaffected. ## What Changed - Extracted the listing-line parse in both providers' `file-sync.ts` into an exported `parseTarVerboseListingLine` that accepts the GNU/busybox dialect and the bsdtar (libarchive) dialect. - The GNU shape now requires the slash-joined `<owner>/<group>` field. This keeps the two shapes mutually exclusive. Without it, a bsdtar line with numeric uid/gid satisfies the loose GNU pattern shifted by one field, which would hide a leading `../` from the traversal check. - Unparseable lines still fail closed. This includes device-node entries, whose size column is `major,minor` in both dialects. - Made the path-traversal fixture in the Daytona suite portable: GNU spells member renaming `--transform`, bsdtar spells it `-s`. - Added a Daytona test that refuses a sandbox-authored tarball carrying a symlink whose target escapes the extraction dir. - Added parser unit tests for both dialects (file, dir, symlink, hardlink, numeric owner, year-form dates, fail-closed lines) to both providers' suites. ## Verification - `pnpm test` in `packages/plugins/sandbox-providers/daytona`: 136/136 pass on a macOS host. On unpatched `master` the round-trip test fails there with the unparseable-entry-listing error. - `pnpm test` in `packages/plugins/sandbox-providers/kubernetes`: the new parser tests pass; no new failures against the `master` baseline on the same host. - `pnpm typecheck` passes in both packages. - CI runs the same suites on Linux/GNU tar and proves the GNU path is unchanged. ## Risks - Low risk. The GNU pattern is one token stricter (`<owner>/<group>` must contain `/`). GNU and busybox tar always print the slash-joined owner field, so accepted GNU listings are unchanged. - The bsdtar branch only widens acceptance on hosts that were failing 100% of syncOuts before, so no working deployment changes behavior. - The check still fails closed on anything neither pattern matches. ## Model Used Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking + tool use). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
67001ec6eb |
chore(db): treat Drizzle migration snapshots as binary in diffs (#11254)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is about 39k lines. > - PR #11240 marked these files `linguist-generated=true`. That collapses the file view and drops the files from language stats. > - But `linguist-generated` does not change the PR line-count badge. Git still counts every snapshot line, so a PR that adds one migration shows a ~39k-line badge (two migrations show ~78k). PR #11237 shows +84,665 for this reason. > - This pull request adds `-diff` to the same files, so git treats them as binary and their lines leave the +/- count. > - The benefit is a pull request badge that reflects the real code change, not generated snapshot noise. ## Linked Issues or Issue Description No existing issue. This is a small follow-up to merged PR #11240. Description follows the enhancement template: **Problem / Motivation** PR #11240 marked `packages/db/src/migrations/meta/**` as `linguist-generated=true`. That attribute collapses the diff in the Files-changed view and removes the files from language stats, but it does not remove their lines from the PR additions/deletions badge. Each Drizzle snapshot is a full copy of the schema (~39k lines), so any PR that adds a migration still shows a huge line count. PR #11237 shows +84,665, of which ~78k are two generated snapshots. **Proposed Solution** Add `-diff` to the same glob. Git then treats the snapshots as binary. `git diff --numstat` reports `-` for these files, so their lines leave the +/- badge and GitHub shows "Binary file not shown" in place of the full JSON. **Alternatives Considered** - Keep only `linguist-generated`: leaves the misleading ~39k/78k badge on every migration PR. - Use the `binary` macro (`-diff -merge -text`): also disables EOL normalization. This repo has open CRLF/LF work, so `-text` is left unset on purpose. ## What Changed - Added `-diff` to `src/migrations/meta/**` in `packages/db/.gitattributes`. - Kept `linguist-generated=true` (language stats) and `-merge` (no auto-merge of generated snapshots). - Left `-text` unset on purpose, so end-of-line normalization stays intact. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Check the attributes and confirm the snapshot is now treated as binary: ``` git check-attr linguist-generated diff merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/0009_fast_jackal.sql git diff --numstat origin/master...origin/feat/adapter-sandbox-login -- \ packages/db/src/migrations/meta/0214_snapshot.json ``` Expected: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: diff: unset packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: diff: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified - - packages/db/src/migrations/meta/0214_snapshot.json ``` The snapshot reports `-` in numstat (binary, not counted). The `.sql` migration keeps normal diff behavior. GitHub reads the rule from the PR tree, so the badge drops on the next PR that touches these files. ## Risks Low risk. The change only affects how git and GitHub render and count generated files. It does not touch application code, the schema, or any migration. - With `-diff`, GitHub and local `git diff` no longer show a text diff for a snapshot. This is intended; the files are generated and are not reviewed by hand. The raw file is still viewable. - `-text` is left unset, so this change does not affect the CRLF/LF handling that other PRs (for example #8922) address. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` and `git diff --numstat` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0aa743fc30 |
build(db): clean dist before drizzle generate (#11241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Database migrations are generated by drizzle-kit, which reads the schema from the db package's built `dist/schema/*.js` > - `tsc` never deletes stale outputs, so a long-lived checkout keeps compiled schema files whose sources were deleted long ago > - A `generate` run in such a checkout sees those ghost tables and sweeps phantom `CREATE TABLE` statements into an unrelated migration > - This pull request makes `generate` clean `dist` before building, so drizzle always diffs against exactly the current schema sources > - The benefit is that no contributor can accidentally resurrect deleted tables inside a new migration ## Linked Issues or Issue Description **What happened?** Running `pnpm --filter @paperclipai/db generate` in a months-old checkout produced a migration re-creating `cloud_upstream_connections`, `cloud_upstream_runs`, and `company_secret_pools` — tables whose schema sources were deleted in #10507. The compiled copies were still in `dist/schema/`, and `drizzle.config.ts` reads the schema from `dist`, so drizzle treated them as new tables missing from the snapshot. **Expected behavior** `generate` diffs the current schema sources only; deleted tables can never reappear in a generated migration. **Steps to reproduce** 1. Build the db package, then delete a schema source file without cleaning `dist`. 2. Run `pnpm --filter @paperclipai/db generate`. 3. The generated migration re-creates the deleted table. ## What Changed - `packages/db/package.json`: the `generate` script runs `pnpm run clean` before `tsc`, so the drizzle-kit input is always a fresh build of the current sources. ## Verification - In a checkout carrying stale `dist/schema/cloud_upstreams.js` / `company_secret_pools.js` artifacts, `generate` produced a phantom migration before this change; after a clean build it reports "No schema changes, nothing to migrate". With this change the clean happens inside `generate` itself. ## Risks - Low risk: `generate` is a developer-only script; the change only adds the existing `clean` step ahead of the existing build, at the cost of a full rebuild per generate. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b5ebda1dca |
fix(grok-local): report real token usage and cost instead of hardcoded zeros (#10433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Cost/usage tracking is core to that: the dashboard shows per-agent spend so a team can see what their AI workforce is costing them > - The `grok_local` adapter (xAI's Grok Build CLI) is a newer adapter than `claude_local`/`codex_local`, and its usage/cost wiring was left incomplete > - Every `grok_local` run persists `usage.inputTokens/outputTokens/cachedInputTokens = 0` and `costUsd = null` in `heartbeat_runs`, unconditionally, even though the underlying `grok` CLI reports real, non-zero token counts and cost per turn in its own JSON stream > - This pull request wires the parser to actually read `usage`/`total_cost_usd` from the CLI's terminal `end` event, threads those values into the adapter's execution result, and marks them `usageBasis: "per_run"` so the heartbeat service doesn't incorrectly delta them against a prior run on a resumed session (matching how `claude_local`/`codex_local` already do this) > - The benefit is accurate cost/usage visibility for any self-hosted Paperclip instance running Grok Build agents, instead of a dashboard that always reads zero ## Linked Issues or Issue Description Fixes: #10432 ## What Changed - `packages/adapters/grok-local/src/server/parse.ts`: `parseGrokJsonl()` now reads `usage.input_tokens` / `usage.output_tokens` / `usage.cache_read_input_tokens` / `total_cost_usd` from the terminal `end` event and returns them on `ParsedGrokJsonl` (previously discarded entirely). - `packages/adapters/grok-local/src/server/execute.ts`: `toResult()` now populates `usage.inputTokens/outputTokens/cachedInputTokens` from the parsed values instead of hardcoded `0`, sets `usageBasis: "per_run"` (each `--single` invocation reports usage for just that process, not a running session total), and surfaces `costUsd` only when `billingType === "api"` (metered) — subscription/OAuth billing has no marginal dollar cost, so it stays `null` there, but token counts are populated for both billing types since usage visibility is useful regardless of billing model. - `packages/adapters/grok-local/src/server/parse.test.ts`: added a test asserting usage/cost extraction from a representative `end` event payload, and updated the existing exact-equality test for the new fields. - `packages/adapters/grok-local/src/server/execute.test.ts`: added a test covering both subscription billing (tokens populated, `costUsd: null`) and API-key billing (tokens populated, real `costUsd`), and asserting `usageBasis: "per_run"` in both cases. ## Verification - `pnpm vitest run packages/adapters/grok-local/src/server/parse.test.ts packages/adapters/grok-local/src/server/execute.test.ts` — 9/9 passed - `tsc --noEmit` on the `grok-local` package — clean - Verified against a real self-hosted Paperclip instance running `grok` CLI `0.2.112` with SuperGrok subscription (OAuth) auth: confirmed the raw CLI stream reports real `usage`/`total_cost_usd` (e.g. `"usage":{"input_tokens":21560,...},"total_cost_usd":0.0564448`) that was previously discarded before ever reaching `heartbeat_runs.usage_json`, which always showed all-zero tokens regardless of real usage. ## Risks - Low risk, additive change scoped entirely to the `grok_local` adapter's usage/cost reporting path — no change to control flow, session handling, or process execution. - `usageBasis: "per_run"` mirrors the existing, already-tested pattern in `claude_local`/`codex_local` execute paths, so the heartbeat service's per-run vs. session-cumulative delta logic is exercised the same way. - `costUsd` is intentionally left `null` for subscription/OAuth billing (no behavior change there beyond now-populated token counts) to avoid implying a dollar cost that doesn't exist for flat-rate billing. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, no extended thinking. Root cause was found by comparing real `grok` CLI JSON stream output (captured directly from a live invocation) against the persisted `heartbeat_runs.usage_json` row for the same run on a self-hosted instance, then reading `parse.ts`/`execute.ts` source to confirm the hardcoded zero values. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/grok-local-usage-cost-tracking`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no user-facing docs reference this internal usage-reporting behavior) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending at time of writing) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (addressed the one P1 raised — `usageBasis: "per_run"`) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c0bdf26633 |
chore(db): collapse Drizzle migration snapshot diffs and block auto-merge (#11240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip stores its state in a Postgres database, managed in `packages/db`. > - The schema uses Drizzle. `drizzle-kit` writes a full-schema snapshot to `packages/db/src/migrations/meta/` for every migration. > - Each snapshot is a large generated JSON file. One snapshot is over 200 KB. > - GitHub shows these files as 30k+ line diffs in a pull request. The diffs add no review value, because a human never edits the files. > - The files also merge badly. `drizzle-kit` computes the `id`/`prevId` chain and the full schema state, so a line-level merge of two snapshots produces a file that no real `generate` run creates. > - This pull request adds a `.gitattributes` file that marks the snapshot directory as generated and blocks its auto-merge. > - The benefit is clean pull request diffs and a loud conflict that forces the correct fix when two branches add a migration. ## Linked Issues or Issue Description No existing issue. This is a small repository-hygiene change. Description follows the enhancement template: **Problem / Motivation** Every migration adds a full-schema snapshot JSON under `packages/db/src/migrations/meta/`. These files are large and generated. GitHub renders them as 30k+ line diffs in pull requests, which buries the real change (the `.sql` migration) in noise. The files also have no meaningful line-level merge: `drizzle-kit` computes each snapshot's `id`/`prevId` chain and full schema state. **Proposed Solution** Add `packages/db/.gitattributes`: - `linguist-generated=true` on `src/migrations/meta/**` — GitHub collapses the diff and drops the files from language stats. - `-merge` on the same glob — git refuses the line-level merge and raises a conflict instead of fabricating an invalid snapshot. **Alternatives Considered** - `-diff` / `binary`: hides the diff completely and blocks text merge, but also blocks any local `git diff` and gives a worse conflict experience. `linguist-generated` keeps the file expandable and text-based, so it is the lighter option. - Do nothing: leaves the noisy diffs and the risk of a silent bad merge. ## What Changed - Added `packages/db/.gitattributes`. - Marked `src/migrations/meta/**` as `linguist-generated=true` to collapse the snapshot and journal diffs on GitHub. - Set `-merge` on the same files so git raises a conflict instead of auto-merging generated snapshots. - Left the `.sql` migration files untouched, so their diffs stay visible for review. ## Verification Run `git check-attr` against the affected files and a control `.sql` file: ``` git check-attr linguist-generated merge -- \ packages/db/src/migrations/meta/0031_snapshot.json \ packages/db/src/migrations/meta/_journal.json \ packages/db/src/migrations/0009_fast_jackal.sql ``` Expected output: ``` packages/db/src/migrations/meta/0031_snapshot.json: linguist-generated: true packages/db/src/migrations/meta/0031_snapshot.json: merge: unset packages/db/src/migrations/meta/_journal.json: linguist-generated: true packages/db/src/migrations/meta/_journal.json: merge: unset packages/db/src/migrations/0009_fast_jackal.sql: linguist-generated: unspecified packages/db/src/migrations/0009_fast_jackal.sql: merge: unspecified ``` The snapshot and journal files carry both attributes. The `.sql` migration keeps its normal diff and merge behavior. GitHub applies the rule from the pull request tree, so the collapse shows on the next pull request that touches these files. ## Risks Low risk. The change only affects git and GitHub display and merge behavior for generated files. It does not touch application code, the schema, or any migration. - `-merge` leaves the current-branch version in the working tree on conflict and marks the file conflicted. It does not insert conflict markers into the JSON. The correct resolution stays "renumber the later migration and regenerate", then commit. - Related open pull requests #879 and #8922 also add `.gitattributes` rules for migration files, but for CRLF/LF hash mismatches on Windows. If either lands, a follow-up can merge the rules into one file. There is no functional overlap with this change. ## Model Used Claude Opus 4.8 (Anthropic), model ID `claude-opus-4-8`, used through Claude Code with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — not applicable; this change touches no code path, only git/GitHub file handling - [x] I have added or updated tests where applicable — not applicable; `.gitattributes` behavior is verified with `git check-attr` (see Verification) - [x] I have updated relevant documentation to reflect my changes — not applicable; the `.gitattributes` file documents its own rules inline - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8f478242f1 |
fix(adapter-utils): close sandbox stdin file race with atomic write and fault-tolerant poller (#11235)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use execution targets to exchange input and output with sandbox processes > - The sandbox input path can expose partial files, and the poller can delete or stop on invalid input > - These timing windows can lose agent input without a clear error > - This pull request makes host writes atomic and makes the poller retry invalid files before it drops them > - The benefit is reliable sandbox input delivery with visible failure after bounded retries ## Linked Issues or Issue Description Closes #10874 ## What Changed - Decode host input into a temporary file, then rename it onto the final JSON path. - Apply the same atomic write pattern to the filesystem client. - Parse each input file before deletion. - Retry parse failures and drop a file after the bounded retry limit with an error event. - Add regression tests for empty, partial, and permanently malformed input files. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts` passes 5/5 tests. - Related sandbox callback, execution target, and sandbox execution suites pass 71/71 tests. - TypeScript checks pass for the changed files. - The regression suite fails on the old code and passes on this change. ## Risks - Low risk. The change affects sandbox input file handling and adds bounded retry behavior. - A permanently malformed file now creates an error event after the retry limit. > This bug fix does not add a core feature, so a roadmap change is not needed. ## Model Used Codex, OpenAI GPT-5, current agent runtime, large context window, tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
23a1b025c2 |
feat(server): chunked resumable company import transfers (#11223)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import moves large packages into an instance, and since the upload cap rose to 1 GB, the transport is the weak point: one HTTP request, buffered fully in memory, with no resume > - A dropped connection at 90% of an 800 MB upload starts the whole transfer over, and a server restart loses all progress > - This pull request adds the server side of chunked resumable import transfers: a durable run ledger and routes that accept the same import zip as verified ~32 MB parts spooled to disk > - An interrupted transfer resumes from the parts already uploaded — across dropped connections, page refreshes, and server restarts — and peak upload memory drops from the whole package to one part > - The benefit is that large imports become reliable on real-world connections instead of all-or-nothing ## Linked Issues or Issue Description **What happened?** Large company imports travel as a single HTTP upload. On a slow or flaky connection, any interruption discards all progress and the upload restarts from zero. The server buffers the entire compressed package in memory during upload. A server restart mid-upload loses the transfer entirely. With the upload cap now at 1 GB, these failure modes govern exactly the imports the cap was raised for. **Expected behavior** A large import upload survives interruptions: already-transferred data is kept and verified, only the missing remainder is re-sent, and the server's memory use during upload is bounded by a part, not the package. **Steps to reproduce** 1. Import a multi-hundred-MB company package over a connection that drops mid-upload. 2. The upload fails; retrying starts from byte zero. 3. Repeat on an unstable connection and the import may never complete. ## What Changed - New `company_transfer_runs` table (drizzle schema + migration) and `companyTransferRunService`: one row per transfer with a content-derived idempotency key, per-part completion recorded atomically and idempotently, resume scoped to actor and direction, completed runs short-circuiting retries of identical content. - New transfer routes beside the existing import routes, same authorization: declare a sliced zip (`POST /import/transfers` — validates cap, 64 MB part ceiling, contiguity, size sums, sha256 format), upload parts (`PUT .../parts/:n` — raw body, hash-and-size verified before an atomic write to a disk spool under the instance root; re-uploads are no-op successes), poll resume state (`GET .../:id` — missing parts recomputed from disk), and apply (`POST .../:id/apply` — requires all parts, re-verifies the assembled zip against the whole-file hash fail-closed, then feeds the existing import pipeline through factored helpers rather than duplicated logic). - Hourly sweep fails and cleans spools idle for 24 h; a swept transfer honestly reports all parts missing on resume. - Strict UUID gating on run ids before any filesystem path construction. - The existing single-shot upload path is untouched; clients arrive in the follow-up PR. ## Verification - Transfer route suite (embedded Postgres): create/upload/status/apply round-trip with a real imported company, out-of-order parts, wrong-hash part rejected and unrecorded, re-upload no-op, apply-with-missing-parts rejection, resume after failure with prior progress intact, assembled-hash mismatch failing closed with spool deletion, actor scoping 404s, async-job apply, sweep followed by honest resume. - Ledger suite (embedded Postgres): part idempotency, actor/direction scoping, completed-run short-circuit, cancelled runs staying cancelled. - Existing portability route suite unchanged and green; server + db typechecks clean. Exact counts in the PR checks. ## Risks - New routes are additive; the existing import path is untouched. The transfer routes carry the same board authorization as the import routes they sit beside. - Disk spool: bounded by the existing upload cap per transfer, cleaned on success, failure, hash mismatch, and by the 24 h sweep. Spool paths are strict-UUID-gated. - The apply step still materializes the assembled zip in memory once (same profile as today's single-shot import at apply time); upload-time memory drops to one part. - Known limitation, deliberate: transfers are keyed on content alone, so identical package content cannot be imported twice without re-exporting (surfaced explicitly to the caller). Acceptable for v1; noted for review. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0044fa8904 |
Let tenants edit env vars on managed sandbox environments; add managed-sandbox-only mode (#11200)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments give each agent run an execution target: the local host, SSH, or a sandbox provider > - A managed deployment can provision one platform-managed sandbox environment through the `PAPERCLIP_MANAGED_CONFIG` `environments` section > - That row is fully locked today. A tenant cannot add environment variables for their agents. There is also no way to hide local execution — run selection falls back to the local row > - A platform that manages the sandbox for its tenants needs both: the tenant adds env vars (and nothing else), and local execution is neither visible nor reachable > - This pull request opens exactly one tenant edit (env vars) on the managed sandbox row, and adds an `enableManagedSandboxOnly` mode that hides local and makes run selection fail closed > - The benefit is a complete managed-sandbox experience with no change for self-hosted instances ## Linked Issues or Issue Description **Subsystem affected** Environments (managed sandbox provisioning, environment routes, run environment selection) and the environments UI. **Problem or motivation** Platform-provisioned sandbox environments (`metadata.managedByPaperclip`) reject every write on cloud-managed instances. Agents often need environment variables inside their sandbox. The tenant has no way to set them on the managed row. Separately, an operator cannot remove local execution: the environment list always shows the local row, and run selection falls back to it when no default is set. **Proposed solution** Allow an envVars-only PATCH on the managed sandbox row, and echo those env vars back for editing. Add a managed-tier feature (`enableManagedSandboxOnly`) that hides the local environment from all read surfaces and redirects local-landing run selection to the managed sandbox environment, failing closed when it is unavailable. **Alternatives considered** UI-only hiding of the local row. This was rejected: it does not stop a run from resolving to local, so it is presentation without enforcement. Full unlock of the managed row was also rejected: name, driver, and config stay platform-owned so boot reconciliation cannot fight tenant edits. ## What Changed - `server/src/routes/environments.ts`: the platform-provisioned write floor admits an envVars-only PATCH on the generalized managed sandbox row (sandbox driver, `managedByPaperclip`, not legacy kubernetes-marker rows). Name, driver, config, status, metadata, and DELETE stay rejected. The read floor stops blanking env vars on that row; credential-shaped config keys stay redacted for every actor. Legacy kubernetes-marker rows keep the full floor. - Same file: under `enableManagedSandboxOnly`, the environments list and the by-id read omit the local row for every actor, including instance admins. - `server/src/services/execution-workspace-policy.ts`: `resolveExecutionWorkspaceEnvironmentId` gains the managed-sandbox-only inputs. A selection that lands on the local environment is redirected to the managed sandbox environment. With no active managed row it throws `ManagedSandboxUnavailableError` — never local. Non-local selections (ssh, user-created sandboxes) are untouched. - `server/src/services/heartbeat.ts`: the run path reads the flag, looks up the managed row (`findManagedSandboxEnvironment`, new read-only finder in `environments.ts`), and passes both to the resolver. Mirrors the forced-kubernetes precedent, which keeps precedence when both regimes are on. - `server/src/services/managed-environments.ts`: after a successful reconcile, the instance default environment moves to the managed sandbox row when the current default is unset, local, or dangling. A tenant-chosen custom environment is never overridden. - `packages/shared`: new `enableManagedSandboxOnly` key (schema default false, catalog tier `managed`, cloudDefault false, selfHostedDefault false) and the matching interface field. - UI: managed rows show a "Managed by Paperclip" lock badge; editing one opens a dedicated env-vars-only editor that sends the one PATCH shape the server admits (the old full form failed with a 403 on save). New `ui/src/lib/managed-sandbox-environment.ts` mirrors the local filter for cached lists (applied in the project picker; the agent picker already excluded local). The experimental settings page gains the toggle at its alphabetical card position. `environmentsApi.update` now declares the `envVars` field it already sent. - Tests: environment route floor coverage (envVars-only accepted, mixed bodies rejected, legacy rows still blanked and locked, local hidden and 404 under the flag, self-hosted unchanged), an embedded-postgres service test pinning that boot reconciliation never touches tenant env vars, resolver redirect/fail-closed cases, managed-environments default-stamping cases, and UI lib/settings tests. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/environment-routes.test.ts src/__tests__/environment-service.test.ts src/__tests__/execution-workspace-policy.test.ts src/services/managed-environments.test.ts` — all pass. - `pnpm --filter @paperclipai/shared exec vitest run` — 425 pass (catalog/schema default parity is pinned by an existing test). - `pnpm --filter @paperclipai/ui exec tsc --noEmit` and the affected UI suites (CompanyEnvironments, InstanceExperimentalSettings incl. card-order test, new lib test) — all pass. - Full workspace `pnpm test`: 3,414 passed. 17 files report failures on this machine; the identical 17 fail on a clean `origin/master` worktree in the same environment (git-worktree/skills/embedded-postgres environment dependencies and plugin-SDK zero-test collections). One additional file (`issue-monitor-scheduler.test.ts`) failed one timing-sensitive test in one of two full-suite runs and passes 7/7 in isolation on this branch — a flake in a domain this diff does not touch. The branch introduces no new failures. - Self-hosted zero-delta: every new behavior is gated on the cloud-managed instance check or the new flag, which defaults to false in schema and catalog; pinned by the "does not floor platform-marked rows on self-hosted instances" and flag-off tests. ## Risks - Behavior is opt-in twice over: the write-floor exception applies only to rows the managed-config provisioner stamps, and the hiding/forcing applies only when `enableManagedSandboxOnly` is on (default false everywhere). Self-hosted instances see no change. - The env-vars echo is scoped to the generalized managed sandbox row; legacy kubernetes-marker rows keep the blanket floor because pre-generalization builds may have written platform values there. - Fail-closed run selection means a managed instance with the flag on and an archived managed row (provider plugin down) refuses runs with a precise error instead of running locally. That is the intended posture. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
7ea2068ef8 |
fix(files): only highlight accessible workspace file links (#11090)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task comments can contain references to files in project and execution workspaces. > - Paperclip detected path-shaped inline code and showed it as an actionable file chip. > - The UI did not first confirm that the current board session could open the file. > - Missing, denied, ambiguous, remote, and unsupported files therefore looked actionable and failed after a click. > - This pull request adds an issue-scoped availability check and promotes only confirmed files to chips. > - The benefit is that the task thread shows a file action only when that action can succeed. ## Linked Issues or Issue Description **What happened?** Task comments promoted path-shaped inline code to file chips before Paperclip checked the file. A chip could point to a missing, denied, ambiguous, remote, or non-previewable file. The action then failed after the user selected it. **Expected behavior** Paperclip must show a file chip only after the server confirms that the current board session can open the exact file reference. All other path-shaped text must stay ordinary inline code. **Steps to reproduce** 1. Add a task comment that contains inline code with a missing or inaccessible workspace path. 2. Open the task thread as a board user. 3. Observe that the path looks like an actionable file chip. 4. Select the chip and observe that the file cannot open. **Paperclip version or commit** `19be4cf927` and earlier. **Deployment mode** Local dev and self-hosted server. **Access context** Board user. ## What Changed - Added shared request, response, and validation contracts for batched workspace-file availability checks. - Added an issue-scoped server endpoint that resolves file references with company, issue, workspace, and preview-access checks. - Added bounded batch concurrency and tests for missing, denied, ambiguous, remote, unsupported, and available files. - Added an issue-scoped UI availability registry that deduplicates, batches, caches, and invalidates file checks. - Changed task-comment markdown rendering so only confirmed files get chip styling and file-viewer behavior. - Bound each chip to the exact workspace target that passed the availability check. ## Verification - `pnpm exec vitest run packages/shared/src/workspace-file-resource.test.ts server/src/__tests__/file-resources.test.ts ui/src/components/MarkdownBody.test.tsx ui/src/components/WorkspaceFileMarkdownBody.availability.test.tsx ui/src/lib/remark-workspace-file-refs.test.ts ui/src/lib/workspace-file-availability.test.ts` — 93 passed, 35 skipped. - `pnpm check:token-gates` — clean. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm test:run` — all server and UI groups passed. One unchanged CLI test saw the run-injected static AWS credentials and expected only its local `AWS_PROFILE`. The same test passed, 8 of 8, after removing only `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` from its process environment. ## Risks - File chips now appear after an asynchronous availability check, so path-shaped text can briefly render as inline code. - Availability results use the existing 30-second file-resource cache window. File-resource invalidation forces a new check. - The endpoint limits each request to 100 references and the client chunks larger sets. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The service did not expose a more specific model ID or context-window size. The agent used high-reasoning mode, repository tools, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
b58ce27a02 |
fix: isolate execution workspace summaries (#10790)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip gives operators a summary for each workspace. > - An execution workspace detail page used the parent project-workspace summary slot. > - Two execution workspaces under one project workspace could therefore show the same summary. > - This pull request gives each execution workspace its own summary scope. > - It also limits the summary snapshot and generated issue to that execution workspace. > - The benefit is that a new or parallel execution workspace cannot inherit unrelated status. ## Linked Issues or Issue Description **What happened?** An execution workspace detail page read and refreshed the summary slot for its parent project workspace. Parallel execution workspaces could show the same status and include issues from each other. **Expected behavior** Each execution workspace must have one isolated summary slot. Its generated snapshot must include only issues assigned to that execution workspace. **Steps to reproduce** 1. Create two execution workspaces under one project workspace. 2. Add different issues to each execution workspace. 3. Generate the summary in the first execution workspace. 4. Open the second execution workspace. 5. Observe that the old implementation could reuse the first summary. **Paperclip version or commit** The problem exists on `master` before this pull request. **Deployment mode** The issue affects both local trusted and authenticated deployments. ## What Changed - Added `execution_workspace` to the shared summary-slot scope contract. - Validated execution-workspace ownership and stored generated summary issues on the correct execution workspace. - Limited execution-workspace snapshots to issues with the matching execution workspace ID. - Updated the execution workspace page to use its own summary slot. - Updated Summarizer instructions, routine options, catalog metadata, documentation, and regression tests. ## Verification - `NODE_ENV=test pnpm exec vitest run packages/shared/src/summary-slot.test.ts server/src/__tests__/summary-slots.test.ts ui/src/pages/ExecutionWorkspaceDetail.test.tsx` — 30 focused tests passed; the embedded-Postgres server tests were run outside the process-restricted sandbox. - `pnpm check:token-gates` — passed. - `pnpm --filter @paperclipai/skills-catalog validate` — passed with 17 catalog skills. - [Latest-head GitHub Actions](https://github.com/paperclipai/paperclip/actions/runs/31491475405) — all 22 jobs passed on `beea14cbaf`, including typecheck, build, server/workspace tests, serialized suites, e2e, canary, and aggregate verification. One unrelated adapter cleanup test initially hit an `ENOTEMPTY` temp-directory race; its single permitted rerun passed. - Greptile — 5/5 confidence on `beea14cbaf`, 12 files reviewed, zero comments added, and zero unresolved threads. ## Risks - Low risk. The new scope is additive. - Existing project and project-workspace summary slots keep their current keys and behavior. - A summary generated for an execution workspace now excludes sibling workspace issues by design. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The deployment does not expose a more specific model ID or context-window value. It used agentic reasoning, repository tools, code execution, and GitHub tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
35aaaa0bd0 |
feat(server): preserve task timestamps and hierarchy through company import/export (#11193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company export/import moves a whole company — agents, tasks, comments — between instances as a portable bundle > - The bundle never carried task timestamps or parent links: the export writes neither, the importer lets database defaults stamp "now", and sub-tasks arrive flattened > - Boards sort by recency, so every imported task showing "created just now" collapses the task list into import order, and the task hierarchy the user built is gone > - This pull request adds created/updated/started/completed/cancelled timestamps and a parent link to the bundle (schema v7), preserves them end to end on import, and keeps comment imports from clobbering a preserved updated time > - The benefit is that an imported company reads like the company the user left: same recency order, same task tree ## Linked Issues or Issue Description **What happened?** After a company import, every task showed as created at import time. Recency sorting collapsed to import order, and parent/child task nesting disappeared. The user called out losing "the meaningful task hierarchy and recency sorting". Cause: the export bundle has no fields for task timestamps or parent links, the importer lets `defaultNow()` win on insert, and the comment importer bumps every touched task's `updatedAt` to now. **Expected behavior** An imported company preserves each task's creation/update/start/completion times and its position in the task tree, so sorting and nesting on the destination match the source. **Steps to reproduce** 1. On a source instance, create tasks over several days, including sub-tasks nested under parents. 2. Export the company and import it into another instance. 3. Every task shows the import moment as its creation/update time and all tasks are top-level. ## What Changed - Export writes `createdAt`/`updatedAt`/`startedAt`/`completedAt`/`cancelledAt` (ISO, only when set) and `parent: <taskSlug>` into each task's bundle extension; a parent outside the export selection drops the edge with an aggregate warning, mirroring the existing blocker-edge warning (`server/src/services/company-portability.ts`). - Bundle schema version 6 → 7. All new fields are optional: v5/v6 bundles import unchanged with a version-aware downlevel warning; bundles newer than the board still fail closed. - Manifest parsing validates the new timestamps like comment timestamps (invalid → warn and ignore, never a hard failure); shared types and the zod validator carry the new optional fields. - Import resolves parent slugs to pre-generated destination ids, drops self-references and cycles from tampered bundles with warnings, and orders rows parents-first because the self-referencing FK is checked per insert chunk. - `importIssues` writes the preserved timestamps (falling back to insert time when absent; `startedAt` stays null unless bundle-carried, per #11191's semantics) and `parentId`. - `addImportedComments` no longer blanket-bumps `updatedAt = now()`; it takes `GREATEST(updated_at, newest imported comment createdAt)`, so a preserved update time never regresses while unpreserved rows keep the old behavior. ## Verification - `pnpm vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts server/src/__tests__/productivity-review-service.test.ts` — 102 passed, 1 pre-existing opt-in benchmark skip. Includes: full round-trip with exact timestamp equality and a 3-deep parent chain against embedded Postgres; v6 back-compat (defaults + warning); forward-compat rejection (v8); cycle/self-reference/invalid-timestamp tampered-bundle handling; comment-bump preserve-awareness in both directions. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/shared typecheck` — clean. ## Risks - **Rollout ordering**: a board on the previous build (max schema v6) refuses bundles exported by this build (stamped v7) — the existing newer-than-supported rejection, working as designed. Cross-instance moves need the importing board upgraded first. Called out here so operators aren't surprised during the transition window. - Parent edges from tampered bundles are dropped with warnings rather than failing the import; blocker relations already behave this way. - Timestamps are data-only; no destination schema migration. Stacked on #11191 (its commit is included here) — merge #11191 first; this PR then shows only the v7 changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
459469638e |
feat(adapter-codex-local): add secure device-login building blocks (#11097)
## Thinking Path > - Paperclip connects AI agents to local and remote runtimes. > - The Codex local adapter needs a safe device-login flow. > - A future sandbox integration needs strict prompt validation, secret protection, cleanup, and private credential storage. > - This pull request adds tested building blocks for that flow. > - The result gives a later Daytona integration a clear security boundary. ## Linked Issues or Issue Description No public issue covers this change. **Problem or motivation** The Codex local adapter has no safe, reusable flow to prove device login inside an isolated sandbox. **Proposed solution** Add parser, runner, credential export, and proof helpers. Validate the prompt, protect login data, store credentials in a private run-scoped home, and dispose all sandbox resources. **Alternatives considered** Do not connect a production Daytona driver in this change. Use an injected sandbox driver and focused tests first. This keeps the security controls testable before live provider integration. **Roadmap alignment** The change extends the Codex local adapter. It does not add a core Paperclip route or duplicate a planned core feature. **Additional context** The flow keeps the login URL, code, and token out of logs, results, and errors. The proof home uses a company-scoped root and a run-scoped private directory. ## What Changed - Add a pure parser for the exact Codex device-login URL and one-time code shape. - Add a sandbox runner with prompt handling, timeout, cancellation, and disposal. - Add a credential export step with company scoping, path checks, payload checks, private modes, locking, and cleanup. - Add redacted device-login fixtures and focused tests for parsing, secret redaction, runner outcomes, credential export, and cleanup. ## Verification - Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`. - Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc --noEmit`. - Review tests for strict URL and code validation, timeout, cancellation, disposal, secret redaction, path safety, payload safety, file modes, and cleanup. ## Risks - This change provides building blocks, not a live Daytona proof. - A later integration must connect the runner to a concrete sandbox driver. - Credential export depends on existing Codex authentication cache helpers. - Incorrect path or payload assumptions can reject valid credentials. ## Model Used Codex, GPT-5, tool use, code execution, and repository review. The Paperclip runtime controls the exact context window and reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5ca752dc81 |
fix(server): raise company import zip upload limit to 1 GB and make it operator-configurable (#11184)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import/export lets an operator move a full company package between instances, with the Import page uploading the package as one compressed `.zip` > - The server caps that upload at 128 MB, and real company packages with attachments now exceed it — imports fail at the preview step > - The failure message tells the user to use the CLI folder import, but that path posts inline JSON capped at 64 MB, so the advice is a dead end for exactly these packages > - This pull request raises the zip upload cap to a 1 GB default, makes it operator-configurable through an environment variable, scales the decompression-bomb guards from the cap in effect, and replaces the misleading hint > - The benefit is that large real-world company packages import successfully, and operators with unusual needs can tune the cap without a code change ## Linked Issues or Issue Description **What happened?** A company import fails at the preview step with `Preview failed: Import package exceeds 134217728 bytes`. The package is a valid Paperclip export. Its compressed size is larger than the 128 MB server cap (one reported package is 257 MB). The error panel suggests the CLI folder import, but that path sends the package as one inline JSON body capped at 64 MB, so it also fails. **Expected behavior** A valid company package of realistic size imports successfully through the Import page. If a package is too large, the error must state the limit clearly and suggest a step that can work. **Steps to reproduce** 1. Export a company with enough attachments to make the compressed package larger than 128 MB. 2. Open the Import page and upload the `.zip`. 3. Click "Preview import". 4. The preview fails with `Import package exceeds 134217728 bytes`. **Deployment mode** Reported from a managed deployment; the limit applies to all deployment modes. ## What Changed - Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default (`server/src/http/body-limits.ts`). - Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid or non-positive values fall back to the default. - Scale the zip decompression-bomb guard from the configured cap at the import route: the aggregate inflated ceiling is 4x the cap. The per-entry ceiling stays at 512 MB because V8's string length limit applies to an entry regardless (`server/src/routes/companies.ts`, `packages/shared/src/portability-zip.ts`). - Report the 422 limit error in MB instead of raw bytes. - Replace the "use the CLI folder import for very large packages" hint on preview failure with advice that works: re-export the package without large attachments (`ui/src/pages/CompanyImport.tsx`). - Update the stale comment in `ui/src/lib/import-preflight.ts` that made the same CLI claim. - Add tests for the new default, the env override, and the invalid-override fallback. ## Verification - `pnpm vitest run server/src/__tests__/body-limits.test.ts packages/shared/src/portability-zip.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — all pass. - `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes, including the updated failure-panel copy assertion. - `pnpm typecheck` — clean across the workspace. - Manual: upload a `.zip` larger than the configured cap; the preview fails with `Import package exceeds the 1024 MB upload limit` and the new hint. A package between 128 MB and 1 GB now previews and imports. ## Risks - Peak per-import memory rises with the cap: the upload is buffered in memory and unzipped in one pass. A 1 GB compressed package can use several GB transiently. Imports are instance-admin actions, so the exposure is a deliberate operator action, not anonymous traffic. Operators on small hosts can lower the cap with `PAPERCLIP_IMPORT_ZIP_MAX_BYTES`. - The aggregate bomb guard moves from a fixed 512 MB to 4x the configured cap. It still bounds expansion far below what a decompression bomb needs. - No migration and no API shape change. The 422 message text changes; no code matches on the old text. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (file edits, local test runs, live-instance inspection). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
cc35c3c395 |
feat: structure and humanize recovery notices (#11075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip posts system comments when automatic run recovery cannot continue > - These comments currently mix the main event with recovery identifiers and routing details > - The task chat shell also renders these comments as large raw text blocks > - Operators need a short explanation first and inspectable evidence on demand > - This pull request emits structured recovery notices and renders them as compact humanized rows > - The benefit is a quieter task thread that keeps the full recovery evidence available ## Linked Issues or Issue Description Related prior extraction source: #11070. This pull request replaces only its structured recovery notice slice with a focused branch based on current master. **What existing behavior does this improve?** Paperclip recovery escalations and the experimental task chat system-comment renderer. **Current behavior** Recovery escalation comments put action identifiers, owner details, run details, and failure codes into the visible markdown body. The task chat shell renders the complete system comment as a large text block. **Proposed behavior** The server emits a short system notice with typed metadata sections. The task chat shell classifies known recovery families and renders one compact row. An operator can expand the row to inspect the full body and metadata. **Reason and benefit** The main thread stays readable during repeated recovery activity. Typed links and evidence remain available without exposing raw failure text in the default view. **Breaking changes** The visible recovery comment body is shorter. Recovery action deduplication now reads the structured metadata and still recognizes legacy body markers. No API schema or database migration changes. ## What Changed - Emit stranded recovery escalations with `system_notice` presentation and typed recovery, owner, run, and failure-code metadata. - Share bounded metadata row builders across recovery notice producers and preserve legacy deduplication compatibility. - Humanize known recovery notice families and render compact expandable task-chat rows. - Route system-authored comments ahead of derived agent authorship so recovery notices do not appear as agent bubbles. - Add focused server and UI regression coverage. ## Verification - `pnpm check:token-gates` — 3/3 clean. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/shared exec vitest run src/validators/issue.test.ts` — 32 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/stranded-notice.test.ts src/__tests__/issue-recovery-actions.test.ts` — 57 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/successful-run-handoff.test.ts src/services/recovery/stranded-notice.test.ts` — 39 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'escalates an exhausted failed successful-run handoff without using generic continuation recovery first|escalates an exhausted successful handoff run that still leaves no disposition'` — 2 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'blocks assigned todo work after the one automatic dispatch recovery was already used'` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/system-notice-humanizer.test.ts src/components/task-chat/TaskChatSystemNotice.test.tsx src/components/task-chat/task-chat-adapter.test.ts` — 15 tests passed. - Storybook visual baselines were not updated because this chat-shell path has no affected snapshot baseline. Focused rendering tests and token gates cover this change. ## Risks - Consumers that parse recovery action identifiers from comment markdown must move to structured metadata. Server deduplication remains backward compatible with legacy comments. - The humanizer uses stable recovery-family phrases. Unknown notices use a generic truncated first-sentence fallback. - The UI changes only the experimental task chat presentation. The stored comment body and expanded metadata remain available. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5, reasoning mode, repository tools, shell execution, and GitHub integration. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d5208d30c1 |
feat(adapter-utils): add host-side pack span to managed-runtime tarball build (#11072)
## Thinking Path > - Paperclip runs AI agents through adapter execution services. > - The adapter runtime records spans for each stage of agent startup. > - The managed runtime packs workspace tarballs before it uploads them. > - These pack operations had no host span, so `stage.sync` omitted pack time. > - This pull request adds one `pack` span around both tarball builds. > - The span nests under the active `stage.sync` step and improves trace detail. ## Linked Issues or Issue Description **What existing behavior does this improve?** The managed runtime workspace sync builds a git-history tarball and a workspace-overlay tarball before upload. **Subsystem affected** The change affects `packages/adapter-utils`, which provides adapter execution and managed runtime support. **Current behavior** The host builds both tarballs without an OpenTelemetry span. The `stage.sync` trace therefore omits the host pack duration. **Proposed behavior** The host wraps both tarball builds in one `pack` span. The executor parents this span under the active startup step. **Reason and benefit** The trace shows the time that the host spends packing workspace data. Operators can use the existing runtime span tree to find sync delays. **Breaking changes** None. The default span runner remains a no-op runner, and the existing control flow remains unchanged. ## What Changed - Add an optional `runtimeSpan` runner to the managed runtime preparation path. - Create one host `pack` span around the two workspace tarball builds. - Parent the `pack` span under the active startup step. - Add unit coverage for span emission and span nesting. ## Verification - Run `tsc --noEmit` for `@paperclipai/adapter-utils`. - Run the `@paperclipai/adapter-utils` Vitest suite. - Confirm the suite reports 448 passed tests and 4 skipped tests. - Confirm the trace test records `pack` under `stage.sync`. ## Risks The change adds optional tracing only. The default no-op runner preserves behavior when tracing is not configured. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The model reviewed the handoff and opened this pull request. The implementation author supplied the commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6b7e0814a0 |
feat(acp): stream Daytona sandbox agent output and remove the host output poll (#11049)
## Thinking Path > - Paperclip is the open source app that manages AI agents for work > - Sandbox providers let agents run in remote and isolated environments > - Daytona session commands need a path that sends agent output to the host without host polling > - Host polling adds delay and repeats provider output work > - This pull request adds typed execute.log notifications and a log sink for incremental output > - This pull request adds an optional ACP session stream with final-result replay protection > - The benefit is lower output delay while the default flags keep current behavior unchanged ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. The change spans the plugin SDK, Daytona provider, adapter utilities, and server execution services. **Problem or motivation** The Daytona ACP bridge polls a host output file while an agent command runs. This adds delay and can repeat work. The host also needs a safe route for provider output chunks. **Proposed solution** Add a typed `execute.log` notification with host-issued invocation correlation. Add an ordered log sink to the environment execute path. Add an optional ACP session-log path that parses newline-delimited JSON frames and removes the host output poll for that path. **Alternatives considered** Keep the output-file poll as the only path. This keeps the current behavior but does not provide timely output. The new path stays behind flags, so the existing path remains the default fallback. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`, including Daytona support. ## What Changed - Add the typed `execute.log` worker-to-host notification and company-scoped host route. - Add ordered `stdout` and `stderr` chunk delivery before the final execute result. - Add the Daytona session log sink and the optional ACP streamed session path. - Add monotonic frame handling so live and final output reach the host once. - Keep `useLogStream` and `streamAgentSessionOutput` off by default. - Add unit and integration coverage for the notification, execution target, runtime, and Daytona paths. ## Verification - Run adapter-utils tests: 445 tests pass locally. - Run server environment tests: 73 tests pass locally. - Run Daytona plugin tests: 131 tests pass locally. - Run TypeScript checks for shared, adapter-utils, and server. - Review the pull request checks after GitHub completes them. - All required GitHub checks pass on the current head. ## Risks The new paths change output delivery only when a feature flag enables them. The final execute result remains available for parsing and fallback. The main risk is a provider stream or frame-order error; the final-result parser limits that risk. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The runtime did not supply a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes No operator documentation change applies because both new flags remain disabled by default. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a511ed1b0 |
feat(apps): support multiple provider connections (#11060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem connects company tools through governed provider connections. > - A company can need more than one account for the same provider. > - The current database constraint and Apps flow assume one named connection per company. > - New quarantined actions also need an explicit review decision before activation. > - This pull request supports multiple provider connections and complete action review decisions. > - The benefit is safer access control and a clear multi-account Apps workflow. ## Linked Issues or Issue Description Refs: #11040 **Subsystem affected** Cross-cutting. This change affects the Apps UI, the tool access API, the shared request contract, and the database schema. **Problem or motivation** The connection name constraint prevents a company from keeping more than one connection for a provider. The Apps UI also reuses an existing OAuth connection when a user asks to connect another account. Action review can enable selected entries without recording a decision for every quarantined action. **Proposed solution** Remove the company and connection name uniqueness constraint. Let users open, count, edit, and create multiple provider connections. Require the finish request to cover every quarantined action exactly once before the server activates reviewed entries. **Alternatives considered** The UI could generate unique internal names and keep the database constraint. This would preserve a one-connection assumption in the data model and would make display names part of identity. The server could also infer review decisions from enabled actions. This would not distinguish a reviewed disabled action from an action that the user did not review. **Roadmap alignment** This change extends the completed MCP Tool Gateway and Apps milestone. It also supports the Connected Apps roadmap item. It follows the navigation and connection management work in #11040. ## What Changed - Remove the company-scoped connection name uniqueness index with an ordered and idempotent migration. - Add a reviewed action list to the finish-app contract and reject incomplete or duplicate review decisions. - Activate reviewed entries and keep unreviewed quarantined entries blocked. - Enable a completed connection and preserve the company and connection scope in all updates. - Show provider connection counts and open the provider setup page from Browse. - Let users edit existing connections or connect another account without reusing an active OAuth connection. - Update focused server and UI coverage for multiple connections and action review. ## Verification - Ran the focused Apps UI suite. All 116 tests passed in 11 files. - Ran the focused server and CLI suite. All 276 tests passed in 3 files. - Ran `pnpm --filter @paperclipai/db check:migrations`. The migration safety check passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One worktree-safety assertion failed because the execution workspace reloads its worktree marker. The same test passed with an isolated non-worktree marker. - Ran `pnpm check:token-gates`. It reports 12 existing violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. - Started the six affected Playwright specifications. Chromium could not start because the host does not provide `libatk-1.0.so.0`. The GitHub e2e jobs will verify these specifications. - GitHub Actions passed every final-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed final commit `9af9200426` at 5/5 with zero review threads. ## Risks - Removing the name uniqueness index permits duplicate display names. Stable connection IDs and UIDs remain unique within a company. - The finish-app endpoint accepts the new review field as optional for backward compatibility. When clients send it, the server requires a complete decision for all quarantined actions. - Multiple OAuth connections depend on the explicit new-connection route flag. Focused tests cover active and draft connection reuse. - The migration is ordered after migration 0210. Its `DROP INDEX IF EXISTS` statement is safe to repeat. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The agent used repository tools, code execution, test execution, and agentic reasoning. The Codex runtime manages the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a71b9cf628 |
feat(skills): add MCP integration preparation skill (#11063)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Skills Store lets a company find and install reusable agent procedures > - The MCP integration preparation procedure existed outside the app catalog > - Paperclip users could not find or install that procedure from the product > - This pull request adds the procedure as an optional software development skill > - The skill keeps research, human approval, and connector delivery as separate gates > - The benefit is a repeatable and governed path from vendor research to one connector pull request ## Linked Issues or Issue Description **What existing behavior does this improve?** The app-shipped Skills Store can install optional skills, but it does not include the MCP integration preparation workflow from `paperclip-content`. **Subsystem affected** `packages/skills-catalog`. **Current behavior** An agent must know where the external workflow lives. The agent cannot find or install it from the Paperclip skills catalog. **Proposed behavior** The optional catalog includes `prepare-mcp-integration`. The installed skill directs agents through cited research, a research-only content pull request, an exact-revision human gate, and one Paperclip connector pull request per approved connection. **Reason and benefit** This change makes the existing integration and connector playbooks available as one installable Paperclip workflow. It also prevents agents from starting connector code before the research gate is approved. **Breaking changes** None. The skill is optional and markdown-only. Related source work: paperclipai/paperclip-content#13. ## What Changed - Add the optional `prepare-mcp-integration` catalog skill under software development - Add Paperclip catalog metadata for roles, requirements, tags, and trust classification - Add a worked Notion MCP research-gate example - Regenerate the checked-in catalog manifest - Add the new key to the shipped optional skill test ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog validate` - `pnpm --filter @paperclipai/skills-catalog test` - Confirm the generated catalog contains `paperclipai/optional/software-development/prepare-mcp-integration` - Confirm the trust level is `markdown_only` and compatibility is `compatible` ## Risks - Low risk. This change adds one optional markdown-only catalog entry. - The workflow can become stale if the two upstream playbooks change. The skill requires agents to read the current playbooks before each phase and to update upstream rules when reusable specifications change. ## Model Used OpenAI Codex with GPT-5.4, reasoning mode, shell tool use, and code editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e76227f12 |
feat(plugin-daytona): stream session command logs behind useLogStream (#11021)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox provider plugins let agents run commands in remote environments. > - The Daytona provider polls the exit code while it waits for command logs. > - Polling delays log delivery and does not support long-lived streamed commands. > - This pull request adds an opt-in Daytona log stream with one reconnect and a poll fallback. > - The benefit is faster log delivery while the existing default path stays unchanged. ## Linked Issues or Issue Description Refs: #10941 **Subsystem affected** packages/plugins — sandbox provider plugins. **Problem or motivation** The Daytona provider polls the command exit code every 50 milliseconds while it waits for logs. This delays output and does not support a long-lived streamed command. **Proposed solution** Add the `useLogStream` provider option. Stream stdout and stderr from the Daytona callback log form, read the exit code after the stream ends, retry the read with bounded backoff, and fall back to the existing poll path after a disconnect. **Alternatives considered** Keep polling for all commands. This keeps the current behavior but does not provide timely logs or a path for long-lived commands. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents, including Daytona. ## What Changed - Add the opt-in `useLogStream` option with a default of `false`. - Stream stdout and stderr from the Daytona callback log form. - Drop replayed log prefixes by delivered byte offset after reconnect. - Retry once after disconnect, then use the existing poll path. - Read the exit code once after a successful stream and retry when the code is not ready. - Add tests for ordered output, exit-code reads, disconnect fallback, and reconnect replay handling. ## Verification - Run `node_modules/.bin/vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`. - Confirm that the full Daytona plugin test project passes with 127 tests. - Confirm that the changed files type-check against Daytona SDK 0.203.0 types. - Review the PR against parent PR #10941 before it reaches `master`. ## Risks - The stream path changes behavior only when `useLogStream` is `true`. - A stream failure can add one reconnect attempt before the existing poll fallback. - The stream path has no command deadline because it supports long-lived commands. - No endpoint, stored data, telemetry shape, authentication rule, or result shape changes. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The context window and reasoning mode are not exposed by this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9485ffea70 |
fix(config): preserve env files during managed updates (#10980)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The CLI and server both update Paperclip values in `.env` files > - The server preserved operator content, but the CLI rebuilt the complete file > - A CLI rerun could remove comments, custom values, ordering, and newline style > - Both paths need one editor with one value encoding and duplicate key policy > - The final integration also needs one regression test across the related setup and sync safety mechanisms > - This pull request moves the editor to the shared package and adds cross-cutting rerun-survival coverage > - The benefit is safe setup and worktree repair reruns that preserve operator edits ## Linked Issues or Issue Description **What happened?** The CLI rebuilt the complete `.env` file when it wrote a managed Paperclip value. This action removed comments, blank lines, custom keys, original quoting, and the original newline style. **Expected behavior** Paperclip must update only the managed assignments. It must preserve all unrelated bytes. It must skip the file replacement when all managed values are current. **Steps to reproduce** 1. Add comments, custom keys, quoted values, and CRLF newlines to the Paperclip `.env` file. 2. Run a CLI path that calls the agent JWT secret setup. 3. Observe that the old writer replaces the complete file. **Paperclip version or commit** The problem exists on `master` before this pull request. Related public context: Refs #437. ## What Changed - Add one shared line-preserving `.env` editor for the CLI and server. - Define minimal and JSON value encodings in the shared helper. - Update every stale duplicate of a managed key and preserve current duplicate encodings. - Preserve comments, ordering, blank lines, unknown keys, export prefixes, trailing comments, and newline style. - Write changed files through a same-directory temporary file and atomic rename. - Limit CLI updates to non-empty `PAPERCLIP_*` entries. - Skip the write when all managed values are current. - Add shared, CLI, and server regression coverage. - Refresh the branch after the related config, sandbox, and skill safety changes landed. - Add a cross-cutting integration test for config, env-file, managed-sandbox, and managed-instructions rerun survival. ## Verification - `pnpm exec vitest run packages/shared/src/env-file.test.ts packages/shared/src/config-schema.test.ts cli/src/__tests__/agent-jwt-env.test.ts cli/src/__tests__/config-store.test.ts server/src/__tests__/config-file.test.ts server/src/__tests__/worktree-config.test.ts` passes 39 tests. - `pnpm exec vitest run server/src/__tests__/rerun-survival.integration.test.ts` passes 4 tests. - `pnpm -r typecheck` passes on the previous head. GitHub CI reruns it on the refreshed head. - The previous head passed the complete general, serialized, workspace, and E2E matrix. GitHub CI reruns that matrix on the refreshed head. - `pnpm build` passes on the previous head. GitHub CI reruns it on the refreshed head. ## Risks - Low risk. The production change only affects managed `.env` assignments. - Existing managed assignments can keep their original quoting when their decoded values are current. - Changed CLI values keep the prior minimal encoding policy. Changed server values keep the prior JSON encoding policy. - Duplicate managed assignments now follow one explicit rule: Paperclip updates each stale occurrence. - The master refresh had one import-block conflict. The resolution keeps both the config merge imports and the env-file imports. - The added integration file is test-only. It has no database, API, or UI contract effect. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family produced this change with reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |