mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
fdbc69172d6e1c8aa77f0575aad9731fdbb4f015
1110
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fdbc69172d |
feat(runner): add PRP v1 schemas and fixtures (#12087)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner needs a language-neutral contract between the server and the runner process. > - A shared contract must exist before TypeScript, Rust, transport, or provider implementations can depend on it. > - Required protocol versions must fail closed, while safe optional fields must remain compatible. > - The contract also needs deterministic fixtures and a drift gate for later cross-language work. > - This pull request adds that contract without adding runtime behavior. > - The benefit is a small, reviewable source of truth for the next implementation pull requests. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This pull request adds a private package contract for later server, TypeScript, and Rust work. **Problem or motivation** Paperclip Runner does not have a small language-neutral protocol boundary on `master`. A runtime implementation without this boundary can drift between languages, accept unsupported required versions, or silently change canonical fixtures. **Proposed solution** Add PRP v1 JSON Schemas, accepted and rejected fixtures, a Codex structured-question fixture, and a generated SHA-256 manifest. Run compatibility and manifest checks during the package build. Keep the package private and export nothing in this pull request. **Alternatives considered** The combined runner branch contains schemas together with providers, SDKs, labs, and server behavior. That change is too large for normal review. Generating TypeScript validators in this pull request would also cross into the next review unit. **Roadmap alignment** This contract supports the governed tool and control-plane direction in `ROADMAP.md`. It does not enable a new production adapter or endpoint. **Additional context** Refs #12084 and #11962. This pull request was reviewed as a stack on #12084, then rebased and retargeted to `master` after #12084 merged. The current delta is 38 files. ## What Changed - Added 20 PRP v1 JSON Schemas with stable identifiers and resolved references, including explicit cross-language conformance input and output schemas. - Added canonical replay, cross-language, and Codex question fixtures. - Added accepted cases for additive optional fields and a rejected case for an unsupported required protocol version. - Added a deterministic manifest with SHA-256 digests for every schema and fixture. - Added package-local schema-instance, schema-reference, compatibility, question-ID, conformance-pair, and drift checks. - Added a private workspace package with no public exports and no production runtime behavior. - Added the package manifest to the Docker dependency-stage inventory required for every workspace package. This does not copy or build runner runtime code into the production image. - Kept the provider descriptor and question fixture Codex-only. No deferred provider package or dependency is present. ## Verification - `pnpm install --frozen-lockfile` passed with Node 24.19.0 and pnpm 9.15.4. No lockfile change is committed. - `pnpm --filter @paperclipai/paperclip-runner check:protocol` passed with 8 tests. - The committed AJV 2020-12 gate accepted every canonical v1 replay, question, and cross-language conformance fixture. It rejected the required v2 fixture, a replay fixture with a missing required command ID, and conformance output with a missing session ID. - `pnpm -r typecheck` passed. - `pnpm build` passed and ran the protocol manifest drift check. - `pnpm check:token-gates` passed. - `node ./scripts/check-docker-deps-stage.mjs` passed. - `git diff --check` passed. - The delta against its declared base is 38 files. - `pnpm test:run` completed with 4,687 passing tests, 19 skipped tests, and 29 failures across 9 unchanged server files. The failures reproduce macOS path aliases, local listener probes, workspace-runtime assumptions, and one connection-retry timeout. No changed-file test failed. Linux CI must pass before this pull request is ready. - `pnpm check:tokens` reports existing personal-name references outside this pull request. A scoped scan of `packages/paperclip-runner` found no secret-like values, internal references, or deferred-provider names. - PR #12084 was squash-merged, and this branch was rebased onto that merge and retargeted to `master`. The first master-base policy run correctly caught the missing Docker dependency-stage manifest copy; commit `4fa1ea7c` fixes that gate, and the complete Linux matrix is green. - Serialized server shard 1 initially hit an unchanged heartbeat test-harness timeout and a later assertion in the same file. Its isolated rerun passed in 3m57s. All other shards passed on their first attempt. - Greptile reviewed the final commit at 5/5 with no blocking failure. Both earlier actionable validation threads are resolved, and no review thread remains open. ## Risks Low production risk. The package is private and has no exports, server adapter, endpoint, or process. AJV is a package-only development dependency that the server workspace already uses. The main risk is contract churn before the TypeScript and Rust consumers land. The generated manifest and compatibility fixtures make that churn explicit. I checked `ROADMAP.md`. This change defines a contract for planned control-plane work and does not add overlapping product behavior. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a14e51d592 |
refactor(environment): classify environment capabilities from static driver definitions (#12045)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environment runtime drivers provide workspace, lease, and custom image behavior > - Runtime code used driver identity checks and several capability-specific members > - These checks spread capability rules across the runtime and made new drivers harder to verify > - This pull request adds one general capability classifier and one static driver support table > - The benefit is one fail-closed capability model that keeps current behavior and supports future drivers ## Linked Issues or Issue Description **What existing behavior does this improve?** Environment runtime capability checks for workspace realization, custom images, lease capabilities, and duplex authorization. **Subsystem affected** Cross-cutting (multiple of the above) **Current behavior** The runtime selects several capability paths from driver identity and separate capability members. Custom image gates also trust provider declarations without checking every matching live worker method. **Proposed behavior** The runtime uses one general capability classifier and one static support table. Custom image gates require both the provider declaration and every matching live worker method. The public capability names remain unchanged. **Reason and benefit** The change keeps capability rules in one place. It removes identity conditions from runtime consumers and makes unsupported drivers fail closed. **Breaking changes** None. The public names sandboxCapabilities, sandboxProviders, and EffectiveSandboxCapabilities remain available. ## What Changed - Add classifyEnvironmentCapabilities and static support definitions for all four driver families. - Add resolveCapabilities to every environment runtime driver. - Move driver traits into environment-driver-traits.ts and migrate runtime consumers. - Require provider declarations and matching live worker methods for all custom image gates. - Migrate duplex authorization to the general resolver and remove the dead sandbox-only member. - Delete the unused resolveEffectiveSandboxCapabilities wrapper and update its test. ## Verification - pnpm --filter @paperclipai/server typecheck - pnpm exec vitest run server/src/__tests__/environment-capability-contract.test.ts server/src/__tests__/environment-runtime.test.ts — 92 tests pass - pnpm exec vitest run server/src/__tests__/environment-driver-traits.test.ts server/src/__tests__/general-capability-classifier.test.ts — 12 tests pass - pnpm exec vitest run server/src/__tests__/environment-custom-images-service.test.ts server/src/__tests__/environment-execution-target-capabilities.test.ts server/src/__tests__/environment-execution-target-duplex.test.ts server/src/__tests__/environment-execution-target-duplex-kill-switch.test.ts — 66 tests pass ## Risks The main risk is a capability gate that denies a valid driver or permits an invalid driver. The static support matrix, live worker method checks, and regression tests reduce this risk. No database, public API, or published type name changes. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The deployment does not provide a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (for example, docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc9e9b704f |
fix: stop teaching agents to curl literal {id} route templates (#12061)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Adapters inject prompt text that teaches agents how to call the
Paperclip API, including copy-pasteable curl examples
> - Some of those URLs contained brace placeholders like
`/api/issues/{id}/checkout`
> - Agents paste such lines verbatim; the placeholder reaches the server
as `/api/issues/%7Bid%7D` and 404s, and request logs show agents doing
exactly that
> - The acpx engine's API note already avoids this by using
`$PAPERCLIP_TASK_ID`, and its test pins `/api/issues/{id}` out of the
prompt
> - This pull request applies the same standard to the gemini adapter,
the shared prompt template, and the openclaw gateway workflow
> - The benefit is that agents stop burning turns on placeholder 404s
and doc examples stay safe to execute as written
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
Server request logs show agents issuing `GET /api/issues/%7Bid%7D` — the
literal, percent-encoded text `{id}` — which 404s. The source is adapter
prompt text: the gemini adapter's API note embeds a curl example with
`/api/issues/{id}/checkout` in the URL, the shared agent prompt template
mentions `/api/issues/{issueId}` endpoints, and the harness checkout
notice names `/api/issues/{id}/checkout`. Models copy these strings into
real requests.
**Expected behavior**
URL paths in prompt text must carry environment variables or real ids,
never brace placeholders, in every string an agent might execute
verbatim. Where a placeholder is unavoidable, the prompt must state
explicitly that the literal text must never be sent.
**Steps to reproduce**
1. Give an agent the gemini adapter's API access note.
2. Watch it call `curl ...
"$PAPERCLIP_API_URL/api/issues/{id}/checkout"` as written.
3. The server logs `POST /api/issues/%7Bid%7D/checkout 404`.
## What Changed
- gemini-local's API note curl example now uses `$PAPERCLIP_TASK_ID` and
tells the agent to substitute a real issue id when that variable is
absent — the same convention as the acpx engine's API note.
- The shared agent prompt template (`server-utils.ts`) uses
`$PAPERCLIP_TASK_ID` in its interaction-creation and resume-endpoint
mentions, and the harness checkout notice names `POST
/api/issues/$PAPERCLIP_TASK_ID/checkout`.
- openclaw-gateway's endpoint workflow keeps its `{issueId}`
placeholders — they are defined by its "determine issueId" step — but
now states explicitly that the literal text must never be sent in a URL.
- `server-utils.test.ts` pins the new form and adds negative pins that
keep `/api/issues/{id}` and `/api/issues/{issueId}` out of the shared
prompt template, mirroring the existing acpx-engine negative pin.
- The `confirmation:{issueId}:plan:{revisionId}` idempotency-key
template is untouched: it is a value-construction pattern, not a URL.
## Verification
- `npx vitest run packages/adapter-utils/src/server-utils.test.ts
packages/adapters/gemini-local packages/adapters/openclaw-gateway` — 152
passed. The single failure (`pre-selects gemini-api-key auth in the
managed HOME for sandbox execution`) is a pre-existing
environment-specific failure on the development machine, unrelated to
prompt text; CI is authoritative for it.
- `pnpm --filter @paperclipai/adapter-utils --filter
@paperclipai/adapter-gemini-local --filter
@paperclipai/adapter-openclaw-gateway typecheck`.
## Risks
- Low risk: prompt-text and test changes only; no runtime logic changes.
- Agents that memorized the old example strings keep working — the
routes are unchanged, only the placeholder text in prompts is.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
c7f4bc1300 |
fix: survive transient sandbox exec failures in the callback bridge worker (#12052)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents on remote sandbox targets reach the Paperclip API through the sandbox callback bridge: a loopback gateway inside the sandbox writes request files, and a host-side worker polls them over the provider's exec channel and forwards them to the server > - The worker's poll loop had one terminal catch: a single reset or slow exec ended the relay for the rest of the run > - The in-sandbox gateway kept queueing requests against the dead worker, so every later API call from the agent stranded, including its final status write > - A relay that dies on one transient fault turns a routine provider hiccup into a lost issue disposition > - This pull request restructures the loop so transient faults back off and retry, while the watchdog remains the escalation path for sustained outages > - The benefit is that one flaky exec no longer severs an agent from the control plane mid-run ## Linked Issues or Issue Description Refs #9904. Refs #8977. Both touch adjacent bridge behavior (curl shim, header forwarding); neither addresses worker-loop lifetime. No public issue exists for this defect. The description below follows the bug report template. **What happened?** During a staging run, the host-side bridge worker hit one failed sandbox exec while relaying requests. The poll loop's only catch is terminal: it failed the pending requests and set the worker to settled, with no restart. The agent's later API calls saw connection-level failures or bridge errors until the run ended. Its final `PATCH status: done` was lost, and the missing-disposition recovery had to repair the issue in a corrective run. **Expected behavior** One transient exec failure must not end the relay for the rest of the run. The worker must back off and retry. A sustained outage must still fail queued requests fast through the watchdog. In-flight request semantics (abort plus 504 backstop, retry-safe 503) must not change. **Steps to reproduce** 1. Start a remote-sandbox run and let the provider exec channel reject or stall one call while the bridge worker polls. 2. The worker hits one `listJsonFiles` failure or one request-attempt timeout, and the loop exits through its terminal catch. 3. Every later bridge request strands. The loopback gateway keeps accepting requests that never complete, and after 64 queued files it answers every request with 503 until run end. ## What Changed - `startSandboxCallbackBridgeWorker`'s poll loop now separates three failure domains: - A failed poll backs off exponentially (capped at 5 s) and retries instead of dying. The first failure of a streak still lands on the run trace through the workerFailed span; later repeats only warn. - A failed or hung request attempt runs the same recovery pass the loop previously died on — abort the in-flight handler (its 504 backstop keeps the caller from stranding) and 503 the unclaimed queued requests — and the loop then continues and serves the caller's retry. - A listing where every file already has an in-flight attempt sleeps one poll interval, like an empty listing. The previous immediate re-list was a hot spin: an exec storm against a real channel, and a pure-microtask loop that starved every timer in the process when the queue client resolves synchronously. - The watchdog, the claim/finalize fences, and the stop/drain semantics are unchanged. ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 38 passed. - New regression test: a request queued behind three consecutive poll failures is still delivered. - Updated tests: the stalled-poll test now expects recovery (the handler's real 200) instead of a terminal 503; the sustained-outage 503 path remains proven by the dedicated watchdog test; the recovery-503-write-retry test triggers the recovery pass through a hung request read, because a hung poll no longer runs that pass. - `pnpm --filter @paperclipai/adapter-utils typecheck`. ## Risks - Behavior change: a transiently failing poll no longer mass-fails queued requests on the first error. Callers wait through the backoff window, bounded by the existing in-sandbox 30 s response deadline, or the watchdog fails them after 20 s of no successful iteration. This trades fast-but-terminal degradation for recovery. - A hard-down channel now retries every ≤5 s for the rest of the run instead of stopping. Each retry is one exec attempt against a channel that already fails. - In-flight mutation safety is unchanged: the guard map and the claim protocol still prevent a double-applied host mutation, and a guarded file is always finalized by its own attempt or by its 504 backstop. ## Model Used - Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended thinking enabled, agentic tool use via Claude Code (CLI harness), 200k context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c62bb4b16b |
feat: environment delete with agent reassignment and consented sandbox destroy (#12053)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Environments define where agent runs execute: local, SSH, or provider sandboxes > - Operators can create and edit environments, but the UI has no way to delete one > - The server already exposes `DELETE /environments/:id` and a delete-blast-radius preflight, but no UI consumes them, and a delete blocked by reusable sandbox leases gives the operator no path forward > - This pull request adds the delete flow to the environment configuration page: a preflight-driven modal that reassigns dependent agents, names the workspaces that hold blocking sandbox leases, and can destroy those sandboxes with explicit consent > - The benefit is that operators can retire stale environments from the UI without database surgery, and dependent agents move to a chosen replacement instead of silently falling back ## Linked Issues or Issue Description Refs #8554 Refs #11124 **Subsystem affected** Environments (server routes, environment runtime service, and the environment settings UI). **Problem or motivation** The environment configuration page has no delete control. The server delete endpoint exists, but nothing in the UI calls it. When reusable sandbox leases block a delete, the 409 error names no owner, so the operator cannot find the blocking workspace. Agents that use the environment as their default lose it silently through the FK `on delete set null`. **Proposed solution** Add a delete button with a confirmation modal on the environment edit page. The modal reads the delete-blast-radius preflight. It offers a dropdown to reassign dependent agents to another environment before the delete. It lists each workspace that holds a blocking reusable sandbox lease, with a link. When those leases are the only blocker, the confirm button destroys the sandboxes inline (`?destroyReusableSandboxLeases=true`) and then deletes. A failed teardown falls back to `pending_cleanup` for the sweep, so no sandbox is orphaned. ## What Changed - `ui/src/pages/CompanyEnvironments.tsx`: delete button on the edit page header, confirmation modal with agent reassignment select, lease-holder list, impact notes, and a consent-labeled destroy-and-delete action - `ui/src/api/environments.ts`: `deleteBlastRadius` and `remove` client methods; `remove` takes an optional `destroyReusableSandboxLeases` flag - `server/src/routes/environments.ts`: `DELETE /environments/:id` accepts `?destroyReusableSandboxLeases=true`; it destroys the environment's reusable sandbox leases first, but only when those leases are the sole delete blocker, then re-checks the blast radius before it deletes - `server/src/services/environment-runtime.ts`: new `destroyReusableSandboxLeasesForEnvironment` — destroys every reusable sandbox lease an environment still owns while the environment config (provider credentials) is still available - `server/src/services/environments.ts`: the delete blast radius now returns `reusableSandboxLeaseHolders` (lease id, workspace, issue) so clients can name what blocks a delete - `packages/shared/src/types/environment.ts`: `EnvironmentDeleteReusableLeaseHolder` type on the blast radius - Tests: route gating for the consent flag (destroy runs, mixed-blocker rejection, surviving-lease rejection), runtime destroy scoped to an environment, blast-radius holder join, and UI tests for the reassignment flow, holder links, and the consent button ## Verification - `npx vitest run server/src/__tests__/environment-routes.test.ts server/src/__tests__/environment-service.test.ts server/src/__tests__/environment-runtime.test.ts ui/src/pages/CompanyEnvironments.test.tsx` - Manual: open Settings → Environments → edit an environment. The trash icon opens the modal. With agents on the environment, pick a reassignment target and confirm; agents move and the environment deletes. With reusable sandbox leases, the modal names the holding workspaces and the confirm button reads "Destroy N sandboxes and delete". ## Risks - The consented path destroys provider sandboxes. It runs only when reusable leases are the sole blocker, so a delete that would still be rejected never destroys anything. A failed teardown routes to `pending_cleanup` and the delete stays blocked until the sweep resolves it. - Agent reassignment issues one PATCH per agent from the client. A mid-sequence failure leaves some agents reassigned; the reassignments are valid on their own and the UI refreshes to the actual state. - Hard blockers (managed local, instance default, pending cleanup) keep the existing 409 behavior and disable the confirm button. ## Model Used - Claude (Anthropic) — Fable 5, model id `claude-fable-5`, extended thinking, agentic tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
63df7ad2b3 |
feat(login): use the login pseudo-terminal for Codex device login and de-Claude the shared channel (#12020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use provider-specific login flows > - Codex device login needs a live pseudo-terminal (PTY), while the shared channel still uses Claude-specific names > - The old streamed-exec path does not provide the prompt transport that Codex needs > - This pull request moves Codex device login to the shared login PTY and removes the dead streamed-exec path > - The benefit is one controlled login transport with fail-closed capability checks and safer credential reads ## Linked Issues or Issue Description **Problem or motivation** Codex device login used a streamed-exec path that did not provide the required prompt transport. The shared login channel also exposed Claude-specific names outside Claude code. **Expected behavior** The host selects a fixed login command from trusted adapter data. Codex login uses the provider login PTY. Providers without that capability fail closed. **Proposed solution** Use a server-controlled session home, create and validate it as a fresh 0700 directory, read credentials from one validated descriptor, and rename shared channel names to the neutral login PTY family. **Alternatives considered** Keep the shared login PTY as the single transport. Do not keep the removed streamed-exec path because it cannot provide the required prompt transport. **Roadmap alignment** This change supports the planned login transport work. It does not add a separate roadmap item. ## What Changed - Route Codex device login through the shared login PTY transport. - Select the login command from a closed internal command key. - Carry a server-controlled session home through the launch contract. - Create and validate the session home as a fresh 0700 directory owned by the login user. - Read the credential file with descriptor-relative, no-follow path walking and final descriptor checks. - Gate the login route and run lease on the provider login PTY capability. - Rename shared channel names to the neutral login PTY family. - Remove the streamed-exec transport value, selector field, driver branch, and related tests. - Hide Codex login in the user interface when the provider lacks the login PTY capability. ## Verification - Server unit suites pass: 89/89. - Adapter-utils suites pass: 262/262. - Codex-local suites pass: 326/326. - Credential-read reader suite passes: 20/20. - Daytona login PTY suite passes: 30/30. - Device-login suites pass: 56/56. - TypeScript checks pass for server, adapter-utils, and UI. - GitHub Actions must pass after pull request creation. - Greptile review must reach 5/5 with no open P2 findings, recommendations, or follow-ups. ## Risks - Providers without a login PTY capability lose Codex login support by design. - The credential read rejects invalid ownership, mode, type, path, and size. - The launch-time sandbox directory race remains outside the threat model because the login runs inside the sandbox and a hostile sandbox already controls its credential. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact context window and reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b59c9315 |
feat(adapter-utils): stream duplex bridge bodies as sequenced chunks with receive-side spill (#12006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use a duplex bridge to send requests and responses across an isolated boundary > - The bridge held each request body and response body in memory on both ends > - Large bodies can exhaust memory and reduce the safe size of adapter traffic > - This pull request sends receive-side bodies as sequenced chunks and spills large bodies to disk > - The benefit is bounded memory use with strict size, order, and cleanup checks ## Linked Issues or Issue Description **What existing behavior does this improve?** The adapter-utils duplex bridge transports request and response bodies across the sandbox boundary. **Current behavior** The bridge stores each complete body in memory on both ends of the duplex channel. **Proposed behavior** The bridge sends body chunks with sequence checks. The receive side keeps bodies up to 1 MiB in memory and spills larger bodies to a temporary file. **Reason and benefit** This change reduces memory pressure and keeps malformed or oversized input on a terminal error path. **Breaking changes** The duplex frame version changes to version 2. The request and response envelopes now carry bodyByteCount, and body_chunk frames carry the body data. ## What Changed - Add version 2 body_chunk frames with 256 KiB raw slices encoded as canonical base64 text. - Add receive-side memory and spill reassembly with per-channel disk and file limits. - Reject malformed, reordered, oversized, truncated, and non-canonical body chunks. - Stream reassembled request bodies to the host forward handler with a web stream and half-duplex request. - Remove spill files on success, failure, channel death, and startup cleanup. ## Verification - Run the adapter-utils type-check. - Run the adapter-utils duplex test suite. - Run all pull request checks. - Run the Greptile review and confirm a 5/5 score with no open findings. ## Risks The wire format changes from version 1 to version 2. Older bridge peers cannot use this protocol. The receive path adds temporary file operations and cleanup paths. The implementation fails closed when a body violates size or sequence rules. ## Model Used OpenAI GPT-5 Codex, tool-enabled coding agent. The exact context-window size and reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/..., fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
05b35d4669 |
feat(duplex): bound aggregate duplex route resource consumption with a process-owned byte ledger (#12003)
## Thinking Path > - Paperclip runs AI agents through adapters and sandboxed execution targets. > - Duplex routes retain bytes across route data, broker messages, decoder buffers, and readiness replay. > - Per-route limits bound each route but do not bound the total retained bytes across many routes. > - A process-owned ledger must charge each retained buffer before allocation and release the charge during cleanup. > - This pull request adds the aggregate ledger, connects it to host and sandbox duplex paths, and adds route coverage. > - The benefit is a fail-closed process-wide byte limit that keeps concurrent duplex work within a safe resource budget. ## Linked Issues or Issue Description **Subsystem affected** This change affects packages/adapter-utils and server duplex orchestration. **Problem or motivation** Many routes can each stay below their per-route limits while their combined retained bytes exceed a safe process budget. **Proposed solution** Add a process-owned aggregate byte ledger. Charge route data, broker bytes, decoder buffers, and readiness replay bytes before allocation. Release each charge during cleanup. Use a separate sandbox_process decoder cap for the in-sandbox path. **Alternatives considered** Keep only per-route limits. This does not bound the combined process use. Set a fixed limit at one call site. This misses retained bytes in other duplex paths. **Roadmap alignment** This is a tightly scoped reliability and resource-safety improvement. It does not duplicate a roadmap feature. **Additional context** The aggregate ceiling uses a safe 256 MiB default. An invalid override falls back to that default and reports the rejected value. ## What Changed - Add a process-owned aggregate byte ledger for duplex route resource use. - Charge and release route data, broker forward and response bytes, decoder buffers, and readiness replay bytes. - Bound host-to-worker pending writes and standard input transport bytes. - Add a separate decoder cap for the sandbox_process path. - Make invalid aggregate-ceiling overrides fall back to the safe default without host startup failure. - Add adapter-utils and server tests for charging, release, rejection, cleanup, and many-route aggregate limits. ## Verification - pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit - pnpm --filter @paperclipai/server exec tsc --noEmit - Run the focused adapter-utils duplex ledger and execution-target tests. - Run the server aggregate-ledger route test. - Confirm all required pull request checks pass on this branch. ## Risks The ledger touches several duplex buffer paths. A missed release could reduce later capacity until process restart. The tests cover charge, release, rejection, cleanup, and route aggregation. The change uses a safe default when configuration input is invalid. ## Model Used OpenAI GPT-5 Codex. The runtime model ID and context window are not exposed to this task. The model used tool calls, shell commands, and code review workflow support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes/Closes/Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issue references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c5050396c7 |
fix(daytona-duplex): chunk host-to-sandbox writes and make a transport close legible (#11986)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agent work through adapters and sandbox providers > - The Daytona duplex path sends host input through a provider pseudo-terminal WebSocket > - Large messages exceed the provider limit, and a transport close can look like a process exit > - This pull request chunks UTF-8 input and carries transport-close state through the duplex path > - The benefit is reliable large input and accurate loss reporting ## Linked Issues or Issue Description **What happened?** The Daytona duplex path sent a full input payload as one WebSocket message. A payload above the provider limit closed the channel. The wait path also mapped a non-numeric exit result to a process exit without exit data. **Expected behavior** The provider must receive large input as ordered UTF-8 chunks. A transport close without exit data must record `transport_closed`, while a numeric exit must record `provider_exit`. **Steps to reproduce** 1. Start a Daytona duplex session. 2. Send an input payload larger than 65536 bytes. 3. Observe that one message closes the provider channel. 4. End a session without a numeric exit code. 5. Observe that the loss reason reports a process exit. **Paperclip version or commit** Commit `1761e79ec9097c65d94f90a8ba20416f8ab718a6`. **Deployment mode** Built from source with the Daytona sandbox provider. ## What Changed - Add a shared UTF-8 byte chunker with a 32768-byte cap. - Route both Daytona pseudo-terminal write paths through the chunker. - Preserve multi-byte UTF-8 sequences across read-side chunks. - Carry an explicit `transportClosed` state through the worker and host wait paths. - Record `transport_closed` for a reason-less transport close and `provider_exit` for a numeric exit. - Keep orderly completion suppression for both exit paths. ## Verification - The Daytona plugin suite passes 194 tests. - The adapter-utils broker, codec, and telemetry suites pass 73 tests. - The plugin SDK duplex and worker RPC host suites pass 37 tests. - The server plugin worker manager duplex suite passes 78 tests. - The execution target sandbox and ACPX execute suites pass 257 tests. - TypeScript checks pass for adapter-utils, plugin SDK, server, and the standalone Daytona plugin. ## Risks The chunk size adds a loop for large input payloads. The 32768-byte cap stays below the provider limit. The optional loss field preserves compatibility for other providers. ## Model Used OpenAI Codex, GPT-5, extended reasoning, tool use, and code execution. The runtime does not expose a separate context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
141b815294 |
fix(adapter-utils): enforce the duplex frame size bound on encode in both codec copies (#11983)
## Thinking Path > - Paperclip uses adapter utilities to move bounded messages between agent processes. > - The duplex frame codec encodes and decodes these messages. > - The decoder rejects frames above the documented byte limit. > - The encoder did not apply the same limit before it sent a frame. > - This mismatch let a sender write a frame that the peer rejected after transmission. > - This pull request applies the same byte limit to both codec copies and keeps the broker channel open. > - The benefit is a local error with stable request telemetry instead of a channel loss. ## Linked Issues or Issue Description No public GitHub issue exists for this change. The problem follows the bug report fields below. **What happened?** The duplex encoder could write a frame larger than `DEFAULT_MAX_DUPLEX_FRAME_BYTES`. The peer decoder then rejected the frame after transmission. In the WebSocket 1009 case, this closed the channel and reported a process exit. **Expected behavior** The encoder should reject an oversized frame before it writes bytes. The gateway should return HTTP 413. The broker should return a bounded terminal response and keep other requests active. **Steps to reproduce** 1. Encode a duplex frame above `DEFAULT_MAX_DUPLEX_FRAME_BYTES`. 2. Send the frame through the gateway or broker. 3. Observe that the old path writes the frame or drops the channel after peer rejection. **Paperclip version or commit** Reproduced from the `master` development line before this change. **Deployment mode** Local dev (`pnpm dev`). ## What Changed - Add `encodeDuplexFrameChecked` to the host and embedded gateway codecs. - Measure encoded JSON bytes without the trailing newline. - Return a typed `frame_too_large` result without throwing. - Return HTTP 413 for oversized gateway requests without writing a frame. - Share one frame bound between broker decode and encode checks. - Return a bounded, non-retryable terminal response for oversized broker responses. - Add encode vectors to the shared wire-compatibility fixture. ## Verification - Run `pnpm --filter @paperclipai/adapter-utils typecheck`. - Run `npx vitest run packages/adapter-utils/src/duplex-frame-codec.test.ts packages/adapter-utils/src/duplex-bridge-broker.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts`. - Confirm the oversized-response broker test keeps the channel open and serves the other in-flight request. - Confirm the gateway test returns HTTP 413 and keeps the channel open. ## Risks The encoder now rejects oversized frames before transmission. This changes an unsafe write into a typed local error. The broker and gateway keep existing frame limits and affect only oversized frames. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The model assisted with review and repository operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc42a67e7e |
fix(adapter-utils): extend the duplex fail-closed run disposition to the CLI lane (#11966)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through adapter execution lanes > - Duplex adapters can lose their control channel before a process completes > - The ACP lane already fails closed, but the CLI lane can report false success > - This pull request applies the same completion rule to the CLI lane and shares the loss code > - The benefit is consistent failure reporting when a duplex channel closes during a run ## Linked Issues or Issue Description **What happened?** A CLI-lane duplex run can lose its control channel before clean process completion. The run can then report `succeeded` with exit code 0 and no error code. **Expected behavior** The execution target must fail closed when the channel dies before clean completion. It must return exit code 1, the typed `duplex_channel_lost` error code, and a short stderr note. **Steps to reproduce** 1. Start a duplex adapter run through the CLI execution lane. 2. Close the duplex control channel before the process completes cleanly. 3. Inspect the run result and error code. **Paperclip version or commit** Commit `5e01523d4eb6df4a20a0bddd05374c9c42225203`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Claude Code, Codex, Cursor, Gemini, Kimi, OpenCode, and Pi local adapters. **Database mode** Not database-related. ## What Changed - Add an optional `errorCode` field to `RunProcessResult`. - Add a one-read completion seam to the execution target process options. - Fail closed when a duplex channel dies before clean process completion. - Add `settleRunDisposition()` to atomically read and mark orderly completion. - Share the typed duplex loss error code across the ACP and CLI lanes. - Mark non-success terminal results as orderly completion before teardown. - Wire the seam through the seven duplex adapters. - Add regression tests for channel loss, clean completion, and non-clean terminal results. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` — 118 passed. - `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "sandbox duplex run-disposition seam"` — 4 passed. - The author confirmed a clean type-check for `@paperclipai/adapter-utils` and the seven duplex adapter packages. - Pre-existing environment failures remain outside this change. They include `EACCES mkdir '/srv/paperclip'` and remote file-size setup failures. ## Risks The change alters terminal status for CLI duplex runs that lose control before clean completion. The typed error code and stderr note keep the failure visible. The broker marks failed, cancelled, and timed-out results as orderly completion to prevent false loss events during teardown. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, with the standard GPT-5 context window. The model assisted with the implementation and test work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
10d2781a29 |
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox adapters provide controlled execution for untrusted provider environments. > - The sandbox channel needs one persistent duplex transport with strict host control. > - The transport must remain off unless the instance setting and provider capability both allow it. > - The host must detect loss, bound resource use, and expose only safe telemetry. > - This pull request adds the broker, gated selection, kill-switch wiring, fixed observability, and real-process proof. > - The benefit is safer sandbox execution with bounded failure behavior and inspectable transport results. ## Linked Issues or Issue Description No public issue exists for this change. The related pull requests are #11738 and #11750. **Problem or motivation** The sandbox duplex channel needs a host-controlled broker, strict transport gates, bounded provider input, and safe loss telemetry. Without these controls, a provider can cause replay, resource growth, unsafe endpoint selection, or data exposure through telemetry. **Proposed solution** Add a host broker with nested time limits, request limits, one-shot loss, and per-id deduplication. Select duplex transport only when the instance setting and provider capability both equal true. Assign the endpoint and nonce on the host. Reject invalid readiness data and use the file bridge on failure. Add fixed redacted telemetry and a real-process end-to-end test harness. **Alternatives considered** Keep the file bridge as the only transport. This avoids new channel behavior but does not provide persistent duplex operation for supported sandbox providers. **Roadmap alignment** This change supports the Cloud / Sandbox agents section in ROADMAP.md. ## What Changed - Add the duplex bridge broker with bounded forward, response, and gateway wait budgets. - Bound concurrent requests, lifetime requests, and request-id bytes before retention or forwarding. - Select duplex transport only when both required gates are true. - Assign the loopback port and nonce on the host and enforce a liveness-only READY frame. - Fall back to the file bridge after invalid readiness, contamination, bind failure, or timeout. - Carry the kill switch through the server, acpx engine, and six local adapters. - Add fixed, redacted duplex telemetry with a provider allowlist. - Add a real-process end-to-end harness for readiness, round trips, loss, and teardown. - Add regression coverage for limits, loss, UTF-8 splits, concurrency, and telemetry dimensions. ## Verification - Adapter-utils, server, and Daytona typechecks pass locally. - Adapter-utils tests pass, including the codec, broker, execution-target sandbox, and real-process harness. - Server kill-switch tests pass. - Live Daytona tests pass with the required provider key and skip without that key. - The root pnpm-lock.yaml file has no diff. - The branch contains ten commits after origin/master. ## Risks - Duplex transport remains disabled unless both gates equal true. - A provider remains an untrusted boundary and needs least-privilege credentials and quotas. - The server telemetry recorder stays deferred; the default recorder does nothing. - A provider that pre-binds the host port causes a fail-closed fallback to the file bridge. - The change adds no database migration and changes no root lockfile. ## Model Used OpenAI GPT-5, exact model family GPT-5, large context window, reasoning, and tool use. The model assisted with Git handoff validation and PR preparation. The implementation commits came from the engineering worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
417336f8be |
fix(workspaces): attach PR preparation to existing branches (#11703)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces isolate an agent task from the primary checkout. > - Pull request preparation can need a branch that already contains completed work. > - The workspace policy could not require an exact existing branch. > - Workspace cleanup also treated worktree creation as branch ownership. > - This pull request adds an exact existing-branch policy and separate branch ownership metadata. > - The benefit is safe pull request preparation that preserves every existing commit and operator-owned branch. ## Linked Issues or Issue Description **What happened?** A pull request preparation run could not pin its execution workspace to an exact existing branch. Workspace reuse and cleanup could also confuse worktree creation with branch ownership. **Expected behavior** The run must attach only to the requested branch in an isolated Git worktree. It must fail if the branch is missing, busy, or inconsistent. Cleanup must not delete a branch that Paperclip does not own. **Steps to reproduce** 1. Create a branch that contains completed work. 2. Configure a pull request preparation task to use that branch. 3. Start the task and observe that the prior policy cannot require the exact branch. **Paperclip version or commit** This behavior reproduces on the base revision before this pull request. **Deployment mode** Local development with isolated Git worktrees. ## What Changed - Add `existingBranch` to the execution workspace policy and shared validation contracts. - Require `existingBranch` to use an isolated Git worktree and reject conflicting branch templates. - Attach to the exact branch without creating, renaming, resetting, or deleting it. - Track branch ownership separately from worktree creation and use that ownership during cleanup. - Return HTTP 422 for invalid existing-branch settings on every issue-producing route. - Add a bounded repair script for existing pull request preparation tasks. - Add focused policy, route, heartbeat, runtime, and ready-comment tests. - Document the exact-branch behavior and safety rules. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspace-policy.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-existing-branch-validation-status.test.ts server/src/__tests__/workspace-runtime.test.ts server/src/services/workspace-runtime-exposure.test.ts server/src/services/workspace-runtime-ready-comment.test.ts` passed 335 tests. - `pnpm -r typecheck` passed for all workspace projects. - `pnpm test:run` passed 4,431 tests. Two unrelated embedded-Postgres setup hooks timed out under aggregate load. Their isolated rerun passed 74 tests. - `pnpm build` passed for all workspace projects. - The two review regressions passed with 139 unrelated tests skipped. - All latest-head CI gates passed after one unrelated timing-sensitive test passed on rerun. - Greptile scored the latest head 5/5 with no unresolved review threads. ## Risks - Invalid workspace settings now return HTTP 422 instead of a generic validation response. - The exact branch must already exist and must not be checked out by another worktree. - The new policy fails closed when it cannot prove branch identity or ownership. - This change has no database migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family assisted with this change. The runtime did not expose its exact deployment ID or context window. The agent used high-reasoning mode, repository tools, shell execution, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fbd20b28d3 |
fix(grok-local): stop defaulting --permission-mode to dontAsk (#11898)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `grok_local` adapter runs the native Grok Build CLI in headless
mode for unattended agent heartbeats
> - Grok CLI 1.0 started to enforce the `dontAsk` permission mode as
deny-by-default, and it takes precedence over `--always-approve`
> - The adapter passes both flags on every run, so each run dies on its
first tool call and is still recorded as a success
> - This pull request removes the `dontAsk` default so unattended runs
rely on `--always-approve` alone
> - The benefit is that `grok_local` agents can execute tools again on
current Grok CLI releases
## Linked Issues or Issue Description
No public issue exists. Description per the bug template:
**What happened?**
Every `grok_local` run on Grok CLI 1.0.x stops on its first tool call.
The stream shows the tool call move from `pending` to `failed` with
"User cancelled the execution for tool `run_terminal_command`", and the
session ends with `stopReason: "cancelled"` after one turn. The CLI
exits 0, so Paperclip records the run as succeeded with no work done,
and the issue lands in missing-disposition recovery.
**Expected behavior**
Unattended runs must auto-approve tool executions. The adapter already
passes `--always-approve` for this.
**Steps to reproduce**
In a clean Linux environment with Grok CLI 1.0.3 and `XAI_API_KEY` set,
run the adapter's exact invocation shape:
`grok --output-format streaming-json --permission-mode dontAsk
--always-approve --disable-web-search --single "Run the shell command:
echo ok"`
The tool call is denied. Drop `--permission-mode dontAsk` (or use
`--permission-mode bypassPermissions`) and the same command executes the
tool. On Grok 0.2.x the original combination worked because the CLI
accepted `dontAsk` without enforcing it; the 0.2.39 embedded docs state
the flag takes effect only for `bypassPermissions` / always-approve.
**Paperclip version or commit**
master (
|
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
5bc6031f79 |
fix(server,ui,claude-local): verify auth on the adapter Test lane and enforce managed-sandbox tenant binding (#11810)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter Test checks whether an agent adapter can run with its configured environment, and every local-driver adapter (Claude, Codex, Gemini, OpenCode, Pi, Cursor, etc.) shares this Test route and its UI resolution logic > - The Claude ACP Test lane could report pass without checking local or remote authentication, and the shared Test route and UI had gaps in environment binding, probe safety, and managed-sandbox resolution that affect every adapter that uses the Test button, not only Claude > - This pull request verifies authentication on every Claude ACP target, and closes the shared Test-route/UI gaps: tenant-binding on the route, a managed-sandbox-only redirect that matches the real run path, and a three-tier environment resolution in the UI > - The benefit is a truthful Test result with safer probe execution and tenant isolation, for Claude specifically and for every other local adapter that shares this Test surface ## Linked Issues or Issue Description **What happened?** The Claude ACP Test lane returned `status: "pass"` without checking authentication for some local and non-sandbox targets. Separately, the shared `/companies/:companyId/adapters/:type/test-environment` route — used by every local-driver adapter, not only Claude — accepted a foreign environment id, and its UI resolution did not mirror the server's managed-sandbox-only redirect. **Expected behavior** The Test lane checks the resolved credential and hello probe for every Claude ACP target. The shared adapter Test route rejects a foreign environment before it reveals environment details or starts a lease, for any adapter type. The Test's environment resolution (UI and server) matches the real run's three-tier resolution, including the managed-sandbox-only redirect. **Steps to reproduce** 1. Run the Claude ACP Test lane against a local target without a valid credential. 2. Run the adapter Test route with an environment id from another company (any adapter type). 3. Observe the pass result on step 1, or the missing tenant-binding rejection on step 2. **Paperclip version or commit** `933749e01f74e82ce5d315c071be534d04e01158` **Deployment mode** Local dev (`pnpm dev`) and server route tests. **Agent adapter(s) involved** Claude Code directly (the ACP auth-verification work). The tenant-binding guard, managed-sandbox-only redirect, and UI three-tier resolution apply to the shared adapter Test route and affect every local-driver adapter (Codex, Gemini, OpenCode, Pi, Cursor, etc.), not only Claude — see "What Changed" below for the split between Claude-only and shared changes. **Database mode** Not database-related. **Access context** Both board and agent paths use the affected Test surface, for every local-driver adapter. **Additional context** Two commits that were previously bundled into this PR — a `plugin-worker-manager` duplex-channel frame-bound fix and a `workspace-runtime` exit-persist crash fix — are unrelated to the adapter Test lane and have been split out into their own PRs: #11860 and #11861. ## What Changed Claude-only (`packages/adapters/claude-local`): - Verify `CLAUDE_CODE_OAUTH_TOKEN` and run the hello probe for every Claude ACP target. - Keep `adapter_auth_missing` sandbox-only and report missing non-sandbox credentials as a warning. - Add a deny-by-default probe environment builder for the ACP and CLI local probes. - Log only fixed probe context and allowlisted classifications. - Seed the host OAuth token into the hello probe environment. Shared, cross-adapter (`server/src/routes/agents.ts`, `ui/src/lib/adapter-test-environment.ts`, `ui/src/components/AgentConfigForm.tsx`, `ui/src/components/OnboardingWizard.tsx`): - Add a company-binding guard and a binding assertion for the generic `/companies/:companyId/adapters/:type/test-environment` route, so a foreign-company environment id is rejected before any secret resolution or sandbox lease, for every adapter type. - Resolve all three server environment tiers (agent default, instance default, local default) in the UI, and add the managed-sandbox-only redirect so the Test probes the same target a real run would use. - Enforce onboarding Test results: block hire on a failed environment test. - Add regression tests for authentication, tenant binding, probe safety, diagnostics, and UI resolution. ## Verification - Adapter suites pass for the Claude local server probe, remote, ACP, auth, probe environment, and config paths. - Server route tests pass, including the five tenant-binding cases. - UI adapter Test environment resolver tests pass for all three resolution tiers. - Adapter package `tsc --noEmit` exits 0. - Full CI must pass on this pull request. ## Risks The probe environment now denies caller variables by default. A required variable that is not on the allowlist could stop a probe from starting. The route now rejects foreign environment ids with a fixed 403 response. The managed-sandbox-only redirect changes where the Test (and the login affordance) probes for every local-driver adapter under that policy, not only Claude — operators running other local adapters under managed-sandbox-only will see their Test target move from local to the managed sandbox, matching what real runs already do. The change limits secret and diagnostic exposure. ## Model Used Original implementation: OpenAI Codex, GPT-5; exact context window not exposed in that run; tool use and code execution. This revision (commit split and title/description correction): Claude, Sonnet 5 (claude-sonnet-5). The original title and description described this PR as Claude-only; review found it also changes the shared adapter Test route and UI resolution used by every local-driver adapter, and carried two unrelated server fixes. Claude split those two commits into #11860 and #11861 via `git rebase --onto` (verified byte-identical to the original tree minus those commits) and rewrote this description to reflect the actual scope. No functional code in this PR was authored by Claude. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1746783d40 |
fix(kimi): align package with Node 24 policy (#11890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip now requires Node.js 24.11.0 or later > - Each workspace package must publish the same Node.js engine requirement > - The Kimi adapter entered `master` after the Node.js upgrade branch started > - Its package still used Node.js 22 types and had no engine requirement > - This pull request aligns the Kimi adapter with the repository Node.js policy > - The benefit is that the Node.js policy check passes again on `master` ## Linked Issues or Issue Description **What happened?** The `pnpm check:node-version` command fails on `master`. The Kimi adapter uses `@types/node` 22 and has no `engines.node` value. **Expected behavior** All workspace packages must use Node.js 24 types and declare Node.js 24.11.0 as the minimum version. **Steps to reproduce** 1. Check out commit `a7e689b3c`. 2. Use Node.js 24.11.0. 3. Run `pnpm check:node-version`. **Paperclip version or commit** `a7e689b3c` **Deployment mode** Local dev (`pnpm dev`). ## What Changed - Update the Kimi adapter to use `@types/node` 24. - Add the repository minimum Node.js engine requirement to the Kimi adapter package. - Keep `pnpm-lock.yaml` out of this pull request. ## Verification - `npx -y -p node@24.11.0 -c 'node --version && pnpm check:node-version && pnpm --filter @paperclipai/adapter-kimi-local typecheck'` - The command reports Node.js `v24.11.0`. - The Node.js policy check passes. - The Kimi adapter typecheck passes. - A broader local suite was started and stopped at the maintainer's request after the focused checks passed. ## Risks - Low risk. This change updates package metadata and development types only. - The lockfile refresh runs in separate repository automation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, with repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3abe9e2134 |
build(deps): bump zod from 3.25.76 to 4.4.3 (#11719)
Bumps [zod](https://github.com/colinhacks/zod) from 3.25.76 to 4.4.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/colinhacks/zod/releases">zod's releases</a>.</em></p> <blockquote> <h2>v4.4.3</h2> <h2>Commits:</h2> <ul> <li>4c2fa95ce3f3390fbc522324e406b4e9e89b88f9 docs: use Zernio primary wordmark for gold sponsor logo</li> <li>2aeec83eb135e3a83756e973ef44845fc5a455d2 docs: prune lapsed gold sponsors and rebalance logo sizing</li> <li>7391be88ac1ee5cd02057f5ccc012a1f5df4efd0 docs: prune lapsed silver/bronze sponsors and add active ones</li> <li>2c703322a21b4e2b12f33f49ea8430c451a68b4f docs: normalize bronze sponsor logos to github avatar pattern</li> <li>9195250cab0e7950efe39c3926d6c203b4b0a170 docs: remove Mintlify from bronze sponsors (churned)</li> <li>b8dffe9e62f17e6571e6249d05cc5102b54d94e4 docs: remove Numeric and Speakeasy (2+ missed monthly cycles)</li> <li>1cab69383fcdeae2a366d5e2a2fc4d8fc765d168 fix(v4): restore catch handling for absent object keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5937">#5937</a>) (<a href="https://redirect.github.com/colinhacks/zod/issues/5939">#5939</a>)</li> <li>c2be4f819064eed62c7c350a2d399b5faecd15f8 fix(v4): generalize optin/fallback to transform; restore preprocess on absent keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5941">#5941</a>)</li> <li>f3c9ec03ba7a28ae72d25cc295f38674bee0f559 4.4.3</li> <li>1fb56a5c18c27102dbc92260a4007c7732a0ccca docs: document release procedure in AGENTS.md</li> </ul> <h2>v4.4.2</h2> <h2>Commits:</h2> <ul> <li>0c62df0ea19fd05abdf90473e9eef7eea530fab2 Clean up docs navigation and stale labels (<a href="https://redirect.github.com/colinhacks/zod/issues/5901">#5901</a>)</li> <li>20cc794895cc8604fe0c87d83a5d1c3f89fad0ac chore: add security policy and refresh tooling deps</li> <li>6fbe07b0177efdd1bf1c0b05160e70d7a0702337 fix(docs): heading anchor links now include the hash so it doesnt scoll all the way up, follows navbar logic (<a href="https://redirect.github.com/colinhacks/zod/issues/5791">#5791</a>)</li> <li>4bbed1b1c73eca4ce9e59b1189ed236aa6c8b5bd Tighten discriminated union option typing</li> <li>bbac3e567e7fccfaaf7cdc97f1ce30c295e2c908 Update PR guidance for agents</li> <li>cf0dc942a32805c292fff59ade20a7ace980735a Merge remote-tracking branch 'origin/main' into fix-discriminated-union-key-constraint</li> <li>292c894a5fd2aa42e527900b83d8d7a3009a709c docs: add Zernio gold sponsor</li> <li>1fc9f311c28dcf80d0bb5a36b177086cbc3d8eca docs: document codec inversion</li> <li>1373c85da9aeff704a9762d27bc58699618aefb7 docs: remove AI disclosure guidance</li> <li>e20d02b473c08e3a4e557bc610b1b5fac079b649 chore: ignore triage notes</li> <li>e58ea4d91b1dfe8194b73508203213cbc7e9c936 docs: test Zod Mini tab code heights</li> <li>905761a5d127e8d5dd2ebb3bc88c75cb0b8149ff docs: document preprocess input type narrowing</li> <li>bf64bac850d4dee2b7dde7e64909d5d796d32043 chore: tighten test guidance in AGENTS.md</li> <li>8ec4e73f4c4693b6361ad591be40fb41eb8a9f95 chore: update play.ts scratch</li> <li>02c2baf7d0d615872fa4528a8020603b71211702 Make z.preprocess defer optionality to inner schema (<a href="https://redirect.github.com/colinhacks/zod/issues/5929">#5929</a>)</li> <li>88015df8e25c44fb5385eb3ef28935119cd5edea fix(docs): drop deprecated <code>baseUrl</code> from tsconfig</li> <li>c59d4474e3b4cad1b323462186cf607178ce8267 4.4.2</li> </ul> <h2>v4.4.1</h2> <h2>Commits:</h2> <ul> <li>481f7be4238c83ed58183f921b2646f340a91c6a ci: gate release publishing on full test workflow</li> <li>95ccab423aec720b2523c3a64cdc7e3204537cc7 test(v3): restore optional undefined expectations</li> <li>cede2c63739a5823d6aa5093d291e9a111da943d fix(v4): reject tuple holes before required defaults (<a href="https://redirect.github.com/colinhacks/zod/issues/5900">#5900</a>)</li> <li>edd0bf0f5ada4a8dc581c259407d7bbad0a71ea7 release: 4.4.1</li> <li>180d83d1dbe6a59260710cc8637a3dea2281ee56 docs: remove Jazz featured sponsor</li> </ul> <h2>v4.4.0</h2> <h2>4.4.0</h2> <p>This is a minor release with a wide set of correctness and soundness fixes. Some fixes intentionally make Zod stricter, so code that depended on previously accepted invalid or ambiguous inputs may need small updates.</p> <h2>Potentially breaking bug fixes</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/colinhacks/zod/commit/1fb56a5c18c27102dbc92260a4007c7732a0ccca"><code>1fb56a5</code></a> docs: document release procedure in AGENTS.md</li> <li><a href="https://github.com/colinhacks/zod/commit/f3c9ec03ba7a28ae72d25cc295f38674bee0f559"><code>f3c9ec0</code></a> 4.4.3</li> <li><a href="https://github.com/colinhacks/zod/commit/c2be4f819064eed62c7c350a2d399b5faecd15f8"><code>c2be4f8</code></a> fix(v4): generalize optin/fallback to transform; restore preprocess on absent...</li> <li><a href="https://github.com/colinhacks/zod/commit/1cab69383fcdeae2a366d5e2a2fc4d8fc765d168"><code>1cab693</code></a> fix(v4): restore catch handling for absent object keys (<a href="https://redirect.github.com/colinhacks/zod/issues/5937">#5937</a>) (<a href="https://redirect.github.com/colinhacks/zod/issues/5939">#5939</a>)</li> <li><a href="https://github.com/colinhacks/zod/commit/b8dffe9e62f17e6571e6249d05cc5102b54d94e4"><code>b8dffe9</code></a> docs: remove Numeric and Speakeasy (2+ missed monthly cycles)</li> <li><a href="https://github.com/colinhacks/zod/commit/9195250cab0e7950efe39c3926d6c203b4b0a170"><code>9195250</code></a> docs: remove Mintlify from bronze sponsors (churned)</li> <li><a href="https://github.com/colinhacks/zod/commit/2c703322a21b4e2b12f33f49ea8430c451a68b4f"><code>2c70332</code></a> docs: normalize bronze sponsor logos to github avatar pattern</li> <li><a href="https://github.com/colinhacks/zod/commit/7391be88ac1ee5cd02057f5ccc012a1f5df4efd0"><code>7391be8</code></a> docs: prune lapsed silver/bronze sponsors and add active ones</li> <li><a href="https://github.com/colinhacks/zod/commit/2aeec83eb135e3a83756e973ef44845fc5a455d2"><code>2aeec83</code></a> docs: prune lapsed gold sponsors and rebalance logo sizing</li> <li><a href="https://github.com/colinhacks/zod/commit/4c2fa95ce3f3390fbc522324e406b4e9e89b88f9"><code>4c2fa95</code></a> docs: use Zernio primary wordmark for gold sponsor logo</li> <li>Additional commits viewable in <a href="https://github.com/colinhacks/zod/compare/v3.25.76...v4.4.3">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~GitHub%20Actions">GitHub Actions</a>, a new releaser for zod since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db4defdfbf |
feat: operator-configurable settings visibility via PAPERCLIP_HIDDEN_SETTINGS (#11823)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The instance settings surface (Access, Plugins, Adapters, General,
Experimental) assumes the person at the keyboard operates the whole
instance
> - Operators who host Paperclip for others — a managed cloud or an
internal shared server — expose settings pages and toggles that do not
apply to their deployment, and the related mutation APIs stay open
> - A hosted tenant can open Plugins or Adapters, try an action, and hit
a confusing failure, because only a few hardcoded platform floors exist
> - This pull request adds a generic, operator-configured visibility
mechanism: one env var hides declared settings surfaces in the UI and
floors their mutation routes with a stable 403 code
> - The benefit is a clean hosted-tenant settings surface for any
operator, with zero behavior change for normal self-hosted instances
## Linked Issues or Issue Description
**Subsystem affected**
Instance settings (server routes and UI), the shared settings registry
in `packages/shared`, and the `/api/health` bootstrap payload.
**Problem or motivation**
An operator who hosts Paperclip for other people cannot hide settings
surfaces that the platform manages. Tenants see Access, Plugins, and
Adapters pages, backup retention, and host-level experimental toggles
that do nothing useful for them. The mutation APIs behind these surfaces
also stay open, so a tenant admin can attempt actions the platform must
control. ROADMAP.md names a cleaner shared deployment story as a goal
("Teams should be able to run the same product in hosted or semi-hosted
environments without changing the mental model").
**Proposed solution**
Add a declarative registry of hideable settings surfaces and one env
var, `PAPERCLIP_HIDDEN_SETTINGS`. The server parses the list at boot,
reports it on `/api/health`, and rejects value-changing writes to hidden
surfaces with a stable `settings_operator_managed` 403 code. The UI
reads the list from the health payload and removes the hidden pages,
sections, and toggles from navigation, routes, and page content. Unknown
keys warn and are ignored, so one list can roll across a fleet with
mixed app versions. With the variable unset, behavior is byte-identical
to today.
**Alternatives considered**
- Hardcode the hidden set for cloud instances in this repo: rejected,
because each hosting operator needs a different policy, and policy does
not belong in shared code.
- Deliver the hidden set through the managed-config document: rejected,
because that channel is cloud-specific and fail-closed on unknown
fields; a plain env var works for any operator, including self-hosted
shared servers.
- Lock the controls with a badge instead of hiding them: rejected for
these surfaces, because they are meaningless to tenants, not merely
platform-controlled; the existing managed-overlay lock stays the right
tool for controlled flags.
**Roadmap alignment**
Supports the "shared deployment story" item in ROADMAP.md: hosted and
semi-hosted deployments keep the same product with a settings surface
that matches what the tenant can actually do.
## What Changed
- New `packages/shared/src/settings-visibility.ts`: registry of hideable
surfaces (every instance settings page — profile, environments, access,
heartbeats, experimental, plugins, adapters; every Instance → General
section; every experimental flag as `instance.experimental.<key>`), the
`PAPERCLIP_HIDDEN_SETTINGS` parser, and the `settings_operator_managed`
error code. The General page stays visible as the settings root and
redirect target.
- New `server/src/services/settings-visibility.ts`: parse-once accessor;
unknown keys log one warning and are ignored.
- `/api/health` reports `hiddenSettings` on every response shape; the
field is omitted when nothing is hidden.
- Server floors on hidden surfaces, with same-value echo tolerance (the
`executionMode` precedent): field-backed general sections and
experimental keys reject value-changing PATCHes, and hiding the whole
Experimental page floors every toggle; plugin lifecycle and config
writes, adapter management writes, and the Access admin routes (reads
included) return 403 `settings_operator_managed`. Reads the app itself
needs (plugin `ui-contributions`, adapter metadata, plugin job trigger)
stay open. Pages without instance-scoped mutation routes are hidden in
the UI only.
- UI: new `useHiddenSettings` hook and `HiddenSettingsPageGate` route
gate (hidden pages redirect to the settings root); the settings sidebar
and tab bar drop hidden entries; remembered settings paths remap to the
default page; `InstanceGeneralSettings` skips hidden sections; every
`ExperimentalToggleCard` now carries its flag key and renders nothing
when hidden.
- Removed the dead `InstanceSidebar` component (referenced only by its
own test).
- Docs: `docs/deploy/environment-variables.md` documents the variable
and the key registry.
## Verification
- `pnpm vitest run` over the new and extended suites: shared registry
and parser, representative floor tests per route class (changed-value
403, same-value echo 200, unset env 200, page-level Experimental
hiding), the health field, the route gate, nav filtering, and
section/card hiding with one hidden example per surface kind — 168 tests
pass.
- Full root `pnpm typecheck` passes.
- Manual: booted a server with the variable set. `/api/health` lists the
keys; an unknown key logs one warning and the server boots; hidden pages
redirect; hidden sections and cards do not render; hidden-field PATCH
returns 403 with `details.code = "settings_operator_managed"`; a
same-value echo returns 200. Unset the variable: the full settings
surface returns and responses are byte-identical to master.
## Risks
- Low risk for self-hosted instances: with the variable unset, the
hidden set is empty, the health field is omitted, and no floor
activates.
- Flooring plugin config writes assumes hosted deployments configure
plugins through the platform. If a future bundled plugin needs
tenant-entered config, the floor needs a narrow carve-out.
- Hidden-key floors tolerate same-value echoes, so API clients that
round-trip full GET responses keep working.
- Hiding a toggle does not change its value; operators pair hiding with
the desired default where the value matters.
## Model Used
Claude Fable 5 (Anthropic, `claude-fable-5`) with extended thinking and
agentic tool use, driven through the Claude Code CLI (file edits, test
execution, and live-server verification loops).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
0fa318b8da | feat(artifacts): bridge Markdown work products into the document review surface (#11822) | ||
|
|
c2cfd55e97 |
fix: exclude thought text from automatic issue comments (#11801)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat system records agent runs and can add a run summary to an issue. > - The ACPX engine receives output text and internal thought text as separate streams. > - The default summary strategy joined both streams and could publish internal text in an issue comment. > - Paperclip already has final-output segmentation for run summaries. > - This pull request makes final-output-only summaries mandatory and removes the configuration bypass. > - The benefit is that automatic issue comments contain the intended final message instead of internal execution text. ## Linked Issues or Issue Description Refs #11761 **What happened?** The ACPX engine used the full summary strategy when an adapter did not set `summaryStrategy`. That strategy joined all text deltas, including thought-stream text and intermediate narration. The heartbeat finalizer could then store that summary as an issue comment. **Expected behavior** An automatic issue comment must use only the final output segment. Configuration must not allow thought-stream text or intermediate narration into that summary. **Steps to reproduce** 1. Run an ACPX adapter without a configured `summaryStrategy`. 2. Emit an output delta, a thought delta, a tool call, and a final output delta. 3. Read the generated run summary. 4. Observe that the old default included all text deltas. **Paperclip version or commit** `54b8bec44417511c623999613f9f1006f8af0517` **Deployment mode** Built from source with a local ACPX adapter. ## What Changed - Limit ACPX run summaries to the final non-empty output segment. - Ignore the legacy full-summary setting so configuration cannot bypass containment. - Update regression tests for the safe default and an attempted unsafe override. ## Verification - Observed the new guard fail before the implementation change because the summary contained thought text. - Ran `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "defaults run summaries to the final output segment without thought text|does not allow configuration to include thought text in run summaries"`. Result: 2 passed. - Ran `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts`. Result: 130 passed. - Ran `pnpm --filter @paperclipai/adapter-utils typecheck`. Result: passed. ## Risks - Run summaries are shorter for adapters that relied on full text aggregation. - The old `summaryStrategy: "full"` setting no longer changes summary behavior. This is an intentional containment change. - The change does not alter run logs or tool events. It changes only the summary selected for downstream use. > This is a focused security and privacy bug fix. It does not add roadmap scope. ## Model Used - OpenAI Codex on the GPT-5 family. The runtime did not expose the exact model ID or context-window size. Reasoning, tool use, terminal execution, and code editing were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant inline documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
67ce516a84 |
build(deps): bump googleapis from 164.1.0 to 174.0.1 (#11723)
Bumps [googleapis](https://github.com/googleapis/google-api-nodejs-client) from 164.1.0 to 174.0.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/googleapis/google-api-nodejs-client/releases">googleapis's releases</a>.</em></p> <blockquote> <h2>googleapis: v174.0.1</h2> <h2><a href="https://github.com/googleapis/google-api-nodejs-client/compare/googleapis-v174.0.0...googleapis-v174.0.1">174.0.1</a> (2026-08-05)</h2> <h3>Bug Fixes</h3> <ul> <li><strong>generator:</strong> respect ignore.json when downloading discovery docs and cleaning up old files (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3965">#3965</a>) (<a href="https://github.com/googleapis/google-api-nodejs-client/commit/7c7bfdd9ce8aff2d8b590994d6bc56573554eb14">7c7bfdd</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/3285a2c50bb5b2b52127fdb5fdef41315fe87364"><code>3285a2c</code></a> chore: release main (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3966">#3966</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/51c4d51eecc2cc9c105a5b3aae0ec35d6030dae9"><code>51c4d51</code></a> chore: ignore broken aiplatform:v1 discovery schema (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3969">#3969</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/7c7bfdd9ce8aff2d8b590994d6bc56573554eb14"><code>7c7bfdd</code></a> fix(generator): respect ignore.json when downloading discovery docs and clean...</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/9967f25bc93acf8c0760796fa2a248fc61f74e83"><code>9967f25</code></a> chore(generator): only commit index files if there are staged changes (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3962">#3962</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/8525611d9cec2017432eca8814a829353f5d8c73"><code>8525611</code></a> chore: ignore broken aiplatform:v1beta1 and analytics:v3 discovery schemas (#...</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/306daab95944d37590fa5bfe48ac7052aad85ecb"><code>306daab</code></a> chore(test): increase kitchen sink system test timeout to 10m (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3960">#3960</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/00487776c3500c407b031237be8de3e156380743"><code>0048777</code></a> chore: release main (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3952">#3952</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/ac153108d76743e8caaa0e4c013bb802d3f5121f"><code>ac15310</code></a> feat: run the generator (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3959">#3959</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/c429a9b6f8272f1a651084f83ffcf4b88124c59f"><code>c429a9b</code></a> feat: run the generator (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3955">#3955</a>)</li> <li><a href="https://github.com/googleapis/google-api-nodejs-client/commit/5950a04a8c33944c5e594e48959868f288b0996e"><code>5950a04</code></a> fix: prevent OOMs and HTTP 408 timeouts during API generation and push (<a href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/3954">#3954</a>)</li> <li>Additional commits viewable in <a href="https://github.com/googleapis/google-api-nodejs-client/compare/googleapis-v164.1.0...googleapis-v174.0.1">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cb0009b097 |
fix: preserve recovery retries across restarts (#11817)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane must keep each active issue on a clear execution or recovery path. > - A missing issue disposition can require more than one bounded repair attempt. > - A server restart could lose that repair path or move source ownership to the recovery owner. > - A parked or expired retry could also make the user interface show a false healthy state. > - Concurrent recovery loops must not schedule the same repair attempt twice. > - This pull request keeps retry state durable, makes scheduling atomic, and keeps source ownership stable. > - The benefit is that recovery continues after a restart and operators see the correct state. ## Linked Issues or Issue Description **What happened?** A run that ended without a valid issue disposition could lose its repair path after a server restart. Manager recovery could also change the source owner. In addition, a parked or expired retry could make the issue look healthy when no active work existed. Concurrent reconciliation could also schedule the same repair attempt twice. **Expected behavior** Paperclip must keep bounded source and manager repair attempts across restarts. Recovery ownership must stay separate from source issue ownership. The server and user interface must report only a live retry as active work. Each repair attempt must be scheduled at most once per company. **Steps to reproduce** 1. Start an agent run on an issue. 2. End the run without a valid issue disposition. 3. Let the first repair attempt schedule a retry. 4. Restart the server, let the retry time pass without a live run, or start two reconciliation loops together. 5. Observe that the repair path can stop, the issue can show a false healthy state, or duplicate retries can be created. **Paperclip version or commit** The problem existed on `master` before candidate head `d8e620fe86bade7df18decac332007f5821ae04f`. **Deployment mode** The problem affects self-hosted servers and local builds that use automatic recovery. ## What Changed - Persist bounded source-owner and manager repair lineages with stable fingerprints and retry limits. - Resume incomplete disposition repairs after a server restart. - Keep recovery ownership separate from source issue ownership and enforce source mutation authority. - Project live retry evidence into issue and blocker summaries. - Show recovery owner, return owner, attempt count, and retry state in the board user interface. - Treat expired or parked retries as attention states unless a queued or running attempt exists. - Atomically deduplicate disposition-repair wake requests with a company-scoped partial unique index. - Reuse the winning run when concurrent reconciliation loses the uniqueness race, without duplicate scheduling activity. - Honor disabled on-demand wake policy before recovery scheduling and again before delayed retry promotion. - Keep the new index migration safe for lagging seeded databases that already contain the index. - Add server and user interface tests for recovery, restart, ownership, retry, concurrency, and blocker states. - Update the implementation and execution semantics documents. ## Verification - Focused server recovery and ownership suites: 282 tests passed on the repaired base candidate. - Focused user interface recovery suites: 128 tests passed on the repaired base candidate. - Atomic-deduplication schema and recovery suites: 111 tests passed on the first Greptile repair. - Recovery and scheduled-retry wake-policy suites: 126 tests passed at `d8e620fe86bade7df18decac332007f5821ae04f`. - The exact lagging-source migration-order test passed after the index migration became idempotent: 1 test passed and 62 unrelated tests were skipped. - `@paperclipai/db` and `@paperclipai/server` typechecks passed at the current head. - Migration generation and migration safety checks passed for migration `0226_tan_colossus.sql`. - `pnpm check:token-gates` passed on the repaired base candidate. - `pnpm -r typecheck` passed on the repaired base candidate. - `pnpm build` passed on the repaired base candidate. - `pnpm test:run` passed 4,540 tests on the repaired base candidate. Four fixed-port cases met listeners that already existed on the host. - The two unchanged fixed-port files passed in an isolated network namespace: 129 tests passed and 27 tests were skipped. - Independent Security and QA reviews approved `63c0423aab54c66f2293a20b0fb3f3b013ee3ba8`; exact-head re-review is required after automated checks settle on `d8e620fe86bade7df18decac332007f5821ae04f`. ## Risks - Recovery orchestration affects issue liveness and ownership. The new paths use bounded attempts, stable fingerprints, row locks, authority checks, and database uniqueness. - A conservative attention state can show more warnings when a scheduled retry has no queued or running attempt. It does not hide stopped work. - Migration `0226_tan_colossus.sql` creates a partial unique index on a known-large table. Migrations run transactionally, so `CONCURRENTLY` is unavailable. The matching disposition-repair key namespace is introduced by this release, so deployed databases have no matching rows before the index is added. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex from the GPT-5 model family used agentic reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window. - Anthropic Claude Opus 5 used a 1M context window, tool use, and code execution for part of the user interface repair, as recorded in the commit history. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
233c12f029 |
feat: add kimi-local adapter for Kimi Code CLI (CLI + ACP engines) (#9967)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters (`claude_local`, `gemini_local`, `grok_local`, …) are the integration surface that lets Paperclip run coding CLIs on the host machine > - The Kimi Code CLI (`kimi`, Moonshot AI) has a documented non-interactive mode, `kimi -p --output-format stream-json` with session resume via `kimi -r`, but Paperclip has no built-in adapter for it > - So Kimi users (especially Kimi membership / OAuth subscribers) cannot onboard their CLI to Paperclip agent teams > - This pull request adds a complete built-in `kimi_local` adapter (both execution engines, session management, instructions + skills delivery, thinking-effort control, environment test, UI and CLI modules, docs) following the established `gemini_local`/`grok_local` package pattern > - Kimi Code ships an ACP server (`kimi acp`), so the adapter runs on Paperclip's shared acpx engine by default (streaming transcript with live tool status, like `claude_local`/`gemini_local`) and falls back to a headless CLI lane (`kimi -p --output-format stream-json`) when ACP prerequisites are unavailable > - The benefit is that Kimi Code becomes a first-class Paperclip agent lane: selectable in the UI, resumable across heartbeats, with the same operating context (instruction bundle, skills, effort) and streaming transcript the other local adapters get ## Linked Issues or Issue Description - Supersedes #9880 (same branch; expanded from the CLI-only lane into a complete adapter with the default ACP engine lane, control-plane skill install, and live transcript wiring) - Refs #9879 (adapter request for Kimi Code CLI, filed with this PR) - Refs #163 (original Kimi support request) Duplicate/related prior PRs, per the dedup search (both appear stale: no updates or maintainer review since May 2026, and both target an older Kimi CLI interface; calling them out for reviewer context per CONTRIBUTING.md): - Refs #6276 (`feat: add kimi-local adapter`): targets an older array-based content format (`{type: think}`/`{type: text}` blocks), not the current documented stream-json schema - Refs #5202 (`feat(adapter): add Kimi CLI local adapter with Wire protocol support`): builds on a `--wire` JSON-RPC interface that current Kimi Code CLI (0.27.0) no longer documents; the current documented headless interface is `-p --output-format stream-json` This PR is a fresh implementation against current master and the currently documented/verified Kimi CLI behavior (see Verification). Happy to fold in anything useful from the earlier attempts if a reviewer prefers. ## What Changed - **New adapter package** `packages/adapters/kimi-local` (`@paperclipai/adapter-kimi-local`), modeled on `gemini-local`/`grok-local`: - `src/server/execute.ts`: spawns `kimi -p <prompt> --output-format stream-json` (argv array, no shell), `-m <model>` only when configured, `-r <sessionId>` when the stored session cwd matches the run cwd, automatic fresh-session retry on unrecoverable-session errors, headless-safe env (`CI=1`, `NO_COLOR=1`, `KIMI_CODE_NO_AUTO_UPDATE=1`, `TERM=dumb`; user-configured values win), full remote (ssh/sandbox) execution lane with runtime install via `@moonshot-ai/kimi-code` - **Instruction bundle delivery**: the prompt path directive now names the sibling instruction files (`./HEARTBEAT.md`, `./SOUL.md`, `./TOOLS.md`) alongside the prepended entry file, and local runs pass `--add-dir <instructions-dir>` so Kimi can actually open them (matching `claude_local`). Without this, only the entry file reached Kimi and agents improvised the operating workflow that `HEARTBEAT.md` documents - **Thinking effort**: a configured `effort` is forwarded as the `KIMI_MODEL_THINKING_EFFORT` operational override (Kimi has no per-invocation effort flag). It is only sent for models that advertise `support_efforts` (currently `kimi-code/k3`) to avoid provider rejections, and `medium` maps to `high` since Kimi has no medium tier (`low`/`high`/`max` pass through) - **Skills delivery**: desired Paperclip skills are delivered via Kimi's `--skills-dir` flag from a dedicated per-run directory (a local snapshot, or the synced snapshot on remote targets), so skills load reliably and in isolation. Paperclip never overwrites the shared `$KIMI_CODE_HOME/skills` home, so skills installed by the operator or other agents are left intact. `--skills-dir` is only passed when at least one skill is desired, so unconfigured agents keep Kimi's default skill discovery - **Live run status**: the adapter now forwards each streamed stream-json line to `onEvent` (assistant `content` as an assistant snippet, `tool_calls` as tool-name events), which drives the issue-thread activity indicator (`currentToolName` / `lastAssistantSnippet` / `lastEventAt`). Previously the adapter only wrote the raw run log, so the issue thread showed a stale "no output for N s" line with no tool or reasoning context while Kimi worked. Tool results are omitted so the last meaningful "Using X" / snippet is not overwritten by a generic label - `src/server/parse.ts`: parses the verified Kimi stream-json event shapes (`assistant` text, `assistant.tool_calls` with JSON-string arguments, `tool` results, trailing `meta.session.resume_hint` for session-id capture) plus failure classifiers (`kimi_auth_required`, transient network, unrecoverable session). A signaled exit (null exit code, not a timeout) is now reported as a failure rather than coalesced to success, and the error message names the terminating signal - `src/server/skills.ts`: lists/syncs Paperclip skills for the adapter's skill-management surface - `src/server/test.ts`: environment test covering CLI resolution + `kimi --version`, cwd check, auth detection (OAuth credential dirs, keyed `[providers.*]` in config.toml, or the `KIMI_MODEL_NAME` + `KIMI_MODEL_API_KEY` env pair), and a live hello probe - `src/ui/` (stdout-line parser for transcripts, config builder) and `src/cli/` (stream event formatter) modules - Root metadata: three managed model aliases (`kimi-code/kimi-for-coding`, `kimi-code/kimi-for-coding-highspeed`, `kimi-code/k3`), effort-capable-model metadata (`EFFORT_CAPABLE_MODELS`, effort mapping helpers), `agentConfigurationDoc` - Tests: 101 tests across parse, execute (args building, resume gating, retry, auth error code, timeout, signaled-exit failure, effort forwarding/gating/mapping, `--add-dir` instructions directive, `--skills-dir` gating, `onEvent` runtime-event forwarding), ACP engine (engine resolution, acpx config build, node-version gate), ACP transcript delegation, environment test, UI parse/build-config - **ACP engine lane (default)** (`src/server/acp.ts` + shared `adapter-utils/acpx-engine`): Kimi Code ships an ACP server (`kimi acp`), so `kimi_local` now runs on Paperclip's shared acpx engine by default, matching `claude_local`/`codex_local`/`gemini_local`. The issue-thread transcript streams live (assistant text deltas, tool calls with a `pending`->`completed` status lifecycle) instead of the CLI lane's bursty complete-message output. Registered `kimi_local -> "kimi"` in `ACPX_ADAPTER_AGENT_IDS` and resolved the built-in agent command to `kimi acp`; `execute.ts` dispatches to the ACP executor first with an automatic CLI fallback when ACP prerequisites fail (`engine=acp` requires ACP, `engine=cli` pins the headless lane); `index.ts` falls back to the shared acpx session codec; the UI/CLI delegate `acpx.*` events to the shared acpx transcript parser and event formatter. The headless CLI lane (above) remains as the fallback - **Registration** (one entry each, mirroring existing adapters): server adapter registry + `BUILTIN_ADAPTER_TYPES`, `AGENT_ADAPTER_TYPES` (shared), UI adapter registry + display registry (`Kimi Code`, Moon icon) + capabilities defaults, CLI adapter registry, `Dockerfile` (package copy + `npm install --global @moonshot-ai/kimi-code@latest`), `vitest.config.ts` workspace, `scripts/release-package-manifest.json` - **Behavioral sets** mirroring `gemini_local` (Kimi resumes sessions the same way): `GIT_SENSITIVE_LOCAL_ADAPTER_TYPES`, `SESSIONED_LOCAL_ADAPTERS` (heartbeat + recovery), `REMOTE_MANAGED_ADAPTERS`, ssh/sandbox execution-target allow-lists, `ADAPTER_DEFAULT_RULES_BY_TYPE` (`timeoutSec: 0`, `graceSec: 15`), and `LEGACY_SESSIONED_ADAPTER_TYPES` + `ADAPTER_SESSION_MANAGEMENT` in adapter-utils - **UI touch-points**: New Agent default-model branch, AgentConfigForm command map (`kimi_local: "kimi"`) + model defaults + a Kimi-specific thinking-effort option list (`Low`/`High`/`Max`, reflecting Kimi's tiers rather than borrowing Claude's), OnboardingWizard (command map, model default, `kimi login` / `KIMI_MODEL_NAME + KIMI_MODEL_API_KEY` auth hints, manual-debug command line), InviteLanding enabled adapters - **Control-plane skill install** (`cli/src/commands/client/agent.ts`): `paperclipai agent local-cli` seeded the Paperclip control-plane skills into `~/.codex/skills` and `~/.claude/skills` so Codex/Claude agents auto-discover the API reference every run. Kimi had no equivalent target, so `kimi_local` agents began each session without the control-plane skill and rediscovered routes (e.g. the company-scoped `POST /api/companies/{companyId}/issues`) by trial and error. Added `~/.kimi-code/skills` (honoring `KIMI_CODE_HOME`) as a third install target for parity. Independent of the per-run `--skills-dir` delivery, which only applies to explicitly configured skills. - **Docs**: `docs/adapters/kimi-local.md` (prerequisites, auth options, config fields including `effort`, session resume, instruction bundle, skills delivery, control-plane skill install) + a row in `docs/adapters/overview.md` Out of scope (deliberately): model profiles, built-in agent `allowedAdapterTypes` additions. ## Verification\n\nCurrent-master rebase verification (OpenAI Codex, 2026-08-03): 13 focused files / 231 tests pass; adapter-utils, server, UI, CLI, and Kimi adapter typechecks pass; full repository build and UI token gates pass. The branch is conflict-free against master at head `1249df117c5e12e5771b9a570a6340866450619e`.\n\nAutomated (all from repo root, pnpm 9.15.4, Node 22): - `vitest run packages/adapters/kimi-local`: 89/89 pass (includes coverage for the instruction `--add-dir` directive, effort forwarding/gating/mapping, `--skills-dir` gating, the signaled-exit failure path, and `onEvent` runtime-event forwarding with cross-chunk line buffering) - `vitest run server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/services/heartbeat-stop-metadata.test.ts ui/src/adapters/adapter-display-registry.test.ts`: 37/37 pass - `vitest run cli/src/__tests__/skills.test.ts`: 13/13 pass (the control-plane skill install target follows the existing Codex/Claude install path, whose symlink logic is unchanged) - `vitest run packages/shared`: 307/307 pass; `vitest run packages/adapter-utils`: pass except one pre-existing, unrelated failure (`mcp-isolation.integration.test.ts` requires Claude CLI ≥ 2.1.207; host has 2.1.185, fails identically on unmodified master) - `pnpm --filter @paperclipai/adapter-kimi-local typecheck|build`, plus typecheck of `server`, `ui`, `cli`, `adapter-utils`: all clean - `pnpm install --frozen-lockfile`: passes (the PR diff itself contains no lockfile changes, per repo policy; verified against a locally regenerated lockfile) - `node scripts/check-no-git-push.mjs` and `node scripts/check-forbidden-tokens.mjs`: pass - CI note: the `policy` job's release-bootstrap step is expected to stay red until a maintainer bootstraps the first npm publish of `@paperclipai/adapter-kimi-local`; see the CI Note for Maintainers comment. All other contributor-actionable checks are green. Manual end-to-end (real Kimi CLI 0.27.0, OAuth login, dev server on an isolated instance): 1. Server `GET /api/adapters` lists `kimi_local` as builtin with correct capability flags; models endpoint returns the three Kimi models 2. `POST .../adapters/kimi_local/test-environment`: all checks pass, including a live `kimi -p` hello probe 3. Created a `kimi_local` agent and invoked two heartbeats: run 1 spawned `kimi -p ... --output-format stream-json`, Kimi used its `Read` tool, produced the expected answer, and the session id was captured from the `session.resume_hint` meta event; run 2 resumed the **same** Kimi session (`sessionIdBefore == sessionIdAfter`) via `-r` 4. UI: adapter appears in the New Agent dropdown; selecting it shows the Kimi command placeholder, the three models, and the Kimi config fields; the run transcript renders Kimi tool calls via the adapter's stdout parser The instruction-bundle, thinking-effort, and `--skills-dir` changes landed after the manual run above. They are covered by the unit tests listed under Automated, and the Kimi CLI flags they rely on (`--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were confirmed against the installed Kimi Code CLI 0.27.0 (`kimi --help`, config-file thinking-effort docs). Screenshots (assets branch on the fork, not part of the diff):       ## Risks - Low risk to existing behavior: the change is additive, one new workspace package plus single-entry registrations alongside existing adapters; no existing adapter code paths are modified. - The adapter invokes the locally installed `kimi` CLI; like other local adapters, run behavior depends on the host's Kimi version. The parser is written against the documented/verified 0.27.0 stream-json schema and degrades gracefully (malformed lines are skipped, failures surface as run errors). - `--skills-dir` overrides Kimi's auto-discovery of user and project skills for the run. This is intentional (paperclip-managed agents get a reproducible, isolated skill set), and it is only passed when at least one Paperclip skill is desired, so unconfigured agents keep default discovery. - Thinking effort is only forwarded to models that advertise `support_efforts` (currently `kimi-code/k3`); `EFFORT_CAPABLE_MODELS` must be extended when more Kimi models gain support, otherwise a configured effort is silently ignored for them. - `Dockerfile` now installs `@moonshot-ai/kimi-code@latest` globally alongside the other agent CLIs, so image size increases slightly. - Maintainer action needed for the npm bootstrap gate: the `policy` job's release-bootstrap step fails until the first npm publish of `@paperclipai/adapter-kimi-local` (the gate from #5146 that every new adapter package has passed through). Enrollment with `publishFromCi: true` is required by the manifest validator (dropping the entry, `false`, or `private` are all rejected), so this is intentionally left to a maintainer. Remaining CI lanes are expected to run once it is done. ## Model Used\n\n- **Current-master rebase, conflict adaptation, and registry-parity coverage:** OpenAI, **GPT-5 Codex** (Codex agent; exact serving model ID and context-window size were not exposed to the runtime), with repository, shell, Git, and GitHub tooling. It preserved Hawik’s commit authorship, reconciled ACPX and environment-capability changes, added current registry tests, and ran the verification above.\n- **Adapter implementation and initial review:** Moonshot AI, **Kimi K3 Coding** (latest), via **Kimi Code CLI v0.27.0** (`kimi-code/k3` alias, 1M-token context window, thinking mode, agentic tool use). The CLI agent explored the repo, wrote the adapter implementation (delegated to a coder sub-agent of the same model), ran tests, and drafted the first version of this PR body. A second model-driven review pass (read-only, same model) audited the diff for security/correctness before submission; its findings (shell-quoting hardening, auth-detection false positive, session-compaction registration, test gaps) were fixed and are included. - **Harness-context fixes and review responses:** Anthropic, **Claude Opus 4.8** (`claude-opus-4-8`) via Claude Code. Diagnosed from run logs that Kimi received only the entry instructions file (not the `HEARTBEAT.md`/`SOUL.md`/`TOOLS.md` bundle) and that `effort` was never wired, then implemented the instruction `--add-dir` delivery, `KIMI_MODEL_THINKING_EFFORT` forwarding, and `--skills-dir` skill delivery, added the accompanying tests and docs, and addressed the automated review comments (preserving external skills on remote sync, treating a signaled exit as a failure). Also extended the `paperclipai agent local-cli` installer to seed the control-plane skills into `~/.kimi-code/skills` for Codex/Claude parity, wired `onEvent` runtime events so the issue-thread activity indicator reflects Kimi's tool and reasoning output live, and built the ACP engine lane (`kimi acp` via the shared acpx engine, default) so the transcript streams with live tool status like the other ACP adapters. The Kimi CLI flags, subcommand, and env var relied on here were verified against the installed Kimi Code CLI 0.27.0. - All CLI behaviors claimed here (`-p`, `--output-format stream-json`, `-r` resume, event shapes, `--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were verified empirically against the installed Kimi CLI, not assumed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green *(only the release-bootstrap step remains red, pending the maintainer npm publish described in Risks)* - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(will address all Greptile comments as they arrive)* - [x] I will address all Greptile and reviewer comments before requesting merge --- ## Maintainer Addendum (2026-08-20) The shared acpx-engine and issue-chat changes (run-summary segmentation, placeholder tool-event coalescing, `ISSUE_CHAT_TRANSCRIPT_MAX_VISIBLE_ENTRIES` 30 → 400, live-reasoning UI) have been **extracted to #11761** so the cross-adapter behavior changes review and revert independently — both commits there preserve @hawikk's authorship. This PR is now the kimi-specific adapter only (60 files, +3,793/−8, essentially pure addition); the only shared-engine touch left is the `kimi acp` command resolution. `publishFromCi` is `true` — the package name is bootstrapped on npm. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Devin Foley <devin@paperclip.ing> |
||
|
|
1d0e826767 |
feat(acpx/ui): adapter-declared capabilities for verbose streaming backends (no behavior change by default) (#11761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters stream their work through the shared acpx engine, which every ACP adapter (`claude_local`, `codex_local`, `gemini_local`, custom ACP) runs on > - Verbose streaming backends break two shared-engine behaviors: the auto-posted run summary concatenates every text delta including the thought stream (a long multi-tool run once auto-posted a ~50k character monologue as an issue comment), and token-by-token tool-argument streaming floods the run log with tens of thousands of placeholder-titled in-progress events per run > - These fixes were developed inside the `kimi_local` adapter PR, where Kimi Code's streaming volume (~16,000 text deltas per run vs ~290 for a comparable Claude run) surfaced both problems > - Changing this behavior for all adapters at once is a fleet-wide risk, and hardcoding adapter identities in shared code does not scale to many adapters (or work at all for externally-shipped plugin adapters) — so the behaviors become invocation-config parameters that an adapter's own acpx config builder sets, with defaults preserving today's behavior byte-for-byte > - The benefit is that the machinery lands fully tested with zero behavior change for existing adapters — pin tests prove it — the engine carries no adapter identities, and any adapter (including `custom_acp` configs for external backends) opts in declaratively ## Linked Issues or Issue Description - Refs #9967 — extracted from the `kimi_local` adapter PR and restructured to be inert by default; the commit preserves the original author's (@hawikk) authorship. **Current behavior** When an acpx-engine run ends without the agent leaving a comment, the auto-posted summary is every streamed text delta concatenated, thought stream included. Backends that stream tool arguments emit tens of thousands of placeholder-titled `in_progress` tool events into the stored run log, pinning the live activity indicator to a generic "tool call". There is no mechanism for an adapter to vary either behavior, and shared code must never branch on adapter identities. **Proposed behavior** Two engine invocation-config parameters, read with behavior-preserving defaults: `summaryStrategy` (`"full"` = existing concatenation, the default; `"lastOutputSegment"` = segment output at tool starts, exclude thought stream, post the last non-empty segment) and `coalescePlaceholderToolUpdates` (`false` = never drop an event, the default; `true` = coalesce placeholder-titled in-progress updates). An adapter opts in from its own acpx config builder — the engine has no per-adapter knowledge, no adapter identity appears anywhere in shared code, and `custom_acp` agent configs can set the same knobs for external verbose backends. **Reason and benefit** Existing adapters are provably unaffected — new pin tests assert the default path's summary and tool-event output byte-for-byte, so any future change that alters behavior for claude/codex/gemini/custom fails the suite. The verbose-backend handling still lands fully tested, activated declaratively by the adapter that needs it (the `kimi_local` adapter PR sets both knobs in its config builder). ## What Changed - `packages/adapter-utils/src/acpx-engine/execute.ts`: the run preparation parses `summaryStrategy` and `coalescePlaceholderToolUpdates` from the invocation config (validated, defaulted); summary accumulation and `emitRuntimeEvent` branch on the prepared values. The default path is the pre-existing code (`textParts.join("")`, no event filtering). `buildAcpxRunSummary` is the exported last-segment strategy. - `packages/adapter-utils/src/acpx-engine/execute.test.ts`: a pin test asserting the default path's exact summary (thought stream included) and full tool-event stream (placeholder-titled in-progress updates present, names restored); opt-in tests for each knob; a `buildAcpxRunSummary` unit test. - `ui/src/adapters/types.ts`: `UIAdapterModule` gains an optional `transcriptPresentation` capability — `maxVisibleEntries` (issue-chat transcript window, default 30) and `liveReasoningView` (`"ticker"` default; `"scrollLog"` renders live reasoning in a scrollable auto-following box with one entry per tool call). - `ui/src/lib/issue-chat-messages.ts` and `ui/src/components/IssueChatThread.tsx`: shared code resolves the hints via `findUIAdapter(adapterType)` with today's defaults as fallback — no adapter identities anywhere. The `scrollLog` rendering component ships here but is unreachable until an adapter declares it. - `ui/src/lib/issue-chat-messages.test.ts`: a capability test registers a synthetic verbose adapter and asserts the wider window; the pre-existing test keeps pinning the default 30-entry window. No adapter declares any of this in this PR — every adapter renders and summarizes exactly as before, and there is no per-adapter data anywhere. The `kimi_local` adapter PR (#9967, stacked on this branch) is the first consumer: it declares `transcriptPresentation` in its own UI module and sets the engine knobs in its own acpx config builder. ## Verification - Engine suite: 130/130 pass (126 existing + 4 new); chat suites (`issue-chat-messages`, `IssueChatThread`, `RunChatSurface`): all pass including the new synthetic-adapter capability test — 237 tests across the touched surfaces - The pin tests are the regression guard: the engine test encodes today's summary text and tool-event sequence for a default-config run, and the existing 30-entry-window test pins the default transcript window, so "nothing changed for Claude/Codex users" is an executable assertion, not a review judgment - `pnpm --filter @paperclipai/adapter-utils --filter @paperclipai/ui typecheck`: clean ## Risks - Low: with no adapter setting the knobs, every code path taken in production is the existing one. The only behavioral surface is additive (an unused strategy and an unused filter), exercised by tests. - The knobs are ordinary invocation-config keys, so a `custom_acp` agent config can also set them — intended: an external verbose backend gets the same handling without code changes. Both knobs only affect that agent's own run summaries and run-log verbosity. ## Model Used - Original implementation authored in #9967 by @hawikk (models documented there: Moonshot AI Kimi K3 Coding via Kimi Code CLI 0.27.0, OpenAI GPT-5 Codex, Anthropic Claude Opus 4.8). The commit preserves that authorship. - Extraction, restructuring into config-declared parameters, pin tests, and verification: Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Hawik <davapa@gmail.com> |
||
|
|
e826188e82 |
refactor(settings): unify settings and speed up exports (#11789)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - Operators use the settings area to control a company and its Paperclip instance. > - The current navigation separates related settings and uses duplicate instance pages. > - Company exports also do independent reads in sequence and do extra work for previews. > - Hardened workspace commands can differ from their saved command after loopback binding. > - This pull request makes these related operator workflows consistent and faster. > - The benefit is one clear settings area, faster exports, and stable runtime command matching. ## Linked Issues or Issue Description Refs #338 Related: #9834 **What existing behavior does this improve?** This improves the company settings UI, company export preparation, and workspace runtime command matching. **Current behavior** Company and instance settings use separate navigation and duplicate pages. Export preparation reads many independent records in sequence. Preview generation can also build an unused organization image. A command with a forced loopback bind can fail to match its saved runtime command. **Proposed behavior** Use one settings navigation and put general instance controls on the company General page. Load independent export data with bounded concurrency, skip unused preview image work, and load the export page only when it is needed. Treat the loopback-bound form of a command as the same runtime command. **Reason and benefit** Operators get one clear settings area. Large company exports need fewer serialized reads. Export previews and initial UI loads do less work. Hardened runtime services remain linked to their saved command definitions. **Breaking changes** The obsolete instance General URL redirects to the unified settings page. Access and Heartbeats remain available, and legacy bookmarks keep their destinations. No API response shape or database schema changes. ## What Changed - Unified company and instance settings navigation and removed duplicate instance settings pages. - Embedded general instance controls in the company General page and kept access-sensitive navigation behavior. - Preserved instance Access and Heartbeats controls in the unified navigation and normalized old bookmarks to those destinations. - Improved environment and access-state handling when workspace seed requests overlap. - Added bounded export reads, a lighter preview path, deferred export preparation, and lazy export-page loading. - Matched loopback-bound runtime commands to their saved command definitions. - Added focused shared, server, and UI regression tests. ## Verification - `pnpm exec vitest run <18 changed test files>`: 18 files and 256 tests passed. - `pnpm check:token-gates`: passed all four token gates. - `pnpm -r typecheck`: passed for all workspace projects. - `pnpm build`: passed for all workspace projects. - `pnpm test:run`: tests ran without a reported failure, but the runner did not close after the server handoff tests. The process closed with status 0 after an interrupt. - Focused latest-head route tests: 2 files and 4 tests passed. - GitHub latest-head checks: all completed without failure. - Greptile: 5/5 with no unresolved review threads. ## Risks - Medium risk: settings routes and navigation changed across several operator roles. - Medium risk: bounded export concurrency increases simultaneous database reads. The limits stay below the normal pool size. - Low risk: runtime command matching accepts only the known Tailscale HTTPS loopback transformation. - No migrations are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with a GPT-5-family coding model. The runtime does not expose the exact deployed model ID or context-window size. Reasoning, tool use, and local code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details Exception: This task requires the existing execution branch. The harness does not permit a branch rename. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9d1f740f0 |
fix(workspaces): seed managed worktrees when the base checkout has no config (#11752)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents do that work in isolated git worktrees, and a managed
worktree runs its own Paperclip instance with a cloned database
> - That clone needs a seed source, and the source must come from
server-owned registration, never from state the workspace itself can
rewrite
> - The seed-source resolver requires the registered base project
workspace to hold its own `.paperclip/config.json`
> - A managed project workspace is a plain `git clone`, and no code
writes that file into it
> - Every isolated worktree provision, deferred seed, and workspace
repair therefore fails on a managed checkout
> - This pull request lets a named source supply the config when the
base checkout has none
> - The benefit is that managed worktrees provision again, and the seed
source stays server-owned
## Linked Issues or Issue Description
No public GitHub issue exists for this problem. It is described below.
**What happened?**
Agent runs that need an isolated worktree fail during provisioning. The
provision command exits with this error (paths redacted):
```
Execution workspace provision command "bash ./scripts/provision-worktree.sh" failed:
Registered base project workspace has no canonical Paperclip config:
<instance-home>/instances/default/projects/<company-id>/<project-id>/<repo>/.paperclip/config.json
```
`resolveRegisteredWorktreeSeedSource` sets `registeredConfigPath` to
`<baseCwd>/.paperclip/config.json` whenever the caller names a
registered base workspace. It then requires that file to exist.
`scripts/provision-worktree.sh` applies the same rule.
A managed project workspace never has that file.
`materializeManagedProjectWorkspace` creates it with `git clone` and a
rename, so the checkout holds repository content only. The control plane
keeps its config at `<home>/instances/<id>/config.json` instead.
The failure reaches three paths: worktree provisioning, deferred seeding
through `worktree ensure-seeded`, and workspace repair.
The behavior changed in #11671. That pull request replaced a fallback
chain with a single hard requirement. Fixture code in
`scripts/__tests__/provision-worktree-self-heal.test.mjs` writes a
config into the fake base workspace, so tests kept passing.
**Expected behavior**
A managed worktree provisions and seeds from the registered source. The
seed manifest still never selects that source.
**Steps to reproduce**
1. Register the Paperclip repository as a project with a `repoUrl`, so
the server materializes a managed checkout.
2. Assign an issue to an agent whose workspace strategy is
`git_worktree`.
3. Watch the workspace operation log for the provision command.
4. The command exits non-zero with the error above.
**Paperclip version or commit**
Reproduced on `master` at
|
||
|
|
ed1db310a7 |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.66.0 to 0.69.0 (#11728)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.66.0 to 0.69.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.69.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.68.0...v0.69.0">0.69.0</a> (2026-08-16)</h2> <h3>Features</h3> <ul> <li>report changed files to AIR (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/450d6b19dc46a041128356a6fa3cfa3ce6a5a382">450d6b1</a>)</li> </ul> <h2>v0.68.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.67.0...v0.68.0">0.68.0</a> (2026-08-14)</h2> <h3>Features</h3> <ul> <li>align typed session failures with AIR protocol (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0581b9cf397ffd88f2830db721c2d8e3689045e4">0581b9c</a>)</li> </ul> <h2>v0.67.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.67.0">0.67.0</a> (2026-08-14)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Update to <code>@anthropic-ai/claude-agent-sdk</code> v0.3.232 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/de0d0e2b7d185c521c3d1ac7b86e4312c919abfe">de0d0e2</a>)</li> <li>expose typed session failures for AIR (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/979">#979</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8157ee113e705750be4eb6cc787bdd12f1db84ff">8157ee1</a>)</li> <li>publish the model fallback as a warning advisory (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/35aaddb3d5b14eca316582ab285af8a660150dae">35aaddb</a>)</li> <li>surface resolved model name in default model option description (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ec73cd8560be7d5e8b9741e404d7d45e17336996">ec73cd8</a>)</li> <li>surface Skill tool calls with name and kind in _meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1f09e9a3cae787b0effdfae0e7205ef7fe22b9dc">1f09e9a</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>preserve task plans across prompts (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/974">#974</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1afa940a2c8c0f4c610f4f64d30c0961642907b0">1afa940</a>)</li> <li>show a pending title while Claude prepares a file (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/978">#978</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3df1ede89f217312bc237124dc1eccc10c860f99">3df1ede</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.68.0...v0.69.0">0.69.0</a> (2026-08-16)</h2> <h3>Features</h3> <ul> <li>report changed files to AIR (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/450d6b19dc46a041128356a6fa3cfa3ce6a5a382">450d6b1</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.67.0...v0.68.0">0.68.0</a> (2026-08-14)</h2> <h3>Features</h3> <ul> <li>align typed session failures with AIR protocol (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0581b9cf397ffd88f2830db721c2d8e3689045e4">0581b9c</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.67.0">0.67.0</a> (2026-08-14)</h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> Update to <code>@anthropic-ai/claude-agent-sdk</code> v0.3.232 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/de0d0e2b7d185c521c3d1ac7b86e4312c919abfe">de0d0e2</a>)</li> <li>expose typed session failures for AIR (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/979">#979</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/8157ee113e705750be4eb6cc787bdd12f1db84ff">8157ee1</a>)</li> <li>publish the model fallback as a warning advisory (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/35aaddb3d5b14eca316582ab285af8a660150dae">35aaddb</a>)</li> <li>surface resolved model name in default model option description (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ec73cd8560be7d5e8b9741e404d7d45e17336996">ec73cd8</a>)</li> <li>surface Skill tool calls with name and kind in _meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1f09e9a3cae787b0effdfae0e7205ef7fe22b9dc">1f09e9a</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>preserve task plans across prompts (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/974">#974</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1afa940a2c8c0f4c610f4f64d30c0961642907b0">1afa940</a>)</li> <li>show a pending title while Claude prepares a file (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/978">#978</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/3df1ede89f217312bc237124dc1eccc10c860f99">3df1ede</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/59a7e9367b3931a50178de4783cf6074b20060cd"><code>59a7e93</code></a> chore(main): release 0.69.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1007">#1007</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/450d6b19dc46a041128356a6fa3cfa3ce6a5a382"><code>450d6b1</code></a> feat: report changed files to AIR (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1001">#1001</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/5de5d4a2c8363e6462c231380c9fc17d80c568cc"><code>5de5d4a</code></a> chore(main): release 0.68.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1000">#1000</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/0581b9cf397ffd88f2830db721c2d8e3689045e4"><code>0581b9c</code></a> feat: align typed session failures with AIR protocol (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/992">#992</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/11dd73b87146478d3d99ed62c2728289db5dff5d"><code>11dd73b</code></a> Restore native provider state after overrides (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/996">#996</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/e4dba808eaf280379a1218280081fdf0346632e1"><code>e4dba80</code></a> chore(main): release 0.67.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/975">#975</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/de0d0e2b7d185c521c3d1ac7b86e4312c919abfe"><code>de0d0e2</code></a> feat(deps): Update to <code>@anthropic-ai/claude-agent-sdk</code> v0.3.232 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/993">#993</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/35aaddb3d5b14eca316582ab285af8a660150dae"><code>35aaddb</code></a> feat: publish the model fallback as a warning advisory (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/990">#990</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/1f09e9a3cae787b0effdfae0e7205ef7fe22b9dc"><code>1f09e9a</code></a> feat: surface Skill tool calls with name and kind in _meta (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/986">#986</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/ec73cd8560be7d5e8b9741e404d7d45e17336996"><code>ec73cd8</code></a> feat: surface resolved model name in default model option description (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/982">#982</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.66.0...v0.69.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
6bf8f1d194 |
build(deps): bump react-dom and @types/react-dom (#11712)
Bumps [react-dom](https://github.com/react/react/tree/HEAD/packages/react-dom) and [@types/react-dom](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/react-dom). These dependencies needed to be updated together. Updates `react-dom` from 19.2.7 to 19.2.8 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/react/react/releases">react-dom's releases</a>.</em></p> <blockquote> <h2>19.2.8 (July 21st, 2026)</h2> <h2>React Server Components</h2> <ul> <li>Performance improvements when decoding (<a href="https://redirect.github.com/facebook/react/pull/37087">#37087</a> by <a href="https://github.com/eps1lon"><code>@eps1lon</code></a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/react/react/commit/1dd4ecbdabf826f527fc9a58c05ea70375b7d170"><code>1dd4ecb</code></a> [FlightReply] Performance improvements when decoding (<a href="https://github.com/react/react/tree/HEAD/packages/react-dom/issues/37087">#37087</a>)</li> <li><a href="https://github.com/react/react/commit/b0d2fdb78bdfae075a7fa02ddcebbf25f90952c2"><code>b0d2fdb</code></a> [19.2.x] Update required references to GitHub repo (<a href="https://github.com/react/react/tree/HEAD/packages/react-dom/issues/36753">#36753</a>)</li> <li>See full diff in <a href="https://github.com/react/react/commits/v19.2.8/packages/react-dom">compare view</a></li> </ul> </details> <br /> Updates `@types/react-dom` from 19.2.3 to 19.2.4 <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/react-dom">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
b4953985c1 |
build(deps): bump react and @types/react (#11721)
Bumps [react](https://github.com/react/react/tree/HEAD/packages/react) and [@types/react](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/react). These dependencies needed to be updated together. Updates `react` from 19.2.7 to 19.2.8 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/react/react/releases">react's releases</a>.</em></p> <blockquote> <h2>19.2.8 (July 21st, 2026)</h2> <h2>React Server Components</h2> <ul> <li>Performance improvements when decoding (<a href="https://redirect.github.com/facebook/react/pull/37087">#37087</a> by <a href="https://github.com/eps1lon"><code>@eps1lon</code></a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/react/react/commit/1dd4ecbdabf826f527fc9a58c05ea70375b7d170"><code>1dd4ecb</code></a> [FlightReply] Performance improvements when decoding (<a href="https://github.com/react/react/tree/HEAD/packages/react/issues/37087">#37087</a>)</li> <li><a href="https://github.com/react/react/commit/b0d2fdb78bdfae075a7fa02ddcebbf25f90952c2"><code>b0d2fdb</code></a> [19.2.x] Update required references to GitHub repo (<a href="https://github.com/react/react/tree/HEAD/packages/react/issues/36753">#36753</a>)</li> <li>See full diff in <a href="https://github.com/react/react/commits/v19.2.8/packages/react">compare view</a></li> </ul> </details> <br /> Updates `@types/react` from 19.2.17 to 19.2.18 <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/react">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
faab2620ad |
feat(sandbox): add a duplex transport for Daytona behind a default-off kill switch (#11750)
## Thinking Path > - Paperclip is an open source app that manages AI agents for work > - Paperclip runs agents in local and remote sandbox environments > - A sandbox needs a bounded channel for commands and asynchronous input > - Daytona needs a real pseudo-terminal transport for this channel > - The sandbox gateway also needs a mode that handles channel loss safely > - This pull request adds the Daytona transport and gateway mode behind a default-off kill switch > - The benefit is a tested foundation for later transport selection ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): sandbox providers, plugin SDK, server settings, and shared types. **Problem or motivation** The merged sandbox protocol has no runtime transport for Daytona. The generated sandbox gateway also has no duplex mode. A later transport-selection change needs both parts and a safe per-run gate. **Proposed solution** Add a Daytona `duplexCommandStream` transport over a raw pseudo-terminal. Add a generated gateway mode named `duplex_v1`. Add the `enableSandboxDuplexBridge` setting with a default value of `false`. Keep transport selection disabled until a later pull request. **Alternatives considered** Keep the protocol unused until the transport-selection change. This would delay provider tests and leave the gateway path without direct coverage. **Roadmap alignment** This change supports the completed Roadmap item for cloud and sandbox agents. It extends the merged sandbox channel foundation in pull request #11738. **Additional context** The Daytona provider remains an untrusted boundary. Deployments must use least-privilege provider credentials and provider-side quota controls. Operators must name an owner for duplex telemetry retention before rollout. ## What Changed - Add the Daytona `duplexCommandStream` capability over a raw pseudo-terminal. - Add a launch wrapper that disables echo and newline translation for NDJSON frames. - Close channels on lease release, destroy, resume of a stopped worker, and worker shutdown. - Declare the capability in the Daytona manifest and set `PLUGIN_VERSION` to `0.1.5`. - Add the worker-to-host notification sink at `ctx.duplexChannel.data` and `ctx.duplexChannel.exit`. - Add the generated sandbox gateway mode `PAPERCLIP_API_BRIDGE_MODE=duplex_v1`. - Add channel-loss results of `409 outcome_indeterminate` and `503 bridge_unavailable`. - Add the per-run setting `enableSandboxDuplexBridge`, with a default value of `false`. - Add unit tests, generated-source codec tests, lifecycle tests, and a credential-gated live Daytona test. ## Verification - Daytona suite: 185 tests pass. - Adapter utilities: 754 tests pass and 4 tests skip. - Plugin SDK: 62 tests pass. - Shared package: 28 tests pass. - Server duplex tests pass. - Shared, plugin SDK, server, and Daytona TypeScript checks pass. - The live Daytona test passes 3 cases when `DAYTONA_API_KEY` is set. - The live Daytona test skips 3 cases without `DAYTONA_API_KEY`. - CI must run the full workspace typecheck, test, and build gates after PR creation. ## Risks - The Daytona control plane and pseudo-terminal remain untrusted boundaries. - The duplex gateway changes behavior only when the mode and per-run setting enable it. - A lost channel fails requests without replay, so callers must handle indeterminate outcomes. - The transport-selection change must require both `duplexCommandStream === true` and `enableSandboxDuplexBridge === true`. - The provider credential and quota limits need operator control before rollout. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d5416fde9a |
fix(runtime): route sandbox git-bundle export through native syncOut (#11749)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox-managed runtimes move files between a host and an isolated sandbox > - The git-bundle export path reads the full bundle into host memory > - The workspace restore path already uses the provider native `syncOut` transfer > - This pull request uses `syncOut` for bundle export and keeps `readFile` as a fallback > - The change reduces host buffering and exposes the transfer to provider tracing ## Linked Issues or Issue Description **What happened?** The sandbox git-bundle export used `client.readFile` even when the provider supported native `syncOut`. The path buffered the full bundle in host memory and used a chunked base64 command loop. **Expected behavior** The export should use one confined native file transfer when the provider supports `syncOut`. Providers without that capability should keep the existing `readFile` fallback. **Steps to reproduce** 1. Prepare a sandbox-managed runtime with native `syncOut` support. 2. Export the sandbox git bundle. 3. Inspect the sync operations and file reads. 4. Confirm that the bundle uses one native file mapping and that the status file still uses `readFile`. **Paperclip version or commit** Commit `3fc88d14be7133867897e87838276e15aba675fd`. **Deployment mode** Built from source. The change applies to sandbox-managed runtime execution. **Additional context** This PR contains a focused bug fix. No public issue matched this change during the duplicate search. ## What Changed - Route bundle export through `nativeSyncOut` when the provider supports it. - Keep the existing `readFile` fallback for providers without native sync support. - Confine the native bundle mapping with `assertSyncOperationsConfined`. - Cover the native path, fallback path, retry path, and confinement checks with unit tests. ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-managed-runtime.test.ts` — 52 tests passed. - `npx tsc --noEmit` in `packages/adapter-utils` — passed. - The PR CI workflow must pass before merge. ## Risks Low risk. The native path runs only when the provider advertises `syncOut`. The existing `readFile` path remains available as a fallback. Native transfer progress reports only start and finish events. ## Model Used OpenAI GPT-5 Codex. Exact runtime model ID: GPT-5 Codex. Context window: not exposed in this run. Capabilities used: code review, shell tools, GitHub CLI, and Paperclip API coordination. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8161244284 |
feat(sandbox): add opt-in duplex command-stream foundation (capability, protocol, bounded host route, frame codec) (#11738)
## Thinking Path > - Paperclip provides a control plane for companies that run AI agents. > - Sandboxed agents need a safe execution path for persistent command streams. > - The existing callback transport does not provide a bounded, generic duplex route. > - The host must control capability access, route identity, protocol limits, and close behavior. > - This pull request adds an opt-in duplex command-stream foundation across the sandbox layers. > - The feature stays inert because no current provider declares the capability. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox command execution needs a persistent host-to-sandbox stream. The current callback bridge uses a file transport and does not provide this generic route. **Proposed solution** Add a fail-closed provider capability, generic worker protocol messages, a host-owned bounded route, cross-layer service mediation, and a versioned newline-delimited frame codec. **Alternatives considered** Keep the file transport and add feature-specific commands. This does not provide one reusable duplex contract or host-owned route bounds. **Roadmap alignment** This work supports the completed Cloud / Sandbox agents roadmap area and the safe autonomy goal in the product definition. **Additional context** The change passed a two-stage security review. The final code review verdict was approve after fixes for active-stream bounds and service-layer capability mediation. ## What Changed - Add the opt-in `duplexCommandStream` provider capability with fail-closed narrowing. - Add duplex open, write, stop, and close requests and data and exit notifications to the plugin worker protocol. - Add a host-owned route with bounds for chunk size, cumulative bytes, lifetime, protocol errors, pending requests, and pre-bind buffering. - Add close acknowledgement handling with worker retirement when the close remains unconfirmed. - Wire `openDuplexChannel` through the execution target, runtime service, and plugin worker. - Add a versioned frame codec with shared wire-compatibility vectors and split UTF-8 handling. ## Verification - `server/src/__tests__/plugin-worker-manager-duplex.test.ts` passes 18 tests. - `server/src/__tests__/environment-execution-target-duplex.test.ts` passes 11 tests. - `packages/adapter-utils/src/duplex-frame-codec.test.ts` passes 38 tests. - `server/src/__tests__/sandbox-capability-contract.test.ts` passes 15 tests. - Setup-token pseudo-terminal regression tests pass 47 tests. - Server TypeScript check passes. - Continuous integration will run the full required test, typecheck, build, and policy checks. ## Risks - Providers that opt into the capability must implement the complete worker protocol. - Route limit defaults can close a stream when a workload exceeds the configured bounds. - The capability remains disabled for current providers, so current production behavior does not change. ## Model Used OpenAI GPT-5 (`gpt-5`), with tool use and code execution. The model reviewed and prepared this pull request from the supplied implementation and verification record. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd059a073d |
fix(workspaces): make managed runtimes reliable across restarts (#11740)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces need isolated databases, ports, and runtime services > - Concurrent workspaces could reuse ports or lose service ownership after a restart > - A markerless worktree also needed seed recovery, but normal markerless instances still needed to boot > - This pull request makes seed, port, and service ownership state explicit and recoverable > - It also checks live process and listener identity before it reclaims shared resources > - The benefit is reliable workspace startup, restart, adoption, and concurrent provisioning ## Linked Issues or Issue Description **What happened?** Managed workspaces could lose runtime service ownership after a control-plane restart. Concurrent worktrees could also reuse a port when their parent paths differed. A seed recovery change made every markerless instance resolve a worktree seed source, so normal instances without a source could not start. **Expected behavior** Paperclip must preserve healthy managed services across restarts. It must reserve unique ports across worktree parents. It must provision a registered markerless worktree, but it must skip seed work for a normal markerless instance. **Steps to reproduce** 1. Start two managed worktrees under different parent paths at the same time. 2. Restart the control plane while a managed service stays alive. 3. Start Paperclip with a config that has no seed markers and no registered worktree source. 4. Observe duplicate port selection, lost service adoption, or a seed-source startup error. **Paperclip version or commit** Current `master` plus the workspace runtime reliability changes in this pull request. **Deployment mode** Local development with managed execution workspaces and embedded Postgres. ## What Changed - Added a shared port registry with lease heartbeats, process identity checks, and live listener probes. - Reserved worktree ports across custom parent paths and repaired duplicate legacy assignments. - Preserved and adopted healthy managed services across control-plane restarts. - Reconciled guest bind modes and verified listener ownership before termination or reuse. - Provisioned registered markerless worktree databases and kept normal markerless instance startup as a no-op. - Added CLI, shared, server, and shell regression tests for seed, port, listener, restart, and adoption behavior. - Updated the worktree development documentation. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts --reporter=verbose` — 63 tests passed. - `pnpm exec vitest run packages/shared/src/worktree-port-registry.test.ts --reporter=verbose` — 5 tests passed. - Focused runtime Vitest set — 199 tests passed across 37 suites. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 10 tests passed. - `git diff --check` passed. ## Risks - Port reservation now depends on lease and process identity data. The fallback listener probe prevents early reclamation when process metadata is incomplete. - Runtime adoption is stricter about bind and owner identity. The tests cover healthy adoption, stale records, PID reuse, and unrelated listeners. - Markerless seed detection now separates registered worktrees from normal instances. The tests cover both paths. - There are no database schema migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
233be4b36c |
feat: parallelize sandbox file-sync behind a provider opt-in capability (#11736)
## Thinking Path > - Paperclip runs AI agents through local and remote execution adapters. > - Sandbox providers move workspace and asset files before and after agent runs. > - Serial file transfers delay startup and teardown when several operations do not depend on each other. > - Providers need an opt-in contract so existing providers keep their serial behavior. > - This pull request adds a bounded scheduler and routes inbound and outbound sync operations through it. > - The benefit is shorter sandbox setup and teardown with stable errors, clear telemetry, and a safe opt-in path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above): packages/shared, packages/adapter-utils, packages/plugins, and server. **Problem or motivation** Sandbox sync processes the workspace, assets, and referenced projects in series. This adds avoidable wait time to agent startup and teardown. **Proposed solution** Add a fail-closed provider capability named concurrentSyncOperations. Use a bounded scheduler with a limit of four operations. Preserve operation order for error reporting. Keep non-opted-in providers on the serial path. **Alternatives considered** Increase the serial transfer speed or add provider-specific schedulers. Those options do not provide one shared contract or stable behavior across providers. **Roadmap alignment** ROADMAP.md lists cloud and sandbox agents as a product area. This change improves sandbox execution without changing the control-plane contract. **Additional context** The Daytona provider opts in. Board trials on this commit showed overlap for inbound sync and outbound restore, with no referenced-project staging failures. ## What Changed - Add the concurrentSyncOperations sandbox capability and fail-closed parsing. - Add a bounded settle-all scheduler with stable input-order errors. - Parallelize inbound workspace, asset, and referenced-project sync operations when the provider opts in. - Parallelize outbound workspace and asset restore operations when the provider opts in. - Surface referenced-project failure text in run logs and server telemetry. - Add Daytona sync spans and the capability declaration. - Preserve in-flight upload scratch tarballs during workspace wipe. - Add unit and regression tests for the scheduler, coordinators, provider behavior, telemetry, and wipe race. ## Verification - Run the adapter-utils and server type checks. - Run the targeted adapter-utils, server, and Daytona test suites. - Run the full automated sweep. - Review six cold Daytona trials, with three serial and three parallel runs. - Confirm that parallel trials show inbound overlap and outbound restore overlap. - Confirm that providers without the capability keep serial behavior. ## Risks - Providers must opt in only when their file operations can run safely at the same time. - A provider that declares the capability incorrectly can expose transfer races. - The scheduler keeps a limit of four to bound resource use. - Providers without the capability keep the prior serial behavior. ## Model Used OpenAI GPT-5 in the Codex runtime. The model used tool calls, code inspection, and GitHub workflow support. The model did not author the implementation commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e0b64529b3 |
feat(auth): normalize agent login in the sandbox onto one session table and a capability contract (#11730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox agents need a safe login path for each supported adapter > - Codex device login and Claude setup-token login used separate session stores and route logic > - Separate stores made session lookup, expiry, and login capability checks harder to keep consistent > - This pull request unifies both flows on one session table and one capability contract > - The benefit is one company-scoped login model with public session identifiers and shared lifecycle rules ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Codex and Claude sandbox login used separate session stores and different route paths. This split increased the risk of inconsistent company scoping, session lookup, and cleanup. **Proposed solution** Use `adapter_auth_sessions` for both login flows. Use public session identifiers for API access. Select login behavior from projected adapter capability data. Share the route spine, lease arguments, runner lifecycle, and reaper rules. **Alternatives considered** Keep two session tables and add matching fixes to both routes. This keeps duplicate logic and does not provide one capability contract, so this pull request uses shared infrastructure. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`. ## What Changed - Unify Codex device login and Claude setup-token login on `adapter_auth_sessions`. - Return and look up sessions with company-scoped public session identifiers. - Enforce one active session for each company, owner, and adapter. - Share the login route spine, sandbox lease arguments, runner lifecycle, and missing-auth check. - Add a standalone setup-token reaper with adapter-specific row selection. - Add optional login capability projection for adapters and drive route and UI selection from that data. - Rename the provider flag to `supportsLoginPty` and validate its deprecated alias. - Remove the old Claude setup-token session table and add the required migrations. ## Verification - Server typecheck passed with `tsc`. - Database typecheck passed. - UI typecheck passed with `tsc -b`. - Codex login service and route suites passed. - Setup-token session, route, and reaper suites passed. - Adapter session schema, plugin validator, capability projection, UI render, and Daytona suites passed. - GitHub Actions must confirm the complete CI gate after pull request creation. ## Risks - The migrations remove short-lived in-flight login rows during deployment. A login that spans the migration can continue until its provider lease expires. - The Codex credential store remains company-scoped. A cross-owner credential race remains a documented, board-accepted risk. - API clients that use internal session row identifiers no longer work. The API accepts only public session identifiers. ## Model Used Codex, GPT-5, exact runtime model ID not exposed in this handoff, large context window, reasoning, and repository tool use. The implementing engineer produced the code with AI assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
98298cb7ff |
build(deps-dev): bump tsx from 4.23.1 to 4.23.12 (#11722)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.1 to 4.23.12. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/privatenumber/tsx/releases">tsx's releases</a>.</em></p> <blockquote> <h2>v4.23.12</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.11...v4.23.12">4.23.12</a> (2026-08-10)</h2> <h3>Bug Fixes</h3> <ul> <li>shim <code>import.meta</code> when tokens are split by comments or newlines (<a href="https://redirect.github.com/privatenumber/tsx/issues/829">#829</a>) (<a href="https://github.com/privatenumber/tsx/commit/ed9d33046a135de13a35fdfce12368b79d1b1518">ed9d330</a>), closes <a href="https://redirect.github.com/privatenumber/tsx/issues/828">#828</a></li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.12"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.11</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.10...v4.23.11">4.23.11</a> (2026-08-07)</h2> <h3>Bug Fixes</h3> <ul> <li>preserve async ESM require fallback (<a href="https://github.com/privatenumber/tsx/commit/55cbecef8ebe839c7110e8c141a1c3bc4da326cd">55cbece</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.11"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.10</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.9...v4.23.10">4.23.10</a> (2026-08-07)</h2> <h3>Bug Fixes</h3> <ul> <li>support nyc coverage discovery (<a href="https://redirect.github.com/privatenumber/tsx/issues/710">#710</a>) (<a href="https://github.com/privatenumber/tsx/commit/ec1bcd5f711e5159b67cb0aea211f06cf2cfce8a">ec1bcd5</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.10"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.9</h2> <h2><a href="https://github.com/privatenumber/tsx/compare/v4.23.8...v4.23.9">4.23.9</a> (2026-08-06)</h2> <h3>Bug Fixes</h3> <ul> <li>map Node test locations (<a href="https://github.com/privatenumber/tsx/commit/2f55884195a8c745fbe64a0288de69bc062ed876">2f55884</a>)</li> <li>support data URLs in tsImport (<a href="https://github.com/privatenumber/tsx/commit/b94f46f6b6a7e6dc575624b0ecc7124318723056">b94f46f</a>)</li> </ul> <hr /> <p>This release is also available on:</p> <ul> <li><a href="https://www.npmjs.com/package/tsx/v/4.23.9"><code>npm package (@latest dist-tag)</code></a></li> </ul> <h2>v4.23.8</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/privatenumber/tsx/commit/ed9d33046a135de13a35fdfce12368b79d1b1518"><code>ed9d330</code></a> fix: shim <code>import.meta</code> when tokens are split by comments or newlines (<a href="https://redirect.github.com/privatenumber/tsx/issues/829">#829</a>)</li> <li><a href="https://github.com/privatenumber/tsx/commit/651f5bec70d9a116d1fe1706000b7ca011a516fc"><code>651f5be</code></a> test: cover CommonJS TypeScript import.meta paths</li> <li><a href="https://github.com/privatenumber/tsx/commit/bd3bc6448e957c1172eb91a0584ec7fec2d6a7ad"><code>bd3bc64</code></a> test: cover CommonJS loader source fallback</li> <li><a href="https://github.com/privatenumber/tsx/commit/55cbecef8ebe839c7110e8c141a1c3bc4da326cd"><code>55cbece</code></a> fix: preserve async ESM require fallback</li> <li><a href="https://github.com/privatenumber/tsx/commit/6c5ba85f7a1e57f06bcb760718d6ba3b978e5a05"><code>6c5ba85</code></a> docs: document CommonJS default interop</li> <li><a href="https://github.com/privatenumber/tsx/commit/ec1bcd5f711e5159b67cb0aea211f06cf2cfce8a"><code>ec1bcd5</code></a> fix: support nyc coverage discovery (<a href="https://redirect.github.com/privatenumber/tsx/issues/710">#710</a>)</li> <li><a href="https://github.com/privatenumber/tsx/commit/b6e5b48a7b0fa4639e67119c8b02150ec8c6cef7"><code>b6e5b48</code></a> docs: clarify CommonJS default imports</li> <li><a href="https://github.com/privatenumber/tsx/commit/2f55884195a8c745fbe64a0288de69bc062ed876"><code>2f55884</code></a> fix: map Node test locations</li> <li><a href="https://github.com/privatenumber/tsx/commit/de935d588b6c0959ac8af71e2cf40d1a61e659e7"><code>de935d5</code></a> docs: document Node source-map stack formatting</li> <li><a href="https://github.com/privatenumber/tsx/commit/b94f46f6b6a7e6dc575624b0ecc7124318723056"><code>b94f46f</code></a> fix: support data URLs in tsImport</li> <li>Additional commits viewable in <a href="https://github.com/privatenumber/tsx/compare/v4.23.1...v4.23.12">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
ec8d7dd21f |
build(deps): bump ws from 8.21.1 to 8.21.3 (#11729)
Bumps [ws](https://github.com/websockets/ws) from 8.21.1 to 8.21.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/websockets/ws/releases">ws's releases</a>.</em></p> <blockquote> <h2>8.21.3</h2> <h1>Bug fixes</h1> <ul> <li>The server now correctly rejects permessage-deflate offers if the incoming <code>client_max_window_bits</code> parameter value is smaller than its configured <code>clientMaxWindowBits</code> (e97a20ea).</li> </ul> <h2>8.21.2</h2> <h1>Bug fixes</h1> <ul> <li>Fixed a test for <a href="https://github.com/nodejs/citgm">CITGM</a> (2eb3be0b).</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/websockets/ws/commit/c791e707eab3c13dd9a261d2479c3cc4a49a6fed"><code>c791e70</code></a> [dist] 8.21.3</li> <li><a href="https://github.com/websockets/ws/commit/e97a20eaa6f2ad7969419eed732a506453251eb9"><code>e97a20e</code></a> [fix] Reject offers with <code>client_max_window_bits</code> below config</li> <li><a href="https://github.com/websockets/ws/commit/787ebf22ce3d091fb6f931d20b4c7e914ba7cf85"><code>787ebf2</code></a> [dist] 8.21.2</li> <li><a href="https://github.com/websockets/ws/commit/b4d62ebad40c3b925c84ff305a47975406015422"><code>b4d62eb</code></a> Revert "[ci] Trust Coveralls Homebrew tap"</li> <li><a href="https://github.com/websockets/ws/commit/e4bb883723a0c18452eea10a74139901ae33c61d"><code>e4bb883</code></a> [security] Use GitHub PVR as main reporting channel</li> <li><a href="https://github.com/websockets/ws/commit/2eb3be0bff2453e2654b1315c5872e8d5d424a50"><code>2eb3be0</code></a> [test] Skip test on Node.js versions where it does not apply</li> <li>See full diff in <a href="https://github.com/websockets/ws/compare/8.21.1...8.21.3">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e1df4c6068 |
fix(workspaces): keep deferred seed databases reliable (#11706)
## Thinking Path > - Paperclip manages agent work in isolated execution workspaces. > - A workspace depends on a valid database seed before it can run. > - Deferred seed failures were hidden behind a successful provision status. > - The seed restore also had two possible owners for the embedded PostgreSQL process. > - That allowed the target database to stop while the restore was still running. > - This pull request makes seed failures visible and gives the seed process sole lifecycle ownership. > - The benefit is that workspace provisioning reports the real result and does not stop its own target database. ## Linked Issues or Issue Description Related: #11684 **What happened?** Initial worktree provisioning could report success before its deferred database seed completed. The seed restore could also reuse a target embedded PostgreSQL process with another shutdown owner. This could stop the target database during the restore. **Expected behavior** Workspace status must show a failed deferred seed as a failure. The seed restore must own the target embedded PostgreSQL process until restore, migration, and validation finish. **Steps to reproduce** 1. Provision a worktree with deferred database seeding. 2. Make the seed manifest end in a failed state while the command exits with code 0. 3. Observe that the provision status remains successful on `master`. 4. Start a seed restore against an already-running target embedded PostgreSQL process. 5. Observe that another lifecycle owner can stop the target during restore. **Paperclip version or commit** `51a843e135` **Deployment mode** Local dev with execution workspaces and embedded PostgreSQL. ## What Changed - Add a first-class `workspace_seed` operation for deferred database seeds. - Require terminal, verified seed evidence before the seed operation succeeds. - Surface the seed phase and failure metadata in workspace status and UI state. - Give the seed process exclusive lifecycle ownership of the target embedded PostgreSQL process. - Suppress imported embedded-Postgres exit hooks without removing existing host listeners. - Record a credential-safe shutdown diagnostic in failed seed manifests. ## Verification - The original deferred-seed commit passed 4 server tests, 24 workspace-status UI tests, shared/server/UI typechecks, and the UI token gate. - The original PostgreSQL-lifecycle commit passed 3 lifecycle tests, 3 ownership/diagnostic tests, 1 real embedded-Postgres seed integration, and the affected package typechecks. - No local tests were rerun after the clean cherry-pick because the operator requested the shortest landing path. - Review the automatic PR checks for the clean `origin/master` replay. ## Risks - A live target database now causes an early error instead of being reused. The error includes recovery guidance. - Workspace consumers must handle the new `workspace_seed` operation type. Shared types and UI state handling are updated in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, high-reasoning mode, with repository, shell, and GitHub tool use. The runtime does not expose a more specific deployment suffix or context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
536d5880c5 |
build(deps): bump @cursor/sdk from 1.0.24 to 1.0.28 (#11520)
Bumps [@cursor/sdk](https://github.com/cursor/cursor) from 1.0.24 to 1.0.28. <details> <summary>Commits</summary> <ul> <li>See full diff in <a href="https://github.com/cursor/cursor/commits">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
03ddca4df7 |
build(deps): bump @agentclientprotocol/codex-acp from 1.1.7 to 1.2.0 (#11515)
Bumps [@agentclientprotocol/codex-acp](https://github.com/agentclientprotocol/codex-acp) from 1.1.7 to 1.2.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/codex-acp/releases">@agentclientprotocol/codex-acp's releases</a>.</em></p> <blockquote> <h2>v1.2.0</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.14...v1.2.0">1.2.0</a> (2026-08-11)</h2> <h3>Features</h3> <ul> <li>expose typed session failures for AIR (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/54987e1c4a4f878af9afad96ec8b6b0b48c7045e">54987e1</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>normalize cwd filters for Windows sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/145ebba5d2030b4aa6d19cbb89d190b7b498d454">145ebba</a>)</li> </ul> <h2>v1.1.14</h2> <h2>What's Changed</h2> <ul> <li>Update codex to 0.147.0 by <a href="https://github.com/acp-release-bot"><code>@acp-release-bot</code></a>[bot] in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/375">agentclientprotocol/codex-acp#375</a></li> <li>feat: support replacing goals through ACP control by <a href="https://github.com/nikita-ashihmin"><code>@nikita-ashihmin</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/376">agentclientprotocol/codex-acp#376</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.13...v1.1.14">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.13...v1.1.14</a></p> <h2>v1.1.13</h2> <p><strong>Full Changelog</strong>: <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.12...v1.1.13">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.12...v1.1.13</a></p> <h2>v1.1.12</h2> <p><strong>Full Changelog</strong>: <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.11...v1.1.12">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.11...v1.1.12</a></p> <h2>v1.1.11</h2> <h2>What's Changed</h2> <ul> <li>build(deps-dev): bump the npm_and_yarn group across 1 directory with 3 updates by <a href="https://github.com/dependabot"><code>@dependabot</code></a>[bot] in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/372">agentclientprotocol/codex-acp#372</a></li> <li>Fix resuming paused goals by <a href="https://github.com/nikita-ashihmin"><code>@nikita-ashihmin</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/374">agentclientprotocol/codex-acp#374</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.10...v1.1.11">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.10...v1.1.11</a></p> <h2>v1.1.10</h2> <h2>What's Changed</h2> <ul> <li>fix: Stop emitting "Conversation interrupted" message by <a href="https://github.com/Rizzen"><code>@Rizzen</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/358">agentclientprotocol/codex-acp#358</a></li> <li>Update codex to 0.146.0 by <a href="https://github.com/acp-release-bot"><code>@acp-release-bot</code></a>[bot] in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/341">agentclientprotocol/codex-acp#341</a></li> <li>build(deps): bump the npm_and_yarn group across 1 directory with 2 updates by <a href="https://github.com/dependabot"><code>@dependabot</code></a>[bot] in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/362">agentclientprotocol/codex-acp#362</a></li> <li>feat: support device code authentication via URL elicitation by <a href="https://github.com/AlexandrSuhinin"><code>@AlexandrSuhinin</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/347">agentclientprotocol/codex-acp#347</a></li> <li>Update codex to 0.146.1 by <a href="https://github.com/acp-release-bot"><code>@acp-release-bot</code></a>[bot] in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/370">agentclientprotocol/codex-acp#370</a></li> <li>feat: expose provider-neutral ACP goal extension by <a href="https://github.com/nikita-ashihmin"><code>@nikita-ashihmin</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/371">agentclientprotocol/codex-acp#371</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.9...v1.1.10">https://github.com/agentclientprotocol/codex-acp/compare/v1.1.9...v1.1.10</a></p> <h2>v1.1.9</h2> <h2>What's Changed</h2> <ul> <li>Throttle ACP plan update snapshots by <a href="https://github.com/nikita-ashihmin"><code>@nikita-ashihmin</code></a> in <a href="https://redirect.github.com/agentclientprotocol/codex-acp/pull/354">agentclientprotocol/codex-acp#354</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/codex-acp/blob/main/CHANGELOG.md">@agentclientprotocol/codex-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.14...v1.2.0">1.2.0</a> (2026-08-11)</h2> <h3>Features</h3> <ul> <li>expose typed session failures for AIR (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/54987e1c4a4f878af9afad96ec8b6b0b48c7045e">54987e1</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li>normalize cwd filters for Windows sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/145ebba5d2030b4aa6d19cbb89d190b7b498d454">145ebba</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/b51bedf60050c60fef78fc669e6ccf2ff61e3f47"><code>b51bedf</code></a> chore(main): release 1.2.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/389">#389</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/2dccf45b53a33089d5b6e82508a1887aec8b20cc"><code>2dccf45</code></a> ci: release-please release flow (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/388">#388</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/54987e1c4a4f878af9afad96ec8b6b0b48c7045e"><code>54987e1</code></a> feat: expose typed session failures for AIR (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/383">#383</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/9edc92458504a9653f539f2a515f59e4a95796a7"><code>9edc924</code></a> build(deps-dev): bump hono in the npm_and_yarn group across 1 directory (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/380">#380</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/145ebba5d2030b4aa6d19cbb89d190b7b498d454"><code>145ebba</code></a> fix: normalize cwd filters for Windows sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/377">#377</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/5faefec5d55ded33c54b68ffec93def4f6c547f5"><code>5faefec</code></a> Release v1.1.14</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/0d45a13c2618e8175a24d0bb080a482c7920f291"><code>0d45a13</code></a> feat: support replacing goals through ACP control (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/376">#376</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/a8cedc8d3789371f992979a4f4facac29f0bd681"><code>a8cedc8</code></a> Update codex to 0.147.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/375">#375</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/ea57892f7d8e305fc1bd489de420adb19de055b0"><code>ea57892</code></a> Release v1.1.13</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/91cbfd30467d57eae06d3a57ccb068320bb7e171"><code>91cbfd3</code></a> Release v1.1.12</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.1.7...v1.2.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
aad97d93fe |
fix(hermes): surface real reasoning text from reasoning.available events (#9237)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When an agent runs through the Hermes gateway adapter, its stdout is parsed line-by-line into transcript entries that the issue chat renders (the UI fetches the adapter's `./ui-parser` from `/api/adapters/:type/ui-parser.js` and runs `parseStdoutLine` client-side) > - Reasoning-capable models emit a `reasoning.available` gateway event carrying the model's reasoning text, and the chat renders `thinking` parts as expandable chain-of-thought > - The gateway parser mapped `reasoning.available` to a hardcoded `"Hermes reasoning available"` string and discarded the event payload, so the "thinking" part had no real content — the indicator looked static and expanding it revealed nothing (#9209) > - This pull request extracts the actual reasoning text from the event payload and uses it as the `thinking` part's text, keeping the old string only as a fallback for payloads that carry no text > - The benefit is that the "Hermes reasoning available" indicator now surfaces the model's real reasoning, which the existing expandable-thinking UI can display ## Linked Issues or Issue Description Fixes: #9209 ## What Changed - `packages/adapters/hermes/src/gateway/ui/parse-stdout.ts`: the `reasoning.available` handler now extracts the reasoning text from the event `data` via a small helper (`extractReasoningText`), checking the plausible field names (`reasoning`, `reasoning_text`, `thinking`, `text`, `summary`, `content`) and recursing one level into nested `data` / `payload` records, with ANSI stripped. The prior `"Hermes reasoning available"` string is kept only as a fallback when no text field is present. - `packages/adapters/hermes/gateway-ui-parser.cjs`: applied the identical logical change to the committed CommonJS mirror (exported as `./gateway/ui-parser`), keeping the two files in sync. - `packages/adapters/hermes/src/gateway/ui/parse-stdout.test.ts` (new): unit tests for the gateway parser (there were none) covering direct-field, `summary`, nested `data`/`payload` extraction, the no-text fallback, and regression guards for `message.delta` and plain stdout. ## Verification Ran from `packages/adapters/hermes`: - `node_modules/.bin/vitest run src/gateway/ui/parse-stdout.test.ts` → **8/8 passed**. - Negative control: stashed the source changes and re-ran the same test file against the current (pre-patch) parser → **4/8 failed** (exactly the reasoning-extraction assertions), then restored — confirming the tests are discriminating, not vacuous. - `npx tsc --noEmit -p .` → clean. Real-behavior proof (driving the actual shipped `gateway-ui-parser.cjs` `parseStdoutLine`) is in the block below. ## Risks - **Low risk.** Behavior is unchanged for events that carry no recognizable text field — the `"Hermes reasoning available"` fallback is preserved (verified). Only the `reasoning.available` branch changed; `message.delta`, `run.failed`/`run.error`, and the generic/system/stdout branches are untouched. - The exact field name in a real `reasoning.available` payload is defined by the external Hermes gateway and is not present anywhere in this repo, so the extraction is intentionally defensive across several plausible field names rather than pinned to one. If the real event nests the text differently than `data` / `payload`, it will fall back to the existing placeholder (i.e. no regression vs. today). Happy to tighten the field list against real gateway traffic if a maintainer can share a sample. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`) via Claude Code, with tool use and local test execution (ran vitest/tsc against the change). Planning, code review, and the real-behavior proof were done with Claude (Opus 4.8) in the same session. ## Real behavior proof **Behavior addressed:** A `reasoning.available` Hermes gateway event now produces a `thinking` transcript part containing the model's real reasoning text, instead of a static `"Hermes reasoning available"` placeholder with no content behind it (#9209). **Real environment tested:** Drove the actual shipped production artifact — `packages/adapters/hermes/gateway-ui-parser.cjs`, the exact module the UI loads via `/api/adapters/hermes-gateway/ui-parser.js` and runs to parse gateway stdout — on Node v24.16.0, macOS. The input is a raw stdout line in the exact format emitted by `packages/adapters/hermes/src/gateway/server/execute.ts` (`[hermes-gateway:event] run=… event=reasoning.available data=…`). Only the external gateway boundary (the raw line) is synthesized; the parser code path is the real one. **Exact steps or command run after this patch:** ``` # BEFORE = git show HEAD:…/gateway-ui-parser.cjs ; AFTER = patched artifact node proof.cjs # requires each parser build and calls parseStdoutLine(line, ts) # line = [hermes-gateway:event] run=run-abc123 event=reasoning.available \ # data={"text":"Checking whether the cache key includes the tenant id before I refactor the lookup."} ``` **Evidence after fix:** ``` ===== BEFORE (master / old code) ===== [ { "kind": "thinking", "ts": "…", "text": "Hermes reasoning available" } ] thinking part carries real reasoning text? -> NO (static placeholder, nothing for the UI to expand) ===== AFTER (this patch) ===== [ { "kind": "thinking", "ts": "…", "text": "Checking whether the cache key includes the tenant id before I refactor the lookup." } ] thinking part carries real reasoning text? -> YES ``` Additional cases through the same shipped artifact after the patch: ``` -- nested payload (data.payload.reasoning) -- {"kind":"thinking","ts":"…","text":"Weighing two migration orders."} -- bare signal, no text field (regression guard) -- {"kind":"thinking","ts":"…","text":"Hermes reasoning available"} # fallback preserved -- message.delta still works (regression guard) -- {"kind":"assistant","ts":"…","text":"Hello","delta":true} ``` **Observed result after fix:** The `reasoning.available` event yields a `thinking` part carrying the model's real reasoning text (top-level or nested), which the existing expandable-thinking rendering in the chat can display. Events with no text field still yield the original placeholder, and unrelated events are unaffected. **What was not tested:** I did not run against a live Hermes gateway — Paperclip's Hermes gateway binary and its credentials aren't available on this machine, and no captured real `reasoning.available` payload exists in the repo, so the exact wire field name is inferred (hence the defensive multi-field extraction + safe fallback). I also did not render the full React chat component in jsdom; the change is confined to the parser, and the chat's expandable `thinking` rendering already exists (`ui/src/components/IssueChatThread.tsx`). CI / unit tests here are supplemental to the runtime proof above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (searched `9209 in:body` and keyword variants — none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/hermes-reasoning-available-payload`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no user-facing docs describe this behavior; none needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs on the PR) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (will address on review) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
120ae5428f |
feat(server): add a one-click relink action for detached custom-image templates (#11641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip environments can use captured custom images for agent runs > - A configuration fingerprint change can detach a valid custom-image template > - Operators need a safe way to confirm that the image still matches the boot source > - This pull request adds a guarded relink action with drift classification and audit logging > - The benefit is a deliberate relink without a new sandbox boot or provider snapshot ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting environment, server, and UI behavior. **Problem or motivation** A custom-image template detaches when the environment configuration fingerprint changes. The runtime then uses the base image, even when the boot source did not change. The only prior remedy required a full re-capture. **Proposed solution** Add an operator-triggered relink action. Classify configuration drift from a server-owned boot-relevant snapshot. Relink knob-only drift without confirmation. Require explicit confirmation for boot-source or unclassified drift. Guard the route for instance administrators and record a safe activity event. **Alternatives considered** Keep requiring a full re-capture. This adds a sandbox boot and provider snapshot for cases where the image remains correct. **Roadmap alignment** The roadmap has no matching custom-image relink item. This change addresses an environment operation gap. **Additional context** The relink response exposes raw drift values only in the transient 409 response to the instance administrator. The service never persists or logs fingerprints or configuration values. Reserved identity-path segments fail closed. ## What Changed - Add `relinkActiveTemplate` with drift classification and conditional fingerprint update. - Persist a server-owned boot-relevant configuration snapshot during capture. - Add the guarded relink route with strict request validation and activity logging. - Add the relink action and confirmation flow to the environment page. - Add service, route, UI, and OpenAPI coverage. ## Verification - Run the focused service suite: `pnpm vitest run server/src/services/environment-custom-images-service.test.ts`. - Run the focused route suite: `pnpm vitest run server/src/routes/environment-custom-image-routes.test.ts`. - Run the focused UI suite: `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx`. - Run server and UI TypeScript checks. - Confirm the OpenAPI snapshot matches the new route. - Confirm all required GitHub checks pass on commit `e46fdcfe94a719be854adf8849d30714e5b70b93`. - Confirm Greptile reports 5/5 with no unresolved review threads. ## Risks The relink action can keep an image after configuration drift. The service requires explicit confirmation for boot-source or unclassified drift. Reserved path segments produce a safe unresolved marker and never enter stored values. ## Model Used OpenAI GPT-5 Codex. The model used repository inspection, GitHub operations, and PR preparation with tool use and code execution. The runtime did not expose a context-window value or a separate reasoning-mode value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
393da0f67c |
fix(adapter-utils): graft unrelated imported histories instead of failing the run (#11638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - At run finalize, the host imports the sandbox git history and reconciles it with the local worktree in `integrateImportedGitHead` > - Transported workspaces are depth-1 shallow clones, so the boundary commit reads as parentless inside the sandbox > - A `git commit --amend` there rewrites the boundary commit into a root commit, and the re-imported history no longer connects to the host history > - `git merge-tree` has no common base to merge against, so the sync throws "Failed to merge concurrent remote git histories" and the run fails with its work stranded in the sandbox > - This pull request grafts the imported tree onto the current head as a single commit instead of failing > - The benefit is that a history rewrite inside the sandbox can no longer lose a run's work ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "unrelated histories", "Failed to merge concurrent", and "shallow". Depends on #11637 (merged; the graft commit reuses its identity constant). This PR is now rebased onto `master`. **What happened?** An agent run amended a commit inside its sandbox workspace to address review feedback. The sandbox clone is depth-1 shallow, so git treated the boundary commit as parentless and the amend produced a root commit. At finalize, the host-side sync failed with `Failed to merge concurrent remote git histories for <sha>` and the run was marked failed. A follow-up run had to repair the branch by hand: fetch the true parent from origin and rebuild the commit with `git commit-tree`. **Expected behavior** The sync must never strand completed work. When the imported history shares no ancestor with the local one, the imported tree should still land on the current head, with the imported message preserved and the graft recorded. **Steps to reproduce** 1. Start a run whose workspace transport uses the shallow clone path (`withShallowGitWorkspaceClone`, depth 1). 2. Inside the sandbox workspace, run `git commit --amend` on the boundary commit. The result is a parentless root commit. 3. Finish the run. The host-side `integrateImportedGitHead` finds no merge base, `merge-tree` fails, and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `createUnrelatedHistoryGraftCommit` helper. It reads the imported head's tree and message, and creates one commit on top of the current head with the deterministic sync identity and a trailer that records the graft and both shas. - `integrateImportedGitHead` (both the remote-git-sync version and the SSH copy in `ssh.ts`): when `merge-base` reports no common ancestor, graft instead of throwing. The ref update keeps the same compare-and-swap and concurrent-retry semantics as the merge path. - The graft is gated on `git merge-base` exiting with status 1 — the no-ancestor signal. Operational failures (timeout, missing object, repository error) keep the loud merge failure instead of rewriting the tip. - New regression tests: one builds the exact shallow-amend shape (a root commit rebuilt from the base tree) and asserts the graft lands on the current head with the imported tree, subject, and graft trailer; one integrates a well-formed sha the repository does not hold and asserts the integration still throws with the branch tip unchanged. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 19/19 pass (includes the new graft test and the merge-base failure-discrimination test). - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks - Behavioral shift: unrelated imported histories previously failed the integration; now they land as a squash-graft. In this degenerate case there is no base to merge against, so the imported tree is taken wholesale and concurrent local-only tree changes are superseded at the tip. The local commits keep their place in the graft's ancestry, and the trailer records both shas, so nothing is unrecoverable. The old behavior lost the imported work instead, which is the worse failure for an autonomous run. - The graft reuses the imported head's commit message, so branch history still reads naturally after a sandbox rewrite. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9ea8143c87 |
fix(adapter-utils): give sync-created merge commits a deterministic git identity (#11637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs execute in transported workspaces; at finalize, the host syncs the sandbox git history back into the local worktree > - When both sides advanced, `integrateImportedGitHead` reconciles them with `git merge-tree` plus `git commit-tree` on the host > - Execution hosts are often containers with no git config and no resolvable hostname, so `commit-tree` fails with "Author identity unknown" > - That one local command failure marks the whole run as failed, even though the run's work succeeded > - This pull request gives sync-created merge commits an explicit, deterministic identity at the call site > - The benefit is that workspace finalize no longer depends on ambient host git configuration ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "Author identity unknown", "unrelated histories", and "commit-tree identity". **What happened?** A run finished its work, but workspace finalize failed. The host-side sync ran `git commit-tree <tree> -p <localHead> -p <importedHead> -m "Paperclip remote git sync merge <sha>"`. Git exited with `Author identity unknown ... fatal: unable to auto-detect email address (got 'node@<container-id>.(none)')`. The adapter recorded the whole run as failed, and the host worktree kept the stale head. Any container deployment without a global gitconfig reproduces this; I observed it on a Paperclip Cloud stack. **Expected behavior** Commits that the sync machinery itself creates must not depend on ambient host git configuration. The merge commit is machine-authored, so it should carry a deterministic Paperclip identity. **Steps to reproduce** 1. Run the Paperclip server in a container with no `user.name`/`user.email` git config and a hostname git cannot turn into an email. 2. Let a run's sandbox branch diverge from the host worktree, so both sides advance. 3. Workspace finalize calls `integrateImportedGitHead`. The `git commit-tree` step fails with "Author identity unknown" and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `GIT_SYNC_COMMIT_IDENTITY_ARGS` (`-c user.name=Paperclip -c user.email=noreply@paperclip.ing`), applied to the `commit-tree` call in `integrateImportedGitHead`. - `ssh.ts`: the SSH-sync copy of `integrateImportedGitHead` applies the same identity args to its `commit-tree` call. - New regression test: builds divergent histories in a repo with no configured identity and asserts the sync merge commit is created with the deterministic identity, correct parents, and merged tree. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 18/18 pass. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Negative proof: with the source fix stashed, the new test fails on the identity assertion. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks Low risk. The change only adds `-c` identity flags to two machine-generated commit invocations. `GIT_AUTHOR_*` / `GIT_COMMITTER_*` environment variables still take precedence over `-c` when an operator sets them, so existing deployments that configure an identity keep their behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6b8e42168e |
Add governed secret alias confirmation cards (#11486)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need scoped secret bindings to use external services safely. > - Agents could not request an existing secret under a new config name without an internal secret identifier. > - Existing binding proposals were only visible in Settings and did not create an issue-thread approval path. > - A confirmation card could record acceptance without proving that the binding was created. > - This pull request extends the existing secret proposal system with safe source references and governed issue-thread confirmation cards. > - The benefit is a one-click flow that creates the binding or shows a clear failure without exposing secret material. ## Linked Issues or Issue Description Related prerequisite: #11482. **Subsystem affected** Cross-cutting: server REST APIs, shared interaction contracts, database proposal schema, and issue-thread UI. **Problem or motivation** An agent can need an existing bound secret under a second config name. The agent cannot safely discover the internal secret identifier. The existing proposal is also easy for the operator to miss because it only appears in Settings. A generic confirmation can record acceptance without executing the binding. **Proposed solution** Let an agent create a binding proposal from one of its existing config paths. Mint a server-owned, human-only confirmation card on the checked-out issue. Recheck the operator's target-agent permission under the proposal row lock. Execute the existing proposal transaction after card acceptance. Store an `executed` or `failed` result on the card. Render the complete lifecycle in the issue thread and attention resolver. **Alternatives considered** A new alias subsystem would duplicate proposal quotas, expiry, authorization, and binding synchronization. A text-only issue comment would not provide a governed action or an execution result. An agent-supplied card payload would permit metadata smuggling. This change uses the existing proposal transaction and a server-owned payload instead. **Roadmap alignment** This change extends the completed "Secrets Manager with per-agent access" roadmap item. It preserves scoped bindings and audited resolution. The required GitHub search found no other open duplicate issue or pull request. ## What Changed - Added safe source-config-path binding proposals and preserved user-secret ownership checks. - Added a proposal-to-interaction link and an idempotent database migration. - Minted human-only `request_confirmation` cards with server-owned `secretProposal` metadata. - Rejected agent-supplied governed metadata and agent addressees. - Rechecked `agent_config:update` authority under the proposal lock before execution. - Recorded `executed` or `failed` results and posted a failure comment when no binding was created. - Settled failed accepted proposals atomically and mirrored rejection, withdrawal, and expiry in both directions. - Emitted `secret.binding.created` for new agent binding writes. - Added a dedicated issue-thread card for pending, executed, failed, rejected, withdrawn, and expired states. - Showed only the source label, target agent, config path, skeptical justification, expiry, and safe failure code. - Replaced resolved attention-query entries immediately with the stitched server result. - Added focused server, database, UI, and state-transition tests. - Added Storybook fixtures for every review state and documented the API and agent behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/components/AttentionInteractionResolver.test.ts` — 58 passed. - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/db check:migrations` - `NODE_ENV=test pnpm exec vitest run server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/secret-proposals-routes.test.ts server/src/__tests__/secrets-routes.test.ts server/src/__tests__/agents-service-secret-bindings.test.ts` — 142 passed. - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/company-secret-proposals-migration.test.ts --silent` — 1 passed. - `pnpm -r typecheck` - `pnpm test:run` — server 4,175 passed, UI 4,109 passed; the CLI AWS-doctor case passes 8/8 with runtime-injected static AWS credential variables unset. - `pnpm build` - `git diff --check origin/master...HEAD` ## Risks - Migration `0221` adds one nullable foreign key and one index. It uses idempotent guards. - The accept route performs a governed write after it records card acceptance. A failed write is visible and settles the proposal as rejected. - Concurrent proposal and card resolution must use proposal-before-interaction lock order. A race test covers direct approval against card rejection. - The new audit event increases activity rows for newly added agent bindings. It does not include secret values or fingerprints. - The card includes only safe proposal metadata. It does not include secret value, fingerprint, version, or internal secret identifiers. - The UI uses the stitched resolution result. Focused tests cover immediate cache replacement and every terminal state. > This work extends an existing completed roadmap capability. The GitHub duplicate search returned no other open related work. ## Model Used - OpenAI Codex with model ID `gpt-5`. The runtime did not expose its context-window size. Reasoning, repository tools, code execution, database integration tests, UI rendering, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b446ff59bf |
refactor(acpx-engine): coordinator-owned ACP run lifecycle with a typed resource ledger (#11576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters run agent sessions through the ACPX engine > - The ACPX engine handled one run attempt as a long implicit procedure > - That shape made resource ownership, cleanup order, and failure behavior hard to verify > - This pull request gives the attempt a coordinator, a typed resource ledger, separate run sites, and explicit turn and settlement sequences > - The benefit is clear ownership, one cleanup path, safer session reuse, and testable failure behavior ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX engine manages startup, turn execution, session reuse, and cleanup inside one large run procedure. **Current behavior** The run procedure owns several resources through implicit control flow. Cleanup and session reuse behavior depend on lane-specific branches and error paths. **Proposed behavior** The coordinator owns the run attempt. A typed ledger records six resources and their states. Host and sandbox run sites own lane-specific acquisition. Turn and settlement sequences expose typed outcomes. The engine emits allowlisted phase telemetry. **Reason and benefit** Explicit ownership makes cleanup and failure behavior easier to inspect. The fault matrix and characterization tests protect the external result while the refactor reduces hidden control flow. **Breaking changes** None to the public adapter contract. The host warm-save path now closes and relaunches the runtime because a transferred runtime could retain a run-scoped credential. A cold session-handshake failure now closes the created runtime. **Additional context** This pull request contains the ACPX engine lifecycle refactor, its tests, and the lifecycle document. ## What Changed - Add a run coordinator for startup, turn execution, settlement, and result reproduction. - Add a typed resource ledger with open, sealed, and consumed states. - Add host and sandbox run sites for lane-specific resource acquisition. - Replace separate runtime maps with a generic session reuse store. - Split session fingerprint identity from the outer session key. - Add typed turn and settlement sequences with one cleanup owner. - Add a closed allowlist for phase telemetry. - Add characterization tests and a 17-case fault matrix. - Add `doc/acp-run-lifecycle.md`. ## Verification - `npx vitest run packages/adapter-utils/src/acpx-engine/` passes 18 files and 286 tests at the submitted commit. - `pnpm --filter @paperclipai/adapter-utils typecheck` reports 0 errors at the submitted commit. - Run the full pull request checks after GitHub starts CI. - Run Greptile review after the pull request opens. ## Risks - The refactor changes internal control flow across the ACPX engine. - Host warm-save behavior now closes and relaunches the runtime. - Settlement changes the handling of a cold session-handshake failure from a leak to a close. - The characterization baselines and fault matrix reduce the risk of an external behavior change. > Paperclip is the open source app people use to manage AI agents for work > The adapter layer runs agent sessions through the ACPX engine > The engine needs explicit lifecycle ownership for reliable cleanup > This pull request adds coordinator-owned phases and a typed resource ledger > The result makes lifecycle behavior easier to test and review ## Model Used OpenAI GPT-5 Codex. Exact model ID: GPT-5. The model used tool execution, repository inspection, and code review support. The implementation author supplied the submitted code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c366a9059 |
fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server authenticates each agent request in `actorMiddleware` before it attributes chat comments > - When an agent bearer token failed verification, the middleware called `next()` with no error and the request continued without an agent actor > - The request then fell back to the local user actor, so the server stored agent replies as user comments > - The task chat UI renders user comments in blue bubbles, so agent messages appeared as blue user bubbles > - This pull request rejects invalid agent credentials with 401 instead of a silent downgrade > - The benefit is that agent messages keep agent attribution, and broken credentials fail loudly with a clear retry message ## Linked Issues or Issue Description **What happened?** A user cancelled an onboarding question card. The agent posted a follow-up reply. The reply appeared in a blue bubble, which the UI reserves for human messages. The agent run held an expired local agent JWT. The auth middleware could not verify the token, called `next()` without an actor, and the request fell back to the local user identity. The server stored the agent comment as a user comment. **Expected behavior** Agent messages always render as agent bubbles. A request with invalid agent credentials must fail with 401 so the adapter can refresh credentials and retry. It must not post content under a human identity. **Steps to reproduce** 1. Start a local Paperclip instance. 2. Give an agent run an expired or malformed agent JWT. 3. Let the agent post an issue comment through the API bridge. 4. Before this change: the comment is stored with the local user identity and renders as a blue bubble. After this change: the request fails with 401 and a message that tells the caller to obtain fresh credentials. ## What Changed - `server/src/middleware/auth.ts`: a bearer token that fails verification now produces a 401 `unauthorized` error instead of a silent fall-through to the anonymous/local-user actor. - The 401 message states the cause: expired token, unverifiable token, empty bearer token, missing agent record, agent record in another company, terminated agent, or agent pending approval. - The API-key path now also rejects an agent record whose company does not match the key. - `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now writes a `comment id: <id>` marker to the run log for each posted issue comment, so misattributed comments can be traced to a run. - `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression test asserts that a recovered `local-board` comment with a derived agent author renders as an agent bubble, not a user bubble. - `server/src/__tests__/agent-auth-middleware.test.ts` and `packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests cover each rejection path and the log marker. ## Verification - Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in `server/` — 14 tests pass. - Run `pnpm vitest run execution-target-sandbox` at the repo root — 44 tests pass. - Run `pnpm vitest run src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests pass. - Manual check: post an issue comment with an expired agent JWT; the API returns 401 with a retry message and no comment is stored. ## Risks - Behavioral shift: requests that previously continued as anonymous or local-user actors after a failed agent-token verification now receive 401. Any caller that relied on the silent downgrade must refresh its credentials. This is the intended fix, and the adapters already handle 401 with a credential refresh. - No schema or migration changes. Low risk otherwise. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with extended thinking and tool use (agent harness with shell, file, and git tools). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |