mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
b872cd3d1b404bdaff70af493a2973ceb7e5d6ec
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dfdfc8664e |
feat(claude-local): add Claude Fable 5.1 support (#12730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude local adapter lets operators select a Claude model for an agent. > - Claude Fable 5.1 was absent from the adapter model lists. > - The adapter runtime also used a Claude Code build that rejected Fable 5.1. > - This pull request adds the direct Anthropic ID and the AWS Bedrock inference profile ID. > - It also updates the Claude ACP runtime and keeps the Paperclip usage and isolation patches. > - The benefit is that operators can select and run Claude Fable 5.1 through the Claude adapter. ## Linked Issues or Issue Description Refs #8810. That issue covers related model ID handling. This change does not change provider-prefixed model IDs. **Agent or provider** Claude Code through the built-in `claude_local` adapter. The requested model is Claude Fable 5.1. **Why this adapter is useful** Operators can use Fable 5.1 without entering an undocumented model ID. The configured model also reaches both supported Claude execution lanes. **How the agent is invoked** The CLI lane sends `--model claude-fable-5-1`. The ACP lane sends `ANTHROPIC_MODEL=claude-fable-5-1` to `@agentclientprotocol/claude-agent-acp`. **Are you willing to implement it?** Yes. This pull request includes the implementation and tests. **Additional context** Claude Code 2.1.232 rejected Fable 5.1 and required version 2.1.251 or newer. ACP package 0.73.0 includes Claude Code 2.1.257. The update keeps Paperclip's usage metadata and isolated-context behavior. ## What Changed - Added `claude-fable-5-1` to the direct Claude fallback list. - Added `us.anthropic.claude-fable-5-1` to the AWS Bedrock list. - Kept the existing default model at the first position in each list. - Updated the Claude ACP dependency from 0.70 to 0.73. - Carried the Paperclip usage and isolated-context changes into the 0.73 patch. - Added a Claude Code 2.1.251 minimum-version preflight for Fable 5.1 when using the standard `claude` executable, surfaced in both adapter Test and execution. Explicit custom wrappers retain their existing compatibility contract. - Kept local adapter Tests from executing caller-selected binaries: when runtime `PATH` selects a different Claude executable than the trusted probe, the Test warns and defers the authoritative version check to execution instead of approving or rejecting the alternate installation. - Added tests for model listing, discovery deduplication, Bedrock filtering, model pass-through in both execution lanes, old-CLI rejection before launch, custom-wrapper compatibility, and local runtime-PATH mismatch handling. ## Verification - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm exec vitest run packages/adapters/claude-local/src/server/execute.remote.test.ts packages/adapters/claude-local/src/server/test.remote.test.ts packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/acp.test.ts server/src/__tests__/adapter-models.test.ts` (72 tests passed) - `node --test scripts/acpx-patch-packaging.test.mjs` (13 tests passed) - `pnpm -r typecheck` - `pnpm build` - A local Paperclip agent run completed with `usageJson.model` set to `claude-fable-5-1` through ACP 0.73.0 and its bundled Claude Code 2.1.257. - `pnpm test:run` completed 5,638 passing tests and 24 skipped tests. It also found 24 failures in unrelated workspace-runtime, path-canonicalization, and runtime-exposure tests on macOS with Node 26. These failures do not touch this diff. Clean pull request CI is the final full-suite gate. ## Risks - The ACP dependency update can change Claude runtime behavior outside model selection. Focused ACP tests, the full typecheck, the production build, and a real local Fable run reduce this risk. - The 0.73 patch must stay aligned with the installed ACP version. Dependency-resolution CI verifies the manifest and patch pair. - Fable 5.1 adds a short `claude --version` preflight to standard CLI-lane Tests and runs. The result is intentionally not cached so an in-place Claude Code upgrade takes effect without restarting Paperclip. Explicit custom wrappers are not version-probed because their output and compatibility contract can differ from the standard executable. - Local Tests preserve the existing deny-by-default probe boundary and do not execute a binary selected by caller-controlled `PATH`. A mismatched runtime binary produces an explicit warning without blocking an otherwise valid setup; execution independently validates the actual runtime-selected CLI before launch. - The AWS Bedrock identifier differs from earlier IDs because Fable 5.1 has no `-v1` suffix. The model-list test locks this exact value. - There is no schema change or migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: OpenAI. Model: GPT-5 Codex. The host did not expose a more specific model ID or context-window size. Capabilities used: agentic reasoning, repository editing, shell execution, web research, and local runtime verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ad0ad43cf4 |
feat(runner): activate qualified Claude ACPX runtime (#12590)
## Thinking Path > - Paperclip Runner already has a hardened ACPX path for Codex. > - Claude can reuse that protocol only with an exact package/model profile and provider-lifetime fencing. > - Pi needs a separately spawned runtime whose executable does not yet have the descriptor-confined verified launch used by the ACP server. > - This pull request therefore activates Claude only and keeps Pi unavailable before installation or process launch. ## Linked Issues or Issue Description **Subsystem affected** Paperclip Runner ACPX driver, runtime host, sidecar, backend factory, package dependency, and provider conformance tests. **Problem or motivation** The production ACPX backend was Codex-only. Claude needs the same fail-closed model, authorization, cancellation, cleanup, and recovery boundaries without exposing an unsafe secondary runtime path. **Proposed solution** Generalize the hardened ACPX runtime for the exact qualified `claude` profile, add the pinned Claude ACP package and reviewed isolation patch, and reject Pi before installation, backend construction, sidecar initialization, Rust session admission, or process creation. **Alternatives considered** Activating Pi in this PR was rejected after security review because its secondary runtime executable was pathname-based and lacked the verified descriptor/snapshot boundary. Pi is deferred to a dedicated follow-up. Replaying the older generic ACPX implementation was rejected because it predates current hardening. **Roadmap alignment** ROADMAP.md does not list a conflicting ACPX-provider project. This extends the existing Runner provider architecture. ## What Changed - Generalized the ACPX backend, driver, runtime adapter, host, and sidecar for the qualified Claude profile. - Added Claude ACPX activation through its exact pinned package/model pair and isolated-settings patch. - Added provider-lifetime fencing for non-Codex qualified ACPX sessions. - Kept Pi dependencies and its patch out of the package and build configuration. - Added fail-closed Pi rejection at driver validation, backend construction, runtime-host admission, sidecar initialization, and Rust session validation. - Added focused tests for Claude selection, model enforcement, lifecycle fencing, cancellation, recovery, and Pi rejection. - Did not change or commit `pnpm-lock.yaml`; CI regenerates the PR lockfile under the existing repository policy. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - CI runs dependency policy, runner package checks, protocol parity, typecheck, build, security, and stack policy. - Local tests were not run because this checkout is resource constrained, per the requested workflow. ## Risks - Claude package behavior can drift from the qualified protocol; the package and patch are pinned and admission verifies the exact profile. - Unsupported providers and models fail closed. - Pi remains unavailable until descriptor-confined verified launch exists for its separate runtime. - Existing Codex ACPX behavior remains covered by shared conformance tests. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs - [x] I have described the issue in the PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name contains no internal task identifier - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [ ] All applicable Paperclip CI gates are green - [ ] Greptile is 5/5 with no actionable findings ## Stack - Position: lowest unmerged PR - Base: `master` - Previous: [#12588](https://github.com/paperclipai/paperclip/pull/12588), merged qualified OpenCode runtime - Next: [#12591](https://github.com/paperclipai/paperclip/pull/12591), native application integration |
||
|
|
9ca24bba3c |
feat(runner): pin the Codex ACPX runtime (#12400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local host boundary is ready for a concrete ACP implementation, but the first production profile is Codex only. > - ACPX must not inherit the server process environment or choose an executable by pathname after admission. > - Codex must not re-enable ambient apps, memory, skills, MCP configuration, or instructions inside its isolated home. > - This pull request pins only the two required production packages and applies narrowly tested host patches. > - The benefit is a minimal dependency boundary that follows the repository's CI-owned lockfile process. ## Linked Issues or Issue Description **Agent or provider** Codex through `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. **Why this adapter is useful** The injected runtime host needs a concrete ACP session manager and the exact reviewed Codex ACP server. Upstream ACPX does not yet expose a host-owned spawn callback, and upstream Codex ACP does not yet apply Paperclip's isolated instruction, MCP, app, memory, and skill boundary. Both behaviors are required before the dependency can execute inside the runner. **How the agent is invoked** The next pull request will adapt these pinned packages to the private runtime host. ACPX receives a host-owned callback that consumes the already verified executable lease. Codex receives only the isolated environment, explicit base instructions, explicit MCP servers, and the skills rooted in its private `CODEX_HOME`. This pull request alone does not spawn either package or register an adapter. **Additional context** This pull request is stacked on #12399. It adds no Pi, Claude, AWS, SDK, lab, browser, or UI dependency. It intentionally does not commit `pnpm-lock.yaml`: the repository policy job regenerates a manifest-only PR lockfile artifact for downstream frozen installs, and the lockfile bot updates master separately. ## What Changed - Pin `acpx` to `0.13.1` and the Codex ACP server to `1.6.2` in the runner package. - Register both patches in the pnpm 9 root configuration and newer-pnpm workspace configuration. - Preserve the existing embedded-Postgres and ACPX 0.12 patch entries used by other packages. - Patch ACPX to evaluate an allowlisted environment at child-spawn time and keep spawn cwd out of provider-visible session identity. - Patch ACPX to accept a host-owned spawn callback with the resolved arguments and options, allowing the verified command lease to own execution. - Patch Codex ACP to retain runner-owned MCP server identity in permission requests. - Patch Codex ACP to pass explicit Paperclip base instructions on both start and resume. - In isolated mode, disable ambient apps, memory, and existing MCP configuration; load skills only from `CODEX_HOME`; and configure only requested servers. - Add a package contract test that enforces exact versions, Codex-only dependency scope, both pnpm patch registries, and every required patch hook. ## Verification - Both patch files dry-apply successfully to fresh published tarballs for `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. - A local no-lockfile install applied both patches; their runtime markers and exact installed versions were inspected. - Runner TypeScript typecheck — passed against the patched packages. - Runner package tests — passed: 16 Node protocol/package tests and 426 Vitest tests. - `pnpm -r typecheck` — passed for all applicable workspaces. - `pnpm build` — passed, including runner binary, server, UI, and workspace packages. - `git diff --check` — passed. - The diff contains 6 files and does not change `pnpm-lock.yaml`, a GitHub workflow, server selection, or UI behavior. ## Risks The primary risk is drift between published package contents and checked-in compiled patches. Exact versions are pinned, both patches are exercised by package-contract gates, and CI performs the authoritative regenerated-lockfile frozen install. The spawn callback does not grant a new executable path: the following adapter must consume the opaque verified command lease. Codex isolation changes activate only when `PAPERCLIP_ACPX_ISOLATED_CONTEXT=1`, so existing direct Codex adapters are unaffected. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing public item or described the issue in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dependency, patch, isolation, and lockfile boundaries - [ ] All applicable GitHub Actions are green - [ ] Greptile is 5/5 with every actionable comment resolved - [x] I will address all review findings before requesting merge |
||
|
|
0834a0c1f7 |
feat(runner): bind ACPX profile boundary (#12387)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip runner needs a safe boundary before it can launch ACP-compatible agents. > - A caller-controlled command, model, environment, or frame could bypass that boundary. > - The ACPX transport contract in #12386 defines the allowed messages but does not bind a runtime profile. > - This pull request defines closed, versioned profiles and validates the launch inputs around that contract. > - The benefit is a small and reviewable trust boundary before any ACPX process can become available. ## Linked Issues or Issue Description **Agent or provider** ACPX sidecar support for the qualified Pi, Claude, and Codex ACP servers. **Why this adapter is useful** The runner needs one bounded process boundary for ACP-compatible providers. A closed profile prevents an untrusted run from selecting an arbitrary executable, package version, or model. **How the agent is invoked** A later pull request will launch an internal sidecar from an exact profile. This pull request only validates profiles, environment values, and protocol frames. It does not add an executable dependency or enable an adapter. **Additional context** This pull request is stacked on #12386. It keeps the existing direct adapters and the Codex runner path unchanged. ## What Changed - Add a closed profile table for the qualified Pi, Claude, and Codex ACP servers. - Require the exact qualified model and return an isolated profile value to callers. - Add an agent-specific environment allowlist with entry and aggregate size limits. - Add strict parsing for bounded sidecar requests and structured plan values. - Reject unknown fields, unsupported protocol versions, invalid identifiers, null bytes, cyclic values, and oversized input. ## Verification - Runner TypeScript typecheck — passed. - Runner TypeScript tests — 40 files and 362 Vitest tests passed; 11 Node contract tests passed. - `pnpm -r typecheck` — passed for all applicable workspaces. - `pnpm build` — passed, including runner binary, server, UI, and workspace packages. - Prettier and `git diff --check` — passed. - The diff contains 6 files and does not change `pnpm-lock.yaml`, a workflow, a package dependency, or a public export. ## Risks The main risk is accepting more launch state than the sidecar needs. The implementation uses an agent-specific allowlist, rejects null bytes, and enforces per-entry and aggregate bounds. This pull request does not launch a process or expose a new adapter, so production and direct-adapter behavior remain unchanged. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing public item or described the issue in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the compatibility and security boundary - [ ] All applicable GitHub Actions are green - [ ] Greptile is 5/5 with every actionable comment resolved - [x] I will address all review findings before requesting merge |
||
|
|
8ef39febd7 |
refactor(db): drop the vendored postgres teardown patch (reverts #12227) (#12234)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - #12227 vendored a pnpm patch of the `postgres` driver to stop a teardown race (`nextWrite` firing after the socket is nulled) from crashing the process and failing green CI shards > - Maintainer call: carrying a vendored driver patch is not worth it for a CI flake — the patch adds a maintenance obligation on every future driver upgrade > - The race is an upstream bug in `postgres@3.4.9`; the plan is to wait for an upstream release that fixes it and bump the dependency instead > - This pull request reverts #12227 in full: the patch file, its `package.json` registration, and the regression test that exercised the patched behavior > - The benefit is an unmodified dependency graph; the known flake signature returns and is retried when it bites ## Linked Issues or Issue Description Reverts #12227. **What existing behavior does this improve?** Dependency hygiene: `postgres@3.4.9` is consumed unmodified again, with no `pnpm.patchedDependencies` entry to re-evaluate on every driver upgrade. **Current behavior** The repo carries `patches/postgres@3.4.9.patch` (null-socket guard in the driver's deferred write flush, plus an `execute()` refusal on socketless connections) and a regression test for it. **Proposed behavior** Plain upstream `postgres@3.4.9`. The teardown race stays an upstream bug: a green test shard can occasionally fail with `Vitest caught 1 unhandled error` and `TypeError: Cannot read properties of null (reading 'write')` at `Immediate.nextWrite`; the remedy is retrying the shard until an upstream driver release fixes the race and we bump. **Reason and benefit** A vendored driver patch is a standing maintenance cost that outweighs the flake it suppressed. ## What Changed - Reverts #12227 (`6c7c0fd1f`) in full: removes `patches/postgres@3.4.9.patch`, its `pnpm.patchedDependencies` registration in root `package.json`, and `packages/db/src/postgres-driver-teardown.test.ts`. No lockfile involvement — the merged commit never touched `pnpm-lock.yaml` and the refresh bot had not yet recorded the patch. ## Verification - `pnpm install` on the reverted tree is coherent; the full `packages/db` suite passes (26 files / 100 tests). - `git revert` applied cleanly with no conflicts. ## Risks - Low. This restores the exact pre-#12227 state. The known flake signature returns; it fails jobs whose tests all passed and is cleared by retrying the shard. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding session with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6c7c0fd1f2 |
fix(db): stop the postgres driver from crashing the process on a write/close race (#12227)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server and its test suites talk to PostgreSQL through the
`postgres` (postgres.js) driver, and tests routinely tear their
databases down while connections still carry traffic
> - The driver flushes small buffered frames from a `setImmediate`, and
that deferred flush calls `socket.write()` without checking that the
socket still exists; a reserved connection whose backend died keeps
accepting queries, so the flush can fire with a null socket
> - The resulting `TypeError` escapes from a timer callback with no
try/catch above it, crashing the process — in CI this fails suites whose
tests all passed ("Vitest caught 1 unhandled error"), and the same crash
is reported against the driver in the wild after ECONNRESET
> - The latest driver release (3.4.9) still has the bug, so this pull
request adds a pnpm patch guarding the flush and normalizing timer state
on close, plus a deterministic regression test
> - The benefit is CI that no longer fails randomly on a teardown race,
and production processes that survive a database connection dying at the
wrong moment
## Linked Issues or Issue Description
No public issue exists; the underlying problem follows the bug-report
template.
**What happened?**
CI jobs fail with all tests passing: vitest reports `Vitest caught 1
unhandled error during the test run` with `TypeError: Cannot read
properties of null (reading 'write')` at `postgres/src/connection.js`
`Immediate.nextWrite`. The attribution points at whichever test file
happened to be running (e.g. `native-codex-runner.integration.test.ts`),
because the throw comes from a process-level timer callback, not from a
test. The identical crash is reported against the upstream driver by
other projects after `ECONNRESET` (e.g. immich-app/immich#25098).
**Expected behavior**
A connection dying between a write being scheduled and its deferred
flush must settle the affected queries through the driver's normal
connection-error path, never throw from a bare timer callback.
**Steps to reproduce**
Run the new `packages/db/src/postgres-driver-teardown.test.ts` with the
patch removed: reserve a connection (`sql.reserve()` — the same surface
`sql.begin()` uses), destroy the backend socket, wait for the client to
process the close, then issue one query on the reserved connection. The
deferred flush fires one tick later with `socket === null` and crashes
the process with exactly the CI signature.
**Paperclip version or commit**
master `198fc8b28`, `postgres@3.4.9` (latest release; bug still present
on the driver's master branch).
## What Changed
- `patches/postgres@3.4.9.patch` (new, wired via
`pnpm.patchedDependencies`): `nextWrite` returns without writing when
`socket === null`, dropping the buffered bytes — the close path has
already settled every in-flight query, so those bytes have nowhere to
go. The `closed()` and `terminate()` handlers additionally reset
`nextWriteTimer`/`chunk` after `clearImmediate`, so a stale cleared
handle cannot silently block a future reconnect's first flush. All three
shipped builds (`src`, `cjs`, `cf`) get the identical change.
- `packages/db/src/postgres-driver-teardown.test.ts` (new):
deterministic reproduction against a minimal in-process fake wire server
(startup auth + an empty result for the `fetch_types` bootstrap).
Asserts the late query settles with `CONNECTION_DESTROYED` through
`sql.end()` instead of crashing the process.
## Verification
- The regression test fails against unpatched `postgres@3.4.9` with the
exact CI signature (verified by running the same scenario against an
unpatched checkout) and passes with the patch.
- Full `packages/db` suite: 27 files / 101 tests pass.
- Spot-checked server suites that exercise the database through the
patched driver.
## Risks
- Low. The behavioral change activates only in a state that previously
crashed the process (write flush with no socket). Dropping the buffered
bytes matches what the connection's close path already promised callers:
every in-flight query has been settled with a connection error.
- The timer/chunk reset in `closed()`/`terminate()` prevents a
theoretical stale-handle hang after reconnect; on the normal path both
were already reset by `nextWrite`.
- The patch pins to `postgres@3.4.9`; a future driver upgrade will
surface the patch for re-evaluation (pnpm fails loudly on version
mismatch), and the guard can be dropped if the fix lands upstream.
## Model Used
Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic
coding session with tool use (driver source analysis, wire-protocol fake
server, local test execution).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
1ee1275f11 |
fix(adapters): persist ACPX process identity for hot restart (#9838)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents for work. > - Local agent heartbeats need durable process identity so the server can supervise them. > - The ACPX runtime owns the child process used by `codex_local` sessions. > - ACPX did not expose the child PID and start time to the Paperclip adapter. > - Warm ACPX runtimes can also serve a later heartbeat without a new spawn event. > - A hot restart could therefore classify a live Codex run as lost because its heartbeat row had no process identity. > - This pull request forwards ACPX spawn identity, reuses it for compatible warm heartbeats, and fails closed when identity cannot be persisted. > - The benefit is reliable hot-restart adoption for eligible local Codex runs. ## Linked Issues or Issue Description No matching public GitHub issue was found. **What happened?** A `codex_local` heartbeat could run through ACPX without a persisted `processPid` or `processStartedAt`. A Paperclip hot restart then had no durable identity for the live ACP child. Recovery could classify the run as `process_lost` even while the child was still alive. **Expected behavior** ACPX reports the real child PID and start time before the first prompt. A compatible warm runtime reports the same known identity to each later heartbeat that reuses the child. ACPX stops the child if the identity is invalid or persistence fails. Hot-restart recovery can then adopt the live run. **Steps to reproduce** 1. Start a `codex_local` heartbeat through the ACPX execution lane. 2. Keep the run active during a Paperclip hot restart. 3. Inspect the heartbeat row before this change. 4. Observe that the process identity can be null and recovery cannot adopt the live child. **Reproduced on** - Paperclip `master` before this change. - Linux source deployment. - `codex_local` with ACPX `0.12.0`. ## What Changed - Add an awaited `onAgentSpawn` lifecycle hook to the patched ACPX runtime. - Forward the ACP child PID and start time through the adapter `onSpawn` callback. - Keep a mutable callback sink for cached runtimes so a later respawn updates the current heartbeat. - Reuse the last known process identity when a compatible warm heartbeat reuses the existing child. - Kill the ACP child and fail session startup when the PID is invalid or identity persistence rejects. - Add ACPX and heartbeat recovery tests for callback ordering, warm reuse, failure cleanup, durable row identity, and hot-restart adoption. - Document the one-time drain required when an installed pre-fix run already lacks process metadata. ## Verification - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-execute-escalated" pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts` — 89 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-recovery-escalated" pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts` — 92 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-remote-smoke-escalated" pnpm exec vitest run packages/adapter-utils/src/acpx-engine/remote-spawn-smoke.test.ts` — 3 passed. - `PAPERCLIP_HOME="$PAPERCLIP_RUN_SCRATCH_DIR/test-home-ci-repro-escalated" pnpm exec vitest run server/src/__tests__/heartbeat-dependency-scheduling.test.ts` — 6 passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - Reverse and forward dry-run application of `patches/acpx@0.12.0.patch` — passed. - `git diff --check` — passed. - `git diff --exit-code origin/master...HEAD -- pnpm-lock.yaml` — passed. - `git diff --exit-code origin/master...HEAD -- .github/workflows` — passed. ## Risks - Runtime risk is low to moderate. ACPX now awaits process-identity persistence during child startup. - ACPX kills the child when persistence fails. This prevents an unsupervised process, but it makes that heartbeat fail visibly. - A compatible warm heartbeat reuses the identity of the existing ACP child. Regression tests verify that identity is persisted before the next prompt. - The change updates the vendored ACPX patch. Package installation must apply that patch. - There are no schema, migration, public API, UI, workflow, or lockfile changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex used GPT-5.3-Codex for the earlier implementation. - OpenAI Codex used GPT-5 for the lifecycle-hook revision and the current fail-closed review fix. The runtime did not expose a more specific snapshot ID or context-window size. Both runs used reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b517b887ad | fix(acpx): decouple host proxy spawn cwd from in-sandbox remoteCwd (#10122) | ||
|
|
d31a28828b |
fix(acpx): support Windows agent spawning (#9980)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local Claude, Codex, Gemini, and custom ACP adapters run through the shared embedded ACPX engine > - That engine wrapped every local agent command in a generated Bash script to inject environment variables and filter child stderr > - Windows cannot directly spawn that Bash wrapper, and npm/pnpm ACP binaries are exposed through `.cmd` shims there > - ACPX 0.12 already supports per-session child environment variables, so the wrapper is unnecessary > - This pull request registers agent commands directly, injects env through ACPX session options, captures child stderr in-process, and adds a real Node ACP spawn smoke on Ubuntu and Windows > - The benefit is one cross-platform spawn path with a reusable smoke test instead of parallel shell-wrapper implementations ## Linked Issues or Issue Description Fixes #9941. Refs #9428 and #9771. **What happened** ACPX-backed local agents failed to start on Windows because Paperclip registered a generated POSIX `.sh` wrapper as the agent command. Windows also needs the `.cmd` npm/pnpm shim when resolving built-in ACP binaries, and symlink creation can fail with `EPERM` for seeded auth/skill files. **Expected behavior** The same ACPX engine path should spawn a real ACP agent on Windows and Linux, forward Paperclip/runtime env without mutating `process.env`, preserve filtered/unfiltered child stderr behavior, and fall back to copies where Windows symlinks are unavailable. **Steps to reproduce** Run a local ACPX adapter on Windows with the prior wrapper path. ACPX attempts to spawn the generated `.sh` file and the agent never initializes. **Deployment mode** Local Paperclip adapters using `packages/adapter-utils/src/acpx-engine/`. ## What Changed - Removed generated Bash agent/env wrappers and registered local commands directly with ACPX. - Passed the resolved child environment through ACPX `sessionOptions.env`, including resume retry paths. - Added a minimal `acpx@0.12.0` package patch exposing child stderr callbacks and allowing documented uppercase env-map keys in persisted session options. - Moved stderr tee/filter behavior in-process: raw stderr remains in the per-run file while benign `nes/close` noise is omitted from live stderr. - Preferred `.cmd` ancestor binaries on Windows and added `EPERM` copy fallbacks for Codex auth seeding and Gemini skill materialization. - Added a real Node ACP echo-agent spawn smoke that can run directly on any supported platform. ## Verification - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts packages/adapter-utils/src/acpx-engine/spawn-smoke.test.ts` — 57 passed. - `pnpm --filter @paperclipai/adapter-utils typecheck` — passed. - `node --test scripts/acpx-patch-packaging.test.mjs scripts/release-lib.test.mjs` — 10 passed. - Full canary release dry run under Node 24.18.0 / npm 11.16.0 — passed in an isolated scratch clone. - `git diff --check` — passed during implementation verification. - One-time GitHub Actions proof: [Ubuntu ACPX spawn smoke](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774579), [Windows ACPX spawn smoke](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774558), and [Canary Dry Run](https://github.com/paperclipai/paperclip/actions/runs/29924348927/job/88937774497) passed on head `f345ac69f2`; the dedicated smoke jobs are intentionally not retained in the recurring PR workflow. ## Risks - The ACPX stderr callback and env persistence exemption are carried as a pnpm dependency patch until ACPX exposes/fixes those behaviors upstream. - Child stderr is synchronously appended to preserve ordering and failure diagnostics; unusually high-volume agent stderr could briefly block the Node event loop. - The Windows-specific `.cmd` resolution and symlink `EPERM` branches are proven by the standalone smoke test and the linked one-time `windows-latest` run rather than a permanent CI gate. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.4 via Codex CLI, medium reasoning, repository/tool execution enabled; context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b668379bc |
Regenerate embedded-postgres vendor patch
Rebuild the patch file with valid unified-diff hunks so pnpm can apply the locale and environment fixes during install. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f352f3f514 |
Force embedded-postgres messages locale to C
The vendor package still hardcoded LC_MESSAGES to en_US.UTF-8. That locale is missing in slim containers, and initdb fails during bootstrap even when --lc-messages=C is passed later. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4ff460de38 |
Fix embedded-postgres patch env lookup
Use globalThis.process.env in the vendor patch so the spawned child process config does not trip over the local process binding inside embedded-postgres. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5602576ae1 |
Fix embedded Postgres initdb failure in Docker slim containers
The embedded-postgres library hardcodes --lc-messages=en_US.UTF-8 and strips the parent process environment when spawning initdb/postgres. In slim Docker images (e.g. node:20-bookworm-slim), the en_US.UTF-8 locale isn't installed, causing initdb to exit with code 1. Two fixes applied: 1. Add --lc-messages=C to all initdbFlags arrays (overrides the library's hardcoded locale since our flags come after in the spread) 2. pnpm patch on embedded-postgres to preserve process.env in spawn calls, preventing loss of PATH, LD_LIBRARY_PATH, and other vars Co-Authored-By: Paperclip <noreply@paperclip.ing> |