mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
bef9288669baec98befaa77bcb7a103df7bcbd10
268
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5db8ce3c44 |
fix(docker): make tini PID 1 in the server image so adopted orphans are reaped (#12137)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runs execute inside the server container, and they spawn many
short-lived descendants: git, the adapter CLI, esbuild, sh
> - The server image sets `ENTRYPOINT ["docker-entrypoint.sh"]`, and
that entrypoint ends in `exec`, so node becomes PID 1
> - Node reaps only the children it spawned itself. It installs no
`SIGCHLD`/`waitpid` handler for orphans that the kernel re-parents onto
PID 1, so those orphans stay as zombies forever
> - Zombies accumulate monotonically. When the cgroup pid limit is
reached, every `fork()` in the container fails and the instance is dead
> - This pull request installs `tini` and makes it PID 1 in front of the
existing entrypoint, adds a behavioural test that proves reaping, and
adds a `pids_limit` backstop to both compose files
> - The benefit is that a long-running container no longer degrades into
total fork failure, and a future regression is caught by CI instead of
by an outage
Depends-on: none — this change is self-contained in the image build and
its tests, and it touches no other in-flight branch
## Linked Issues or Issue Description
No public GitHub issue exists for this defect. It was found on a live
long-running instance. Description follows the bug report template.
**What happened?**
The server container ran for 22 hours and reached 2039 of 2048 pids in
its cgroup. Of 1760 processes, 1731 were zombies, and all 1731 had PID 1
as their parent. PID 1 was `node --import
./server/node_modules/tsx/dist/loader.mjs server/dist/index.js`. Zombies
accrued at about 79 per hour and were never reaped. The oldest zombie
was 20.8 hours old against a container uptime of 22.0 hours, so nothing
had been reaped since boot. Once the pid limit was reached, `git` and
`gh` failed with `pthread_create failed: Resource temporarily
unavailable`.
**Expected behavior**
PID 1 reaps orphaned processes that the kernel re-parents onto it. The
pid count of a long-running container stays flat instead of growing
without bound.
**Steps to reproduce**
1. Start the server image without `docker run --init` and without `init:
true`.
2. Run agent work that spawns descendants which outlive their immediate
parent.
3. Read `/sys/fs/cgroup/pids.current` and count processes in `Z` state
over several hours.
4. The zombie count grows monotonically and every zombie has PPID 1.
**Relevant logs or output**
```
cgroup pids.current / pids.max : 2039 / 2048
total processes : 1760
zombies : 1731 (98.4%)
parent of every zombie : PID 1 (1731/1731)
PID 1 cmdline : node --import .../tsx/dist/loader.mjs server/dist/index.js
container uptime : 22.0 h
oldest zombie : 20.8 h median: 14.4 h
zombie names : git 717, claude 280, MainThread 167, sleep 141,
esbuild 138, postgres 76, sh 65, sccache 50
```
**Additional context**
The fix pattern is already in this repository.
`docker/agent-runtime/Dockerfile.base` installs `tini` and sets
`ENTRYPOINT ["/usr/bin/tini", "--"]`. It was never applied to the server
image.
## What Changed
- `Dockerfile`: install `tini` in the `base` stage and set `ENTRYPOINT
["/usr/bin/tini", "--", "docker-entrypoint.sh"]`. The entrypoint stays
in the exec chain, so UID/GID remapping, `gosu`, and graceful shutdown
are unchanged.
- `scripts/assert-orphan-reaping.sh` (new): a behavioural probe. It
spawns a leader that forks a grandchild, exits the leader, and asserts
that the orphaned grandchild leaves `Z` state instead of persisting. It
fails closed if the grandchild is not re-parented onto PID 1, so a pass
cannot mean the check ran too early.
- `.github/workflows/docker.yml`: run that probe against the pushed
image after the publish step. The publish step is multi-arch with `push:
true`, so nothing is loaded into the runner daemon and the pushed tag is
the only thing to test. The cloud variant is `FROM production` and
inherits the same `ENTRYPOINT`.
- `scripts/docker-build-test.sh`: run the same probe against a local
build.
- `docker/docker-compose.yml` and
`docker/docker-compose.quickstart.yml`: add `pids_limit: 2048` as a
backstop, so a future leak dies visibly at its own ceiling instead of
starving the host of pids.
- `server/src/__tests__/container-init-reaping.test.ts` (new): 13
assertions that guard the configuration the probe depends on.
No per-orchestrator init lever was added. The image owning PID 1 covers
compose, plain `docker run`, the quadlet units, and the ECS task
definition in one place. Adding `init: true` in compose or
`initProcessEnabled` on the ECS task would nest a second init around
`tini`, and `tini` then warns on every boot that it is not PID 1. The
new test asserts the absence of both levers across all three manifests,
so the decision survives the next edit.
## Verification
| Check | Result |
|---|---|
| `scripts/assert-orphan-reaping.sh` against a real init | Grandchild
re-parented to PPID 1, then reaped. Exit 0. |
| Same probe forced against a genuine zombie | Reports `Z` and fails.
The failure branch is not vacuous. |
| Config guard against the pre-fix files | Exactly the 3 relevant
assertions turn red. |
| Config guard with `tini` removed from `apt-get` but the comments kept
| Red. It checks the install, not a mention of the name. |
| `cd server && npx vitest run
src/__tests__/container-init-reaping.test.ts` | 13 passed |
| `npx tsc --noEmit -p server` | Clean |
| `node scripts/check-docker-deps-stage.mjs` | PASS |
| `node --test scripts/release-verify-workflow.test.mjs` | 8 passed |
Not verified locally: no container runtime is available in the authoring
environment, so the probe has not run against a build of this image. The
new `docker.yml` step runs it against the pushed image on this PR.
## Risks
Low risk, but it is an image and entrypoint change, so it affects
deployments.
- `tini` adds one small package to the `base` stage.
`docker/agent-runtime/Dockerfile.base` already installs it from the same
Debian archive.
- Signal handling changes shape: `tini` receives `SIGTERM` and forwards
it to the entrypoint, which `exec`s node. `tini` forwards signals to its
direct child by default, and the exec chain keeps node as that child, so
graceful shutdown is preserved. A reviewer should confirm this on a real
stop.
- `pids_limit: 2048` is new for compose users. A deployment that
legitimately needs more than 2048 processes would now hit the ceiling.
The measured steady state on a busy instance was under 400.
- If a deployment already passes `--init` or `init: true`, `tini` runs
under another init and prints a warning that it is not PID 1. Reaping
still works because the outer init handles it. The compose files in this
repository do not set `init: true`.
## Model Used
Claude Opus 5 (`claude-opus-5`), extended thinking, with tool use and
code execution in an agent harness.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local issues or links
- [x] My branch name describes the change and contains no internal
ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: zannis <1011451+zannis@users.noreply.github.com>
|
||
|
|
ffff1fe6e3 |
feat(runner): define package API and verification boundary (#12129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package now has protocol, transport, provider, catalog, and authorization foundations. > - Its first upstream package boundary should expose only the implemented runtime and test-helper surfaces. > - Rust correctness belongs in the repository existing build verification, without introducing a parallel release process. > - Direct package creation must build the files declared by the package manifest. > - This pull request defines the minimal package API and verifies the optimized runner binaries in the existing PR and release Build jobs. > - The benefit is a production-ready runner package boundary with minimal build-process change. ## Linked Issues or Issue Description Refs #11962 This pull request replaces one bounded part of the archived large runner change. It follows the package-local authorization change in #12126. ## What Changed - Export only `@paperclipai/paperclip-runner` and `@paperclipai/paperclip-runner/testing`. - Keep Node-only fixture loading and semantic conformance helpers out of the runtime root. - Add a provider-neutral semantic conformance kit with stable JSON comparison and fail-closed input checks. - Keep deferred SDK, eval, browser, React, lab, and command surfaces private. - Pin the runner Rust toolchain to 1.97.1 with the minimal profile and `rustfmt`. - Run the Rust workspace tests in release mode. - Launch the optimized `paperclip-runnerd` and fake-harness binaries in process-level integration coverage. - Add one `pnpm --filter @paperclipai/paperclip-runner check:all` step to each existing PR and release Build job. - Make the existing server `prepack` lifecycle run its existing build after it prepares UI assets. - Document that no production adapter starts runnerd yet. This revision adds no standalone GitHub Actions job. It adds no server runner dependency or runner vendoring. It adds no Docker bootstrap or clean-consumer harness. It does not change `pnpm-lock.yaml`. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` - 66 TypeScript tests - 8 protocol contract tests - 56 Rust unit and integration tests - Release-mode integration coverage launches the optimized runnerd and fake-harness binaries. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-package-build-script.test.ts` (2 tests) - Clean `pnpm pack` from `server/` rebuilt the server and produced both `package/dist/index.js` and `package/dist/index.d.ts`. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` (8 tests) - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `git diff --check` - No `pnpm-lock.yaml` diff. - The diff changes 12 files. ## Risks The runner adds Rust work to the existing Build jobs. These jobs can take longer on a cold cache. The pinned toolchain makes contributor and CI behavior reproducible. Cargo tests use `--release` to verify optimized executables. The server prepack lifecycle now performs the build that its published entry points require. This can make direct server packing slower. This pull request does not wire runnerd into the server. It does not select runnerd for any adapter. Existing application execution and finalization paths remain unchanged. ## Model Used OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a01444514 |
test(release-smoke): cover the background-service leg of onboarding (#12151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release pipeline gates each nightly and beta on a smoke suite that onboards the published npm artifact and drives the golden path > - That smoke runs onboarding inside a Docker container, and containers have no service manager, so the background-service leg of onboarding has zero automated coverage > - v2026.824.0 shipped a service install that crash-looped on a missing shim, and every smoke check stayed green (#12148 fixed the defect itself) > - This pull request adds a `smoke_service` job that runs the same published artifact directly on the runner VM's systemd and requires the installed service to end up serving > - The benefit is that a release with a broken service install can no longer pass the release smoke suite ## Linked Issues or Issue Description Refs #12148 — the fix for the defect this coverage gap let through. The gap: the release smoke runs `onboard` with `--yes` inside Docker, which both skips the service prompt and lacks systemd, so no CI job ever executed `manager.install()` against a real service manager. ## What Changed - New `scripts/service-onboard-smoke.sh`: onboards the published artifact with `--yes --install-service` on a systemd host, then fails unless the managed shim exists and is executable, `paperclipai.service` is active, and `/api/health` answers. A health response while the unit is not active also fails, because that is the signature of something other than the service serving. The script refuses to run over an existing managed install unless `SMOKE_FORCE=true`, and cleans up after itself by default so it is safe to run locally. - New `smoke_service` job in `.github/workflows/release-smoke.yml`: starts a user systemd session on the hosted runner (`loginctl enable-linger` + exported `XDG_RUNTIME_DIR`/`DBUS_SESSION_BUS_ADDRESS`), runs the script against `inputs.paperclip_version`, and uploads `systemctl status` + journal output as diagnostics. - No `release.yml` changes needed: `smoke_nightly` and `smoke_beta` call this reusable workflow, and a `workflow_call` result aggregates all jobs, so the new job gates nightly promotion automatically. ## Verification - `bash -n scripts/service-onboard-smoke.sh` passes and the workflow YAML parses. - End-to-end: dispatched this branch's Release Smoke workflow against the published canary that contains #12148; the `smoke_service` job onboards, installs the service, and verifies the service serves health. (Run link in PR comments.) - Negative case: the same assertions fail against v2026.824.0 — reproduced in a systemd container during the #12148 investigation: shim missing, unit in a 203/EXEC restart loop. ## Risks - Low risk to the product: no application code changes. - Pipeline risk: a flaky user-session setup on the hosted runner would block nightly promotion. Mitigated by validating the job end-to-end from this branch before merge, a 30-minute job timeout, and diagnostics uploaded on every run. - The service leg only covers systemd. launchd (macOS) still has no CI coverage; a macOS runner job is a possible follow-up. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
890ab9acfe |
feat(release): thorough notes skeletons — nest each PR's summary at creation (#12124)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release workflow drafts the upcoming stable's notes skeleton the moment a beta publishes > - That skeleton was a bare list of commit subjects, so the notes only reached the shipped stable's depth after a later authoring pass during the soak > - Stable release notes are consistently verbose and thorough; the initial draft should start that way too > - This pull request nests each referenced PR's own summary under its subject line at creation time, and states the density bar in the authoring skill > - The benefit is a thorough raw document from day one of the soak, with no LLM tokens in Actions ## Linked Issues or Issue Description **What existing behavior does this improve?** The `draft_stable_notes` skeleton generated at beta publish (`scripts/draft-stable-notes.sh`). **Current behavior** The skeleton groups bare commit subjects by conventional-commit type. All substance arrives later, when a maintainer or agent rewrites it — reviewed maintainer feedback: stable notes are a lot more verbose, and the initial beta notes should be consistent with that. **Proposed behavior** Each subject that references a PR carries that PR's own summary nested beneath it — the PR template's "What Changed" bullets, else the first prose lines — fetched best-effort via `gh` and skipped silently when unavailable. The release-changelog skill now states the density bar explicitly: the beta-keyed draft ships verbatim as the stable's notes and is written at the previous stable's depth from the first pass. **Reason and benefit** The notes author starts from a thorough raw document instead of a commit list, and beta-time notes match the verbosity the stable will ship with. ## What Changed - `scripts/draft-stable-notes.sh`: `enrich_pr` nests PR summaries under subjects; best-effort (`gh` failure or `DRAFT_NOTES_SKIP_PR_ENRICHMENT=1` degrades to today's output); pipefail-safe when a "What Changed" section has no bullets. - `.github/workflows/release.yml`: the `draft_stable_notes` step gets `GH_TOKEN` so `gh` can read PR bodies. - `.agents/skills/release-changelog/SKILL.md`: "write at full stable depth from the first pass" guideline. - `scripts/draft-stable-notes.test.mjs`: three new tests — enrichment rendering via a fake `gh`, silent degradation without one, and the sparse-body case that previously killed the script under `set -o pipefail`. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 11 pass. - Live run against the real repository for the current beta (`2026.818.0-beta.1`, 172 commits): exit 0, 439 nested summary lines; spot-checked entries carry the correct PRs' What Changed bullets. - `bash -n` on the script; `release.yml` re-parsed as YAML. ## Risks - Low: the publish path is untouched; enrichment is read-only `gh` calls in the post-publish draft job and degrades to the current skeleton on any failure. Roughly one API call per commit in the range (~170 today) — well inside the token's rate budget, adds a couple of minutes to a job with a 10-minute timeout. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0dfa0fb988 |
ci: refresh general-server shard duration manifest (#12075)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its slowest check sets the feedback time for all contributors > - The general-server test lane splits its vitest suites across five runners with a duration-weighted partition (`scripts/general-server-shard.mjs`) > - The partition reads a duration manifest that was sampled on 2026-08-04, when the lane had 279 suites and 946s of serial time > - The lane has since grown to 405 suites and 1274s; 126 suites had no recorded duration and one suite grew from 37s to 123s > - The stale weights made the partition uneven: in the fully green actions run 32708351172, "General tests (server (2/5))" ran 364s and was the slowest check in the whole run, while sibling shards ran 292-330s > - This pull request refreshes the manifest with per-suite durations measured from that same run > - The benefit is a level five-shard split (255s ±1s of predicted suite time per shard), which removes ~50s from the slowest PR check ## Linked Issues or Issue Description **Describe the current behavior** In the fully green PR actions run [32708351172](https://github.com/paperclipai/paperclip/actions/runs/32708351172) (2026-08-24), the check "General tests (server (2/5))" completed in 364s. Its test step ran 315s while sibling shards ran 241-276s. It was the slowest check in the run. **Describe the improvement** The duration manifest `scripts/general-server-shard-durations.json` is stale. It holds 279 suites sampled on 2026-08-04, but the lane now has 405 suites. The 126 unknown suites fall back to the median weight (~1.3s), and `server/src/__tests__/workspace-runtime.test.ts` grew from 37.4s to 123.3s. The partition therefore predicts a level split but produces an uneven one. Refreshing the manifest restores the level split without any code change. **Expected impact** All five server shards level at ~255s of predicted suite time (~310s job time). The slowest PR check drops from 364s to about 317s, so the PR critical path improves by roughly 50s. ## What Changed - Regenerated `scripts/general-server-shard-durations.json` from actions run 32708351172 (2026-08-24): 405 suites, 1274s total serial time (was 279 suites, 946s from 2026-08-04) - Updated the `$comment` field to name the new sample run and date - No code changes; the partition logic in `scripts/general-server-shard.mjs` is untouched ## Verification - Parsed all five "General tests (server (n/5))" job logs from run 32708351172 with the consecutive-completion-timestamp method described in the manifest `$comment`; asserted that the parsed suite set equals the exact file list that `run-vitest-stable.mjs` collects (405/405, no misses, no extras) - Ran `node scripts/run-vitest-stable.mjs --mode general --group general-server --shard-index N --shard-count 5 --dry-run` for N=0..4 with the new manifest: each shard predicts 255s (±1s) of suite time, and the five shards form a complete, non-overlapping cover of all 405 suites - Ran `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs`: 13/13 pass ## Risks - Low risk. The change is data-only. Wrong weights cannot break correctness: the partition always covers every suite exactly once, so the worst case of a bad weight is an uneven shard, which is the current state. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable (no code change; existing partition tests pass) - [x] I have updated relevant documentation to reflect my changes (manifest `$comment` updated) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #11528 (balanced the serialized server shards by recorded duration), #10923 (split serialized tests into five shards), #11156 (split workspaces-a into two shards). Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
633e102971 |
fix: verify issue-update writes instead of inferring success (#12051)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents report task state to the control plane with `PATCH
/api/issues/{id}` at the end of each heartbeat
> - On remote sandbox targets those writes cross a relay that can fail
at the connection level
> - An agent that pipes its status curl through `head` cannot see that
failure; the write is lost but the run reports success
> - The issue then stays `in_progress` with no disposition, and the
missing-disposition recovery must repair it
> - This pull request makes the issue-update helper verify every write,
and it teaches the shared skill to require verified writes
> - The benefit is that a lost status write becomes a visible, retried
failure instead of a silent success
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
A sandboxed heartbeat run answered its issue in a comment. It then sent
`PATCH /api/issues/{id}` with `status: done` through `curl -sf ... |
head -c 400`. The relay dropped the connection. The `-f` flag suppressed
the error output, and the pipe replaced curl's exit code with the exit
code of `head`. The agent saw empty output and exit 0. It reported the
write as an "empty 2xx" success and exited. The issue stayed
`in_progress`, and the successful-run recovery had to close it in a
corrective run.
**Expected behavior**
A status write that does not reach the server must surface as a failure.
The helper script must retry transient failures. It must exit non-zero
when the write is unconfirmed. Skill guidance must forbid write patterns
that hide failures.
**Steps to reproduce**
1. Point `PAPERCLIP_API_URL` at an endpoint that drops connections
intermittently.
2. Finalize an issue with `curl -sf -X PATCH
"$PAPERCLIP_API_URL/api/issues/$ID" -d '{"status":"done"}' | head -c
400`.
3. Observe exit code 0 with empty output while the server never received
the PATCH.
## What Changed
- `scripts/paperclip-issue-update.sh` now captures `%{http_code}`,
retries a retryable failure (connection-level, 429, 5xx) once — two
attempts total, which matches the shared bounded-write-retry rule —
rejects an empty 2xx body, and confirms the response echoes the
requested status before it exits 0.
- Failure output states plainly that the write was NOT saved, so the
calling agent reports it accurately.
- `skills/paperclip/SKILL.md` Step 8 adds a required "Verify writes —
never infer them" rule: a successful PATCH always returns the updated
issue JSON, disposition writes must never run through `head`/`tail`
pipelines, and an unconfirmed write must be reported as FAILED.
- `server/src/__tests__/paperclip-skill-utils.test.ts` pins the new
skill rule; a new `paperclip-issue-update-helper.test.ts` exercises the
helper's behavior end-to-end.
## Verification
- `bash -n scripts/paperclip-issue-update.sh`
- `server/src/__tests__/paperclip-issue-update-helper.test.ts` runs the
helper end-to-end against a local HTTP server: confirmed-echo success
(exit 0), empty 2xx (exit 1), wrong echoed status (exit 1), 422 reject
(exit 1, exactly one request), 503 then success (two requests),
connection refused (two attempts, then exit 1 with a "NOT saved"
report).
- `npx vitest run
server/src/__tests__/paperclip-issue-update-helper.test.ts
server/src/__tests__/paperclip-skill-utils.test.ts
server/src/__tests__/cli-invocation-safety.test.ts` — 50 passed.
## Risks
- Low risk. The success-path output is unchanged (the updated issue
JSON).
- The helper now exits non-zero on unconfirmed writes. Callers that
previously missed silent failures now see explicit errors. That is the
intended behavior change.
- The single retry re-sends the PATCH after a retryable failure. If the
first request committed and only its response was lost, an attached
comment can post twice. The duplicate is visible and benign; the prior
behavior lost the write silently.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
dc5b070709 |
fix(runtime): guard empty Bash 3.2 array expansion (#11891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in isolated worktrees with a separate Paperclip runtime. > - Runtime provisioning uses a Bash script on macOS hosts. > - macOS ships Bash 3.2, where an empty array expansion fails under `set -u`. > - The source-config argument array is empty when the base workspace already has a config. > - This pull request guards that expansion and tests the normal base-config path on Bash 3.2. > - The benefit is that managed worktree provisioning no longer fails before database seeding. ## Linked Issues or Issue Description No public GitHub issue exists for this problem. PR #11752 added the conditional source-config argument that exposed the failure. **What happened?** `scripts/provision-worktree-runtime.sh` expands an empty `source_config_args` array while `set -u` is active. Bash 3.2 reports `source_config_args[@]: unbound variable` and stops provisioning when the registered base workspace already has `.paperclip/config.json`. **Expected behavior** Runtime provisioning must call `worktree ensure-seeded` without a source override when the base workspace config exists. It must work with the Bash 3.2 version that macOS supplies. **Steps to reproduce** 1. Use macOS system Bash 3.2. 2. Create a base workspace with `.paperclip/config.json`. 3. Run `scripts/provision-worktree-runtime.sh` with `set -u` active in the script. 4. Observe the unbound-variable error before `worktree ensure-seeded` runs. **Paperclip version or commit** Reproduced on `origin/master` before this change. **Deployment mode** Local managed worktree runtime on macOS. ## What Changed - Guard all three optional source-config array expansions with Bash 3.2-compatible parameter expansion. - Add a regression test that uses the base-config path and verifies that no `--from-config` argument is sent. - Document the Bash 3.2 compatibility requirement in the runtime script. ## Verification - `/bin/bash -n scripts/provision-worktree-runtime.sh` - `node --test --test-name-pattern='runtime provisioning invokes ensure-seeded once|runtime provisioning omits the source override|runtime provisioning guards every optional source-config expansion' scripts/__tests__/provision-worktree-self-heal.test.mjs` - `git diff --check` ## Risks Low risk. The change only affects expansion of an optional two-element CLI argument array. The regression tests cover both the empty and non-empty paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model ID `gpt-5`. The runtime did not expose the context-window size. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
417336f8be |
fix(workspaces): attach PR preparation to existing branches (#11703)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces isolate an agent task from the primary checkout. > - Pull request preparation can need a branch that already contains completed work. > - The workspace policy could not require an exact existing branch. > - Workspace cleanup also treated worktree creation as branch ownership. > - This pull request adds an exact existing-branch policy and separate branch ownership metadata. > - The benefit is safe pull request preparation that preserves every existing commit and operator-owned branch. ## Linked Issues or Issue Description **What happened?** A pull request preparation run could not pin its execution workspace to an exact existing branch. Workspace reuse and cleanup could also confuse worktree creation with branch ownership. **Expected behavior** The run must attach only to the requested branch in an isolated Git worktree. It must fail if the branch is missing, busy, or inconsistent. Cleanup must not delete a branch that Paperclip does not own. **Steps to reproduce** 1. Create a branch that contains completed work. 2. Configure a pull request preparation task to use that branch. 3. Start the task and observe that the prior policy cannot require the exact branch. **Paperclip version or commit** This behavior reproduces on the base revision before this pull request. **Deployment mode** Local development with isolated Git worktrees. ## What Changed - Add `existingBranch` to the execution workspace policy and shared validation contracts. - Require `existingBranch` to use an isolated Git worktree and reject conflicting branch templates. - Attach to the exact branch without creating, renaming, resetting, or deleting it. - Track branch ownership separately from worktree creation and use that ownership during cleanup. - Return HTTP 422 for invalid existing-branch settings on every issue-producing route. - Add a bounded repair script for existing pull request preparation tasks. - Add focused policy, route, heartbeat, runtime, and ready-comment tests. - Document the exact-branch behavior and safety rules. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspace-policy.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-existing-branch-validation-status.test.ts server/src/__tests__/workspace-runtime.test.ts server/src/services/workspace-runtime-exposure.test.ts server/src/services/workspace-runtime-ready-comment.test.ts` passed 335 tests. - `pnpm -r typecheck` passed for all workspace projects. - `pnpm test:run` passed 4,431 tests. Two unrelated embedded-Postgres setup hooks timed out under aggregate load. Their isolated rerun passed 74 tests. - `pnpm build` passed for all workspace projects. - The two review regressions passed with 139 unrelated tests skipped. - All latest-head CI gates passed after one unrelated timing-sensitive test passed on rerun. - Greptile scored the latest head 5/5 with no unresolved review threads. ## Risks - Invalid workspace settings now return HTTP 422 instead of a generic validation response. - The exact branch must already exist and must not be checked out by another worktree. - The new policy fails closed when it cannot prove branch identity or ownership. - This change has no database migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family assisted with this change. The runtime did not expose its exact deployment ID or context window. The agent used high-reasoning mode, repository tools, shell execution, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
233c12f029 |
feat: add kimi-local adapter for Kimi Code CLI (CLI + ACP engines) (#9967)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters (`claude_local`, `gemini_local`, `grok_local`, …) are the integration surface that lets Paperclip run coding CLIs on the host machine > - The Kimi Code CLI (`kimi`, Moonshot AI) has a documented non-interactive mode, `kimi -p --output-format stream-json` with session resume via `kimi -r`, but Paperclip has no built-in adapter for it > - So Kimi users (especially Kimi membership / OAuth subscribers) cannot onboard their CLI to Paperclip agent teams > - This pull request adds a complete built-in `kimi_local` adapter (both execution engines, session management, instructions + skills delivery, thinking-effort control, environment test, UI and CLI modules, docs) following the established `gemini_local`/`grok_local` package pattern > - Kimi Code ships an ACP server (`kimi acp`), so the adapter runs on Paperclip's shared acpx engine by default (streaming transcript with live tool status, like `claude_local`/`gemini_local`) and falls back to a headless CLI lane (`kimi -p --output-format stream-json`) when ACP prerequisites are unavailable > - The benefit is that Kimi Code becomes a first-class Paperclip agent lane: selectable in the UI, resumable across heartbeats, with the same operating context (instruction bundle, skills, effort) and streaming transcript the other local adapters get ## Linked Issues or Issue Description - Supersedes #9880 (same branch; expanded from the CLI-only lane into a complete adapter with the default ACP engine lane, control-plane skill install, and live transcript wiring) - Refs #9879 (adapter request for Kimi Code CLI, filed with this PR) - Refs #163 (original Kimi support request) Duplicate/related prior PRs, per the dedup search (both appear stale: no updates or maintainer review since May 2026, and both target an older Kimi CLI interface; calling them out for reviewer context per CONTRIBUTING.md): - Refs #6276 (`feat: add kimi-local adapter`): targets an older array-based content format (`{type: think}`/`{type: text}` blocks), not the current documented stream-json schema - Refs #5202 (`feat(adapter): add Kimi CLI local adapter with Wire protocol support`): builds on a `--wire` JSON-RPC interface that current Kimi Code CLI (0.27.0) no longer documents; the current documented headless interface is `-p --output-format stream-json` This PR is a fresh implementation against current master and the currently documented/verified Kimi CLI behavior (see Verification). Happy to fold in anything useful from the earlier attempts if a reviewer prefers. ## What Changed - **New adapter package** `packages/adapters/kimi-local` (`@paperclipai/adapter-kimi-local`), modeled on `gemini-local`/`grok-local`: - `src/server/execute.ts`: spawns `kimi -p <prompt> --output-format stream-json` (argv array, no shell), `-m <model>` only when configured, `-r <sessionId>` when the stored session cwd matches the run cwd, automatic fresh-session retry on unrecoverable-session errors, headless-safe env (`CI=1`, `NO_COLOR=1`, `KIMI_CODE_NO_AUTO_UPDATE=1`, `TERM=dumb`; user-configured values win), full remote (ssh/sandbox) execution lane with runtime install via `@moonshot-ai/kimi-code` - **Instruction bundle delivery**: the prompt path directive now names the sibling instruction files (`./HEARTBEAT.md`, `./SOUL.md`, `./TOOLS.md`) alongside the prepended entry file, and local runs pass `--add-dir <instructions-dir>` so Kimi can actually open them (matching `claude_local`). Without this, only the entry file reached Kimi and agents improvised the operating workflow that `HEARTBEAT.md` documents - **Thinking effort**: a configured `effort` is forwarded as the `KIMI_MODEL_THINKING_EFFORT` operational override (Kimi has no per-invocation effort flag). It is only sent for models that advertise `support_efforts` (currently `kimi-code/k3`) to avoid provider rejections, and `medium` maps to `high` since Kimi has no medium tier (`low`/`high`/`max` pass through) - **Skills delivery**: desired Paperclip skills are delivered via Kimi's `--skills-dir` flag from a dedicated per-run directory (a local snapshot, or the synced snapshot on remote targets), so skills load reliably and in isolation. Paperclip never overwrites the shared `$KIMI_CODE_HOME/skills` home, so skills installed by the operator or other agents are left intact. `--skills-dir` is only passed when at least one skill is desired, so unconfigured agents keep Kimi's default skill discovery - **Live run status**: the adapter now forwards each streamed stream-json line to `onEvent` (assistant `content` as an assistant snippet, `tool_calls` as tool-name events), which drives the issue-thread activity indicator (`currentToolName` / `lastAssistantSnippet` / `lastEventAt`). Previously the adapter only wrote the raw run log, so the issue thread showed a stale "no output for N s" line with no tool or reasoning context while Kimi worked. Tool results are omitted so the last meaningful "Using X" / snippet is not overwritten by a generic label - `src/server/parse.ts`: parses the verified Kimi stream-json event shapes (`assistant` text, `assistant.tool_calls` with JSON-string arguments, `tool` results, trailing `meta.session.resume_hint` for session-id capture) plus failure classifiers (`kimi_auth_required`, transient network, unrecoverable session). A signaled exit (null exit code, not a timeout) is now reported as a failure rather than coalesced to success, and the error message names the terminating signal - `src/server/skills.ts`: lists/syncs Paperclip skills for the adapter's skill-management surface - `src/server/test.ts`: environment test covering CLI resolution + `kimi --version`, cwd check, auth detection (OAuth credential dirs, keyed `[providers.*]` in config.toml, or the `KIMI_MODEL_NAME` + `KIMI_MODEL_API_KEY` env pair), and a live hello probe - `src/ui/` (stdout-line parser for transcripts, config builder) and `src/cli/` (stream event formatter) modules - Root metadata: three managed model aliases (`kimi-code/kimi-for-coding`, `kimi-code/kimi-for-coding-highspeed`, `kimi-code/k3`), effort-capable-model metadata (`EFFORT_CAPABLE_MODELS`, effort mapping helpers), `agentConfigurationDoc` - Tests: 101 tests across parse, execute (args building, resume gating, retry, auth error code, timeout, signaled-exit failure, effort forwarding/gating/mapping, `--add-dir` instructions directive, `--skills-dir` gating, `onEvent` runtime-event forwarding), ACP engine (engine resolution, acpx config build, node-version gate), ACP transcript delegation, environment test, UI parse/build-config - **ACP engine lane (default)** (`src/server/acp.ts` + shared `adapter-utils/acpx-engine`): Kimi Code ships an ACP server (`kimi acp`), so `kimi_local` now runs on Paperclip's shared acpx engine by default, matching `claude_local`/`codex_local`/`gemini_local`. The issue-thread transcript streams live (assistant text deltas, tool calls with a `pending`->`completed` status lifecycle) instead of the CLI lane's bursty complete-message output. Registered `kimi_local -> "kimi"` in `ACPX_ADAPTER_AGENT_IDS` and resolved the built-in agent command to `kimi acp`; `execute.ts` dispatches to the ACP executor first with an automatic CLI fallback when ACP prerequisites fail (`engine=acp` requires ACP, `engine=cli` pins the headless lane); `index.ts` falls back to the shared acpx session codec; the UI/CLI delegate `acpx.*` events to the shared acpx transcript parser and event formatter. The headless CLI lane (above) remains as the fallback - **Registration** (one entry each, mirroring existing adapters): server adapter registry + `BUILTIN_ADAPTER_TYPES`, `AGENT_ADAPTER_TYPES` (shared), UI adapter registry + display registry (`Kimi Code`, Moon icon) + capabilities defaults, CLI adapter registry, `Dockerfile` (package copy + `npm install --global @moonshot-ai/kimi-code@latest`), `vitest.config.ts` workspace, `scripts/release-package-manifest.json` - **Behavioral sets** mirroring `gemini_local` (Kimi resumes sessions the same way): `GIT_SENSITIVE_LOCAL_ADAPTER_TYPES`, `SESSIONED_LOCAL_ADAPTERS` (heartbeat + recovery), `REMOTE_MANAGED_ADAPTERS`, ssh/sandbox execution-target allow-lists, `ADAPTER_DEFAULT_RULES_BY_TYPE` (`timeoutSec: 0`, `graceSec: 15`), and `LEGACY_SESSIONED_ADAPTER_TYPES` + `ADAPTER_SESSION_MANAGEMENT` in adapter-utils - **UI touch-points**: New Agent default-model branch, AgentConfigForm command map (`kimi_local: "kimi"`) + model defaults + a Kimi-specific thinking-effort option list (`Low`/`High`/`Max`, reflecting Kimi's tiers rather than borrowing Claude's), OnboardingWizard (command map, model default, `kimi login` / `KIMI_MODEL_NAME + KIMI_MODEL_API_KEY` auth hints, manual-debug command line), InviteLanding enabled adapters - **Control-plane skill install** (`cli/src/commands/client/agent.ts`): `paperclipai agent local-cli` seeded the Paperclip control-plane skills into `~/.codex/skills` and `~/.claude/skills` so Codex/Claude agents auto-discover the API reference every run. Kimi had no equivalent target, so `kimi_local` agents began each session without the control-plane skill and rediscovered routes (e.g. the company-scoped `POST /api/companies/{companyId}/issues`) by trial and error. Added `~/.kimi-code/skills` (honoring `KIMI_CODE_HOME`) as a third install target for parity. Independent of the per-run `--skills-dir` delivery, which only applies to explicitly configured skills. - **Docs**: `docs/adapters/kimi-local.md` (prerequisites, auth options, config fields including `effort`, session resume, instruction bundle, skills delivery, control-plane skill install) + a row in `docs/adapters/overview.md` Out of scope (deliberately): model profiles, built-in agent `allowedAdapterTypes` additions. ## Verification\n\nCurrent-master rebase verification (OpenAI Codex, 2026-08-03): 13 focused files / 231 tests pass; adapter-utils, server, UI, CLI, and Kimi adapter typechecks pass; full repository build and UI token gates pass. The branch is conflict-free against master at head `1249df117c5e12e5771b9a570a6340866450619e`.\n\nAutomated (all from repo root, pnpm 9.15.4, Node 22): - `vitest run packages/adapters/kimi-local`: 89/89 pass (includes coverage for the instruction `--add-dir` directive, effort forwarding/gating/mapping, `--skills-dir` gating, the signaled-exit failure path, and `onEvent` runtime-event forwarding with cross-chunk line buffering) - `vitest run server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/services/heartbeat-stop-metadata.test.ts ui/src/adapters/adapter-display-registry.test.ts`: 37/37 pass - `vitest run cli/src/__tests__/skills.test.ts`: 13/13 pass (the control-plane skill install target follows the existing Codex/Claude install path, whose symlink logic is unchanged) - `vitest run packages/shared`: 307/307 pass; `vitest run packages/adapter-utils`: pass except one pre-existing, unrelated failure (`mcp-isolation.integration.test.ts` requires Claude CLI ≥ 2.1.207; host has 2.1.185, fails identically on unmodified master) - `pnpm --filter @paperclipai/adapter-kimi-local typecheck|build`, plus typecheck of `server`, `ui`, `cli`, `adapter-utils`: all clean - `pnpm install --frozen-lockfile`: passes (the PR diff itself contains no lockfile changes, per repo policy; verified against a locally regenerated lockfile) - `node scripts/check-no-git-push.mjs` and `node scripts/check-forbidden-tokens.mjs`: pass - CI note: the `policy` job's release-bootstrap step is expected to stay red until a maintainer bootstraps the first npm publish of `@paperclipai/adapter-kimi-local`; see the CI Note for Maintainers comment. All other contributor-actionable checks are green. Manual end-to-end (real Kimi CLI 0.27.0, OAuth login, dev server on an isolated instance): 1. Server `GET /api/adapters` lists `kimi_local` as builtin with correct capability flags; models endpoint returns the three Kimi models 2. `POST .../adapters/kimi_local/test-environment`: all checks pass, including a live `kimi -p` hello probe 3. Created a `kimi_local` agent and invoked two heartbeats: run 1 spawned `kimi -p ... --output-format stream-json`, Kimi used its `Read` tool, produced the expected answer, and the session id was captured from the `session.resume_hint` meta event; run 2 resumed the **same** Kimi session (`sessionIdBefore == sessionIdAfter`) via `-r` 4. UI: adapter appears in the New Agent dropdown; selecting it shows the Kimi command placeholder, the three models, and the Kimi config fields; the run transcript renders Kimi tool calls via the adapter's stdout parser The instruction-bundle, thinking-effort, and `--skills-dir` changes landed after the manual run above. They are covered by the unit tests listed under Automated, and the Kimi CLI flags they rely on (`--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were confirmed against the installed Kimi Code CLI 0.27.0 (`kimi --help`, config-file thinking-effort docs). Screenshots (assets branch on the fork, not part of the diff):       ## Risks - Low risk to existing behavior: the change is additive, one new workspace package plus single-entry registrations alongside existing adapters; no existing adapter code paths are modified. - The adapter invokes the locally installed `kimi` CLI; like other local adapters, run behavior depends on the host's Kimi version. The parser is written against the documented/verified 0.27.0 stream-json schema and degrades gracefully (malformed lines are skipped, failures surface as run errors). - `--skills-dir` overrides Kimi's auto-discovery of user and project skills for the run. This is intentional (paperclip-managed agents get a reproducible, isolated skill set), and it is only passed when at least one Paperclip skill is desired, so unconfigured agents keep default discovery. - Thinking effort is only forwarded to models that advertise `support_efforts` (currently `kimi-code/k3`); `EFFORT_CAPABLE_MODELS` must be extended when more Kimi models gain support, otherwise a configured effort is silently ignored for them. - `Dockerfile` now installs `@moonshot-ai/kimi-code@latest` globally alongside the other agent CLIs, so image size increases slightly. - Maintainer action needed for the npm bootstrap gate: the `policy` job's release-bootstrap step fails until the first npm publish of `@paperclipai/adapter-kimi-local` (the gate from #5146 that every new adapter package has passed through). Enrollment with `publishFromCi: true` is required by the manifest validator (dropping the entry, `false`, or `private` are all rejected), so this is intentionally left to a maintainer. Remaining CI lanes are expected to run once it is done. ## Model Used\n\n- **Current-master rebase, conflict adaptation, and registry-parity coverage:** OpenAI, **GPT-5 Codex** (Codex agent; exact serving model ID and context-window size were not exposed to the runtime), with repository, shell, Git, and GitHub tooling. It preserved Hawik’s commit authorship, reconciled ACPX and environment-capability changes, added current registry tests, and ran the verification above.\n- **Adapter implementation and initial review:** Moonshot AI, **Kimi K3 Coding** (latest), via **Kimi Code CLI v0.27.0** (`kimi-code/k3` alias, 1M-token context window, thinking mode, agentic tool use). The CLI agent explored the repo, wrote the adapter implementation (delegated to a coder sub-agent of the same model), ran tests, and drafted the first version of this PR body. A second model-driven review pass (read-only, same model) audited the diff for security/correctness before submission; its findings (shell-quoting hardening, auth-detection false positive, session-compaction registration, test gaps) were fixed and are included. - **Harness-context fixes and review responses:** Anthropic, **Claude Opus 4.8** (`claude-opus-4-8`) via Claude Code. Diagnosed from run logs that Kimi received only the entry instructions file (not the `HEARTBEAT.md`/`SOUL.md`/`TOOLS.md` bundle) and that `effort` was never wired, then implemented the instruction `--add-dir` delivery, `KIMI_MODEL_THINKING_EFFORT` forwarding, and `--skills-dir` skill delivery, added the accompanying tests and docs, and addressed the automated review comments (preserving external skills on remote sync, treating a signaled exit as a failure). Also extended the `paperclipai agent local-cli` installer to seed the control-plane skills into `~/.kimi-code/skills` for Codex/Claude parity, wired `onEvent` runtime events so the issue-thread activity indicator reflects Kimi's tool and reasoning output live, and built the ACP engine lane (`kimi acp` via the shared acpx engine, default) so the transcript streams with live tool status like the other ACP adapters. The Kimi CLI flags, subcommand, and env var relied on here were verified against the installed Kimi Code CLI 0.27.0. - All CLI behaviors claimed here (`-p`, `--output-format stream-json`, `-r` resume, event shapes, `--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were verified empirically against the installed Kimi CLI, not assumed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green *(only the release-bootstrap step remains red, pending the maintainer npm publish described in Risks)* - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(will address all Greptile comments as they arrive)* - [x] I will address all Greptile and reviewer comments before requesting merge --- ## Maintainer Addendum (2026-08-20) The shared acpx-engine and issue-chat changes (run-summary segmentation, placeholder tool-event coalescing, `ISSUE_CHAT_TRANSCRIPT_MAX_VISIBLE_ENTRIES` 30 → 400, live-reasoning UI) have been **extracted to #11761** so the cross-adapter behavior changes review and revert independently — both commits there preserve @hawikk's authorship. This PR is now the kimi-specific adapter only (60 files, +3,793/−8, essentially pure addition); the only shared-engine touch left is the `kimi acp` command resolution. `publishFromCi` is `true` — the package name is bootstrapped on npm. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Devin Foley <devin@paperclip.ing> |
||
|
|
a9d1f740f0 |
fix(workspaces): seed managed worktrees when the base checkout has no config (#11752)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents do that work in isolated git worktrees, and a managed
worktree runs its own Paperclip instance with a cloned database
> - That clone needs a seed source, and the source must come from
server-owned registration, never from state the workspace itself can
rewrite
> - The seed-source resolver requires the registered base project
workspace to hold its own `.paperclip/config.json`
> - A managed project workspace is a plain `git clone`, and no code
writes that file into it
> - Every isolated worktree provision, deferred seed, and workspace
repair therefore fails on a managed checkout
> - This pull request lets a named source supply the config when the
base checkout has none
> - The benefit is that managed worktrees provision again, and the seed
source stays server-owned
## Linked Issues or Issue Description
No public GitHub issue exists for this problem. It is described below.
**What happened?**
Agent runs that need an isolated worktree fail during provisioning. The
provision command exits with this error (paths redacted):
```
Execution workspace provision command "bash ./scripts/provision-worktree.sh" failed:
Registered base project workspace has no canonical Paperclip config:
<instance-home>/instances/default/projects/<company-id>/<project-id>/<repo>/.paperclip/config.json
```
`resolveRegisteredWorktreeSeedSource` sets `registeredConfigPath` to
`<baseCwd>/.paperclip/config.json` whenever the caller names a
registered base workspace. It then requires that file to exist.
`scripts/provision-worktree.sh` applies the same rule.
A managed project workspace never has that file.
`materializeManagedProjectWorkspace` creates it with `git clone` and a
rename, so the checkout holds repository content only. The control plane
keeps its config at `<home>/instances/<id>/config.json` instead.
The failure reaches three paths: worktree provisioning, deferred seeding
through `worktree ensure-seeded`, and workspace repair.
The behavior changed in #11671. That pull request replaced a fallback
chain with a single hard requirement. Fixture code in
`scripts/__tests__/provision-worktree-self-heal.test.mjs` writes a
config into the fake base workspace, so tests kept passing.
**Expected behavior**
A managed worktree provisions and seeds from the registered source. The
seed manifest still never selects that source.
**Steps to reproduce**
1. Register the Paperclip repository as a project with a `repoUrl`, so
the server materializes a managed checkout.
2. Assign an issue to an agent whose workspace strategy is
`git_worktree`.
3. Watch the workspace operation log for the provision command.
4. The command exits non-zero with the error above.
**Paperclip version or commit**
Reproduced on `master` at
|
||
|
|
b5a3a863c3 |
feat(release): bootstrap new npm packages with a placeholder publish (#11757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its release pipeline publishes a set of npm packages from CI with npm trusted publishing (GitHub OIDC), gated by `scripts/release-package-manifest.json` > - A brand-new package name cannot be published by CI directly: the PR bootstrap gate requires the name to resolve on npm, and a trusted-publisher rule can only be configured after the package page exists > - The current bootstrap helper closes that gap by building the package locally and publishing its real output from a maintainer machine — before the PR that adds the package has passed CI or review > - This pull request replaces that flow: the helper now publishes a minimal deprecated placeholder at version `0.0.0` that only reserves the name, so every real version ships from CI > - The benefit is that unreviewed build output never reaches npm, and the bootstrap runs from any checkout (including `master`, before the new package's PR merges) with no local build ## Linked Issues or Issue Description **What existing behavior does this improve?** The one-time npm bootstrap for a brand-new release package (`pnpm run release:bootstrap-package`). **Current behavior** The helper builds the target package locally and publishes the real build output from a maintainer machine. That content has not passed repository CI or review at publish time. The helper also requires the new package to exist in the local workspace, so it must run from the (unmerged) PR branch that adds the package. **Proposed behavior** The helper publishes a three-file placeholder at version `0.0.0` (manifest, README, and an `index.js` that throws a descriptive error), waits for the registry to show the package, then deprecates it. The PR bootstrap gate (`scripts/check-release-package-bootstrap.mjs`) only requires the name to resolve on the registry, so the placeholder satisfies it. The first real calver release from CI supersedes the placeholder, and a stable release moves `latest` off it — the same `latest` window that existed under the old flow, but containing an explicit inert stub instead of unreviewed code. **Reason and benefit** Real package content only ever reaches npm from CI, after review and merge. The bootstrap becomes safer (scope guard refuses names outside `@paperclipai/`, already-published names are rejected) and simpler (no local build, no workspace state, runs from any checkout). **Breaking changes** None at runtime. The helper's CLI surface changes: it now takes a package name only (no directory selector) and drops `--skip-build`. `doc/PUBLISHING.md` is updated to match. ## What Changed - `scripts/bootstrap-npm-package.mjs`: replaced the build-and-publish flow with a placeholder publish — stages `package.json` + `README.md` + throwing `index.js` at version `0.0.0` in a temp directory, previews with `npm publish --dry-run`, and publishes only with `--publish`. One-time passwords are prompted interactively (never passed as arguments, since they are single-use and would land in shell history), with re-prompt on a rejected or expired code. After publishing, the helper polls the registry until the package is visible (a first publish can lag by minutes; verified live at ~5 minutes), requiring two consecutive sightings before prompting for a second code and deprecating the placeholder so accidental installs warn loudly; on timeout or failure it prints the exact manual `npm deprecate` command. Added an `@paperclipai/`-scope guard and a fail-fast error when `--publish` runs without an interactive terminal. Removed the workspace-plan dependency so it runs from any checkout. - `scripts/bootstrap-npm-package.test.mjs`: rewrote for the new interface — argument parsing, scope validation, the generated placeholder files (manifest shape, throwing entry point, README), the OTP re-prompt loop, and the registry poll (consecutive-sighting requirement, timeout, transient-error tolerance) via injected fakes. - `doc/PUBLISHING.md`: rewrote the "One-time bootstrap sequence for a new package" section for the placeholder flow, including the `latest` dist-tag window and the trusted-publishing setup ordering (placeholder publish → trusted publisher rule → `"publishFromCi": true`). - `.github/scripts/check-pr-release-bootstrap.mjs` (+ test, + wiring in `run-quality-gates.mjs`): new informational commitperclip notice on PRs that need this bootstrap. It fires when the PR newly release-enables a package that is missing from npm, or adds an unpublished `publishFromCi: false` package that published packages declare a `workspace:*` dependency on, and names the exact maintainer command — so contributors know the red `policy` check is not theirs to fix. It never fails the gate (the `policy` job remains the enforcer), only looks up scope-validated names on the registry, and stays quiet on registry errors. ## Verification - `node --test scripts/bootstrap-npm-package.test.mjs`: 13/13 pass - `node --test .github/scripts/tests/*.test.mjs`: 147/147 pass (10 new for the PR notice) - `pnpm run test:release-registry`: 82/82 pass - Replayed the new PR notice against a real historical PR's live API data (files, manifest at base and head refs): with the registry in its pre-bootstrap state it produces the exact maintainer instruction; with the package bootstrapped it stays silent - Full live end-to-end run: the flow bootstrapped `@paperclipai/adapter-kimi-local` for real — dry-run preview (634-byte, 3-file tarball), publish, registry visibility after ~5 minutes of propagation lag, deprecation confirmed via `npm view ... deprecated` - Guards verified live: an already-published name is rejected, an out-of-scope name (`left-pad`) is rejected, unknown options (including the removed `--otp`) are rejected, and `--publish` in a non-interactive shell fails fast before any network call ## Risks - The `latest` dist-tag points at the deprecated `0.0.0` placeholder until the first stable release supersedes it. This window also existed under the old flow (which parked `latest` at a locally built version); internal consumers are unaffected because release version rewrites pin exact calver versions. - The registry poll caps at ~10 minutes. If propagation is slower than that, the helper prints the exact `npm deprecate ... --otp <code>` command to run manually once `npm view` resolves. - The helper no longer validates the name against the workspace release plan, so a typo within the `@paperclipai/` scope would reserve a wrong name. The dry-run preview shows the exact name before any publish. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It analyzed the existing bootstrap flow and the release scripts (`release-package-map.mjs`, `check-release-package-bootstrap.mjs`, `release.sh` dist-tag handling), wrote the replacement script and tests, updated the documentation, and ran the verification above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bd059a073d |
fix(workspaces): make managed runtimes reliable across restarts (#11740)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces need isolated databases, ports, and runtime services > - Concurrent workspaces could reuse ports or lose service ownership after a restart > - A markerless worktree also needed seed recovery, but normal markerless instances still needed to boot > - This pull request makes seed, port, and service ownership state explicit and recoverable > - It also checks live process and listener identity before it reclaims shared resources > - The benefit is reliable workspace startup, restart, adoption, and concurrent provisioning ## Linked Issues or Issue Description **What happened?** Managed workspaces could lose runtime service ownership after a control-plane restart. Concurrent worktrees could also reuse a port when their parent paths differed. A seed recovery change made every markerless instance resolve a worktree seed source, so normal instances without a source could not start. **Expected behavior** Paperclip must preserve healthy managed services across restarts. It must reserve unique ports across worktree parents. It must provision a registered markerless worktree, but it must skip seed work for a normal markerless instance. **Steps to reproduce** 1. Start two managed worktrees under different parent paths at the same time. 2. Restart the control plane while a managed service stays alive. 3. Start Paperclip with a config that has no seed markers and no registered worktree source. 4. Observe duplicate port selection, lost service adoption, or a seed-source startup error. **Paperclip version or commit** Current `master` plus the workspace runtime reliability changes in this pull request. **Deployment mode** Local development with managed execution workspaces and embedded Postgres. ## What Changed - Added a shared port registry with lease heartbeats, process identity checks, and live listener probes. - Reserved worktree ports across custom parent paths and repaired duplicate legacy assignments. - Preserved and adopted healthy managed services across control-plane restarts. - Reconciled guest bind modes and verified listener ownership before termination or reuse. - Provisioned registered markerless worktree databases and kept normal markerless instance startup as a no-op. - Added CLI, shared, server, and shell regression tests for seed, port, listener, restart, and adoption behavior. - Updated the worktree development documentation. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts --reporter=verbose` — 63 tests passed. - `pnpm exec vitest run packages/shared/src/worktree-port-registry.test.ts --reporter=verbose` — 5 tests passed. - Focused runtime Vitest set — 199 tests passed across 37 suites. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 10 tests passed. - `git diff --check` passed. ## Risks - Port reservation now depends on lease and process identity data. The fallback listener probe prevents early reclamation when process metadata is incomplete. - Runtime adoption is stricter about bind and owner identity. The tests cover healthy adoption, stale records, PID reuse, and unrelated listeners. - Markerless seed detection now separates registered worktrees from normal instances. The tests cover both paths. - There are no database schema migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e434829889 |
fix(release): draft-notes baseline survives candidate-cut stables (#11647)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release workflow drafts the upcoming stable's notes skeleton at
beta publish, ranging from the last stable to the beta's source commit
> - Stables cut from candidate branches leave their tag off master's
lineage, so the ancestor-only baseline rule skips them and falls back a
whole release too far
> - The first live draft did exactly that: it ranged from v2026.722.0
instead of the just-shipped v2026.817.0 and re-included everything
already released
> - This pull request replaces the baseline rule with a merge-base walk
over stable tags and adds a candidate-branch regression fixture
> - The benefit is a correct draft range for every stable lineage:
master-cut, candidate-cut, and promoted-older-source
## Linked Issues or Issue Description
**What happened?**
The first `draft_stable_notes` run (for beta `2026.818.0-beta.1`)
generated `releases/beta/v2026.818.0-beta.1.md` with the range
`v2026.722.0..664052f8e` — 311-commits-worth of already-shipped
v2026.817.0 content re-included.
**Expected behavior**
The draft ranges from the point the shipped stable's content diverges
from the beta source. For v2026.817.0 (cut from
`candidate/release-2026.817.0`) that is the promoted source commit
`8f7b8b3fd`, giving the 172 commits of genuinely-new work.
**Steps to reproduce**
Ship a stable from a candidate branch (tag lands off master's lineage),
then publish a beta from master and read the generated skeleton's range
line.
**Paperclip version or commit**
master at `664052f8e`.
Related (not duplicates): #11567 introduced the generator; its
ancestor-only rule was itself a review fix for the newest-by-version
rule, and this PR is the second iteration with the lineage case the
first fix missed.
## What Changed
- `scripts/draft-stable-notes.sh`: the range start comes from walking
stable tags newest-first and taking the first whose merge-base with the
beta source is a proper ancestor of the source. Candidate-cut stables
resolve to the promoted commit, master-lineage stables to the tag
itself, and tags containing the source are skipped (an older-source
promotion cannot produce an empty range). The skeleton header prints a
runnable short-sha range with the stable tag as a labeled baseline.
- `scripts/draft-stable-notes.test.mjs`: new regression fixture with the
stable tag on an unmerged candidate branch; the existing older-source
and fallback fixtures still pass unchanged.
## Verification
- `node --test scripts/draft-stable-notes.test.mjs` — 8 pass, including
the new fixture.
- Regenerated the live `2026.818.0-beta.1` draft against the real
repository: range start resolves to `v2026.817.0 (merge-base
|
||
|
|
664052f8ea |
feat(release): draft stable notes at beta publish, read them from master at promotion (#11567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system promotes builds canary → nightly → beta → stable, and stable releases publish a GitHub Release from `releases/vYYYY.MDD.P.md` > - The stable lane requires that notes file to exist inside the promoted source commit, but the file is named for the promotion date, which is unknown when the source commit is created > - A promoted beta can therefore never pass the notes check: every happy-path stable is forced through the candidate-branch fix path, with a soak-gate justification, for a notes-only change > - This pull request drafts the notes automatically when the beta is published and lets the stable promotion read them from `master` > - The benefit is a walkable stable happy path: the soak gate stays exact, notes get a real review window during the soak, and the justification path returns to its real purpose (cherry-picked fixes) ## Linked Issues or Issue Description **What existing behavior does this improve?** The stable promotion path in the release channel system (`release.yml`, `scripts/release.sh`). **Current behavior** `release.sh stable` requires `releases/vYYYY.MDD.P.md` in the checked-out source tree, and `publish_stable` checks out the exact promoted SHA. The soak gate requires a `beta/v*` tag to point at that same SHA. No commit can satisfy both for a promoted beta, so a stable promotion must cut a candidate branch with a notes-only commit and bypass the soak gate with a written justification. Release notes are also written at promotion time, under time pressure, with no review window. **Proposed behavior** When a beta publishes, a `draft_stable_notes` job generates a grouped notes skeleton at `releases/beta/v<beta-version>.md` and pushes it to a machine-owned branch; a human opens the PR and edits it during the 3-day soak. The stable preflight resolves notes before the `npm-stable` approval gate: source-tree notes first (the candidate fix path, unchanged), then the merged beta-keyed file on `master`; it fails early with the missing path named when neither exists. After the stable ships, a canonicalization job pushes a branch that moves the file to `releases/vYYYY.MDD.P.md`. Related (not duplicates): #11006 and #11008 introduced the nightly and beta lanes this builds on; older changelog PRs (for example #10669) authored notes manually at promotion time, which is the flow this replaces. **Reason and benefit** The happy path becomes: promote the exact soaked SHA, no justification, notes reviewed during the soak instead of written at the gate. The `releases/vYYYY.MDD.P.md` invariant still holds durably via the canonicalization PR. ## What Changed - `scripts/release.sh`: new `--notes-file PATH` (stable only) overrides where the pre-publish notes check looks, so notes can live outside the source checkout without dirtying the worktree. - `scripts/create-github-release.sh`: same `--notes-file` override for the GitHub Release body. - `scripts/draft-stable-notes.sh` (new): deterministic skeleton generator — commit subjects from the newest stable tag (falling back to the previous beta, then full history) to the beta's source commit, grouped into Features / Fixes / Other. - `.github/workflows/release.yml`: - `draft_stable_notes` job after `publish_beta`: runs the generator and force-pushes `release-notes/v<beta-version>`; the job summary links the compare page. It recreates the beta tag locally if the tag push was rejected (the known workflows-permission case), so drafting is not blocked on manual tag recovery. - `preflight_stable`: computes the target stable version (`release.sh stable --print-version`) and resolves the notes source (`source_tree` → `master_beta` → fail early / warn on dry run); new outputs. - `publish_stable`: materializes `master`-side notes into `RUNNER_TEMP` and passes `--notes-file` to both scripts; outputs the published stable version. - `canonicalize_stable_notes` job: pushes the `git mv` branch after a stable that used `master`-side notes. - `doc/RELEASING.md`, `doc/RELEASE-CHECKLIST.md`: document the drafted-notes flow, the preflight resolution order, and the canonicalization step; the LLM changelog flow now targets the draft branch during the soak. - `.agents/skills/release-changelog/SKILL.md`, `.agents/skills/release-changelog-discord-message/SKILL.md`: the notes-authoring skills now describe this flow — range ends at the beta source commit (not `HEAD`), the file is beta-keyed on the `release-notes/v<beta-version>` branch (seeded with `scripts/draft-stable-notes.sh` for betas that predate the automation), and the canonicalization link caveat is called out for announcements. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 6 tests, temp git-repo fixtures: grouping, stable-tag range, previous-beta and full-history fallbacks, default output path, malformed version, missing tag. - `node --test scripts/release-lib.test.mjs` — unchanged suite still green. - `bash -n` on both changed shell scripts; `release.yml` re-parsed as YAML. - `./scripts/release.sh stable --print-version` unchanged (prints the next stable version); `--notes-file` on a non-stable channel fails with a clear error. - Not exercised end-to-end: the new workflow jobs need a real beta publish to run. The first beta after merge is the live test; the draft job is additive and cannot affect the publish result (it runs after `publish_beta` completes). ## Risks - Low risk to publishing itself: `--notes-file` defaults preserve today's behavior everywhere; the draft and canonicalization jobs are additive and run after the publishes succeed. - The preflight now fails a real stable run when no notes are found. That is the intended fail-early behavior (it previously failed later, inside `publish_stable`, after the `npm-stable` approval). - `draft_stable_notes` force-pushes only the machine-owned `release-notes/v<beta-version>` branch; a beta re-cut regenerates it cleanly. - The stable version computed at preflight could differ from the published one if a run crosses UTC midnight between the two jobs; the materialized notes are passed by path, so the publish still succeeds, and the canonicalization job uses the actually-published version. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
49217aadf0 |
refactor: balance serialized server shards by recorded suite duration (#11528)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ec984502a |
fix(release): stop smoke_beta silently skipping on promote-mode betas (#11582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system re-smokes every published beta as post-publish verification (`smoke_beta`) > - The candidate-branch beta lane (#11209) added `verify_beta_candidate` to `publish_beta`'s needs; that job is skipped on every normal promote-mode beta > - `smoke_beta`'s condition has no status-check function, so GitHub attaches an implicit `success()` that evaluates the needs chain transitively — a skipped ancestor makes it false > - This pull request makes the condition explicit so promote-mode betas smoke again, and pins the shape in the workflow wiring test > - The benefit is that the post-publish beta gate actually runs instead of silently skipping ## Linked Issues or Issue Description **What happened?** Beta `2026.818.0-beta.0` (run 32082007439) published successfully, but its post-publish `smoke_beta` job was skipped. No configuration or input asked for that: the run was a plain `channel: beta` dispatch with `dry_run` at its default `false`, and the same expression `!inputs.dry_run` evaluated true inside `publish_beta`'s own steps (the Docker dispatch step ran). **Expected behavior** Every non-dry-run beta publish is followed by the release smoke suite against the exact published version, as documented in `doc/RELEASING.md` and `doc/RELEASE-CHECKLIST.md`. **Steps to reproduce** Dispatch `release.yml` with `channel: beta` promoting a nightly (promote mode). `verify_beta_candidate` is skipped by design; `publish_beta` runs through its explicit `!cancelled()` condition; `smoke_beta` then skips because its implicit `success()` sees the skipped ancestor in the transitive needs chain (actions/runner#2205 semantics). The beta published on 2026-08-11 predated #11209, so this never surfaced before. **Paperclip version or commit** master at `43ab441f0` (workflow file, current head). Related (not duplicates): #11209 introduced the candidate lane whose skipped job triggers this; #11208 covers the adjacent tag-push failure playbooks. ## What Changed - `smoke_beta`'s condition becomes `!cancelled() && needs.publish_beta.result == 'success' && !inputs.dry_run` — an explicit status-check function suppresses the implicit `success()`, and the result check keeps the dependency on a successful publish. - A comment above the job records why the explicit form is load-bearing. - `scripts/__tests__/release-verify-workflow.test.mjs` pins the new shape so the implicit form cannot silently return. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 8 pass, including the new assertion. - `release.yml` re-parsed as YAML. - The exact skip is visible on run 32082007439 (`smoke_beta: skipped` after `publish_beta: success`); the coverage gap for that beta was closed manually by dispatching `release-smoke.yml` with `paperclip_version: beta` (run 32084880767). - Not exercised end-to-end: the corrected condition needs the next real promote-mode beta to demonstrate; the expression change is minimal and the semantics are the documented actions/runner behavior. ## Risks - Low risk: condition-only change on one job plus a test. Dry runs still skip the smoke (`!inputs.dry_run` retained). Candidate-mode betas, where `verify_beta_candidate` actually runs, behave as before. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
4c349fe6b7 |
feat(runtime): managed Tailscale HTTPS lifecycle, durable runtime leases, and bounded control recovery (#11525)
<!-- Simplified Technical English (ASD-STE100). --> > **Stacked pull request.** This targets #11524. Merge #11524 first. Review only the second commit, `feat(runtime): managed Tailscale HTTPS lifecycle...`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services, so an agent's branch can be previewed while the agent works > - The previous pull request added the host broker, the shared contract, and the database columns, but no code used them > - A managed runtime can only be exposed over HTTPS if it holds a stable loopback port pair for the whole life of the service. The current control path cannot promise this: two controls can race the same execution workspace, a stranded control can stay `running` forever, and a start can adopt a port it does not own > - This pull request adds the HTTPS lifecycle and the control-path hardening that the lifecycle depends on > - The benefit is that a managed preview becomes reachable from another device, and a managed control now always reaches a terminal state ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, workspace operations, the execution workspace routes, and the workspace runtime UI. **Problem or motivation** A managed runtime service is reachable only on loopback, so a preview cannot be opened from a phone or a second computer. Exposing it safely needs an exclusively held port pair. Three existing gaps block that. Overlapping controls can race the same workspace. A control whose owner dies stays `running` and blocks the lane forever. Port allocation does not confirm that the process holding a port is the process Paperclip spawned. **Proposed solution** Add the exposure lifecycle on top of the broker from #11524: reserve before spawn, expose after readiness, validate the public URL, and remove on stop. In the same change, make managed controls mutually exclusive per workspace, give each control a durable issue-owned lease and a terminal state, and verify port ownership before use. **Alternatives considered** - Add HTTPS exposure without the control hardening. This was rejected because a raced or stranded control makes exposure point at the wrong process. - Guard the lane with an in-memory lock only. This was rejected because the lock does not survive a server restart, so the lane can be lost or double-claimed. - Trust the requested bind address. This was rejected because a checkout that predates managed HTTPS overwrites `PAPERCLIP_BIND` from its own `--bind` argument, and then binds the wildcard address. **Roadmap alignment** This completes the managed workspace runtime capability that already exists. It adds no new product surface beyond the HTTPS link. **Additional context** This is the second of three pull requests. The third adds central mediation of leased port pairs. ## What Changed Exposure lifecycle: - Add the server-side broker client and the exposure lifecycle manager. The manager reserves the mapping before spawn, exposes after backend readiness, validates the public URL, and removes the mapping on stop. - Default managed worktree runtimes to `tailscale_https`, read exposure intent from legacy `expose` blocks, and backfill runtimes that are still HTTP-only. - Verify listener ownership for the app port and its Vite HMR companion before the broker is asked to expose anything. An unrelated listener on either port fails the start closed. - Force the loopback bind through argv instead of environment hints. Leave a non-Paperclip service's `--bind` argument alone. - Probe loopback for readiness instead of the public URL, and give Vite HMR its own loopback-bound server in middleware mode. - Preserve operator-declared Serve mappings across the managed lifecycle, so cleanup never removes a mapping that Paperclip did not create. - Name which listener predicate denied an expose, so an operator can act on the message. Control-path hardening: - Make `start`, `stop`, `restart`, and job `run` mutually exclusive per execution workspace. An overlap gets `409 workspace_runtime_control_in_progress`, and authorization is still checked first. - Take a durable exclusivity lease on the execution workspace, owned by the controlling issue. A different issue gets `409 workspace_runtime_lease_conflict` before any operation is recorded. Board and operator actions bypass the lease. - Give every control a terminal state. Each control stamps its owning process and pid, heartbeats while it runs, and has a wall-clock ceiling. Recovery of a stranded control uses a compare-and-swap on `updated_at`, so a live owner is never stolen. - Bound readiness probes, verify allocated port ownership on POSIX and Windows, harden sibling port allocation, and reconcile desired runtimes on server startup. - Surface exposure state and bounded runtime errors in the workspace runtime UI. - Record the new behavior in `doc/DEVELOPING.md`. ## Verification Focused checks, all run on this branch: - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, exactly the count on `master`. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. - `pnpm --filter @paperclipai/ui typecheck` — clean. - Server suites, 177 tests pass across 9 files: `workspace-runtime.test.ts`, `workspace-runtime-leases.test.ts`, `workspace-runtime-control-recovery.test.ts`, `execution-workspace-runtime-control-conflict.test.ts`, `execution-workspace-runtime-lease-route.test.ts`, `workspace-operations-reconciliation.test.ts`, `workspace-runtime-start-terminality.test.ts`, `app-hmr-port.test.ts`, and `workspace-runtime-ready-comment.test.ts`. - Exposure unit suites, 77 tests pass: `src/services/runtime-exposure/` and `workspace-runtime-exposure-backfill.test.ts`. - UI: `WorkspaceRuntimeControls.test.tsx` and `WorkspaceServiceControlBar.test.tsx` — 34 tests pass. **One suite is red on the development host and is expected to be green in CI.** `server/src/services/workspace-runtime-exposure.test.ts` has 10 failures on the machine used to write this branch. The cause is host contamination, not the code. That machine already runs an HTTPS canary that holds ports 42000, 42001, 52000, and 52001 on a tailnet address. The suite allocates from the same range, so the new listener-ownership check correctly reports: ``` listener_ownership_mismatch — port 42000 is bound to 100.123.243.20, 127.0.0.1, fd7a:115c:a1e0:0:0:0:dd3a:f314 ... instead of loopback only ``` A CI runner has no listener on those ports, so the check sees loopback only and the suite passes. Please confirm this from the CI result on this pull request rather than from a local run on a host that already exposes a managed runtime. This is a real weakness of the current test fixture, and the third pull request in the series removes it by allocating the pair through a central mediator instead of a stubbed availability check. `workspace-runtime-https-live-exercise.test.ts` needs a live `tailscale` host and was not run locally. ## Risks - This is the behavior-bearing pull request of the three, so it carries the most risk. - Two new `409` responses appear on managed control routes. A caller that assumed a control always starts must handle a conflict. Board and operator actions are deliberately exempt, so an agent lease cannot lock an operator out. - Managed worktree runtimes now default to `tailscale_https`. If the host has no working broker, the start fails closed and reports the exposure failure instead of silently serving plain HTTP. This is intended, and it is the reason the failure message names the denying predicate. - Startup reconciliation touches persisted runtime rows. It is scoped to desired state and does not resurrect a service that never came up. - The lease has a 30-minute time to live and explicit release paths, so a crashed owner cannot hold a lane forever. - No migration runs in this pull request. The tables and columns land in #11524. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass, with the one host-contaminated suite explained above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d2665ff6b4 |
fix(ui): align the mobile task chat composer with the thread (#11296)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task thread is where people read work and guide agents. > - The mobile composer should use the same content width as the thread. > - The composer kept the desktop 80% width on mobile, so its edges did not align with the thread. > - Long assignee-aware placeholder text could also clip inside the mobile editor. > - Some extracted style tokens used legacy HSL wrappers around complete semantic colors, which made those declarations invalid. > - This pull request makes the composer full width on mobile, preserves the narrower desktop layout, wraps the placeholder, and repairs the invalid color compositions. > - The benefit is a stable mobile composer that aligns with the task thread and keeps its intended visual styles. ## Linked Issues or Issue Description Related work: Refs #11263. **What happened?** At mobile widths, the task chat composer used the same 80% width as the desktop composer. Its horizontal edges did not align with the full task thread. A long assignee-aware placeholder could clip on one line. The composer's extracted shadow also used a legacy `hsl(var(...))` wrapper around complete semantic color values, so the browser could reject the declaration. **Expected behavior** The composer must match the task thread width on mobile. It must stay narrower on larger screens. Long placeholder text must wrap inside the editor. Semantic color tokens must form valid shadows and gradients. **Steps to reproduce** 1. Open a task with the chat-style thread on a mobile viewport. 2. Compare the composer edges with the task thread edges. 3. Select an assignee whose placeholder text wraps to two lines. 4. Inspect the computed composer shadow and the extracted semantic color styles. **Paperclip version or commit** The change is based on `dc6fcd1ff1` from `master`. **Deployment mode** Local build from source. The behavior also applies to packaged web builds. ## What Changed - Made the task chat composer full width below the medium breakpoint and kept the 80% desktop width. - Matched the composer dock padding to the task thread padding. - Allowed long composer placeholders to wrap and reserved enough mobile editor height for two lines. - Replaced invalid legacy HSL wrappers around full semantic colors in extracted shadows, gradients, and approval styles. - Added a token gate that prevents legacy `hsl(var(--token))` wrappers from returning. - Added focused regression tests for responsive width, padding, placeholder wrapping, mobile height, and semantic shadow validity. ## Verification - `pnpm check:token-gates` — all four gates pass. - `pnpm --filter @paperclipai/ui exec vitest run src/components/TaskChatThread.test.tsx src/components/task-chat/TaskChatComposer.test.tsx src/components/task-chat/TaskChatComposerStyles.test.ts` — 37 tests pass. - `pnpm --filter @paperclipai/ui typecheck` — passes. - `pnpm --filter @paperclipai/ui build` — passes. The build prints existing CSS optimizer and bundle-size warnings. ## Risks - Low risk. The width change is limited to the mobile breakpoint. The desktop 80% layout remains in place. - The semantic token fixes can affect shadows and gradients that were previously invalid. The new gate prevents the invalid wrapper pattern from returning. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The context-window size is not exposed in this environment. The model used reasoning, repository tools, code execution, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a8d118a779 |
Prefer public base URL for generated invite links (#7619)
Fixes #7623 ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company invites are part of the access subsystem and must produce URLs that recipients can open from outside the host machine. > - Paperclip already has public/auth base URL configuration for deployments behind a public hostname, Tailscale, or a reverse proxy. > - Invite URL composition was still deriving its origin from the incoming request host, so loopback-bound servers emitted `http://127.0.0.1:3100/invite/...`. > - A loopback invite URL is not shareable with a remote human or agent, even when the token itself is valid. > - This pull request makes invite URL builders prefer the configured public base URL and keep the existing request-host fallback when it is unset. > - The benefit is that copied invite links use the reachable deployment origin without changing local-only behavior. ## Linked Issues or Issue Description Fixes #7623 No duplicate or related PRs/issues were found in a GitHub search for invite URL, loopback, public base URL, and `authPublicBaseUrl` terms. ## What Changed - Added base URL resolution in `server/src/routes/access.ts` that strips trailing slashes and prefers configured `authPublicBaseUrl` over the request-derived host. - Threaded `authPublicBaseUrl` through invite summary, invite onboarding manifest, onboarding text, access routes, `createApp`, and server startup wiring. - Added `server/src/__tests__/invite-url-public-base-url.test.ts` covering configured public-base precedence, unset fallback behavior, and trailing-slash normalization. - Registered the invite public-base URL test in the serialized Vitest server runner. ## Verification ```bash pnpm install --frozen-lockfile pnpm exec vitest run server/src/__tests__/invite-url-public-base-url.test.ts pnpm run test:run:serialized ``` Local results from the rebased PR branch: - `pnpm install --frozen-lockfile` exited 0. - Targeted invite URL test exited 0: 1 file, 3 tests passed. - Serialized server suite exited 0: 106 serialized suites completed; the new invite URL test passed inside that runner. Manual check after deployment: set `PAPERCLIP_AUTH_PUBLIC_BASE_URL` or equivalent public base URL config, create a company invite, and confirm the returned/copied invite URL uses that public origin instead of `127.0.0.1`. ## Risks Low risk. The new public base URL parameter is optional and falls back to existing request-derived behavior when unset. The main operational risk is misconfigured public base URL input; the implementation only trims trailing slashes and otherwise trusts the configured origin. ## Model Used - Original implementation: Anthropic `claude-sonnet-4-6`, 200k context, tool use and test execution. - Conflict repair and verification: OpenAI Codex GPT-5.5, coding agent with shell, git, GitHub CLI, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Coder (Claude) <coder-claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip Coder (Claude) <lad-agent@paperclip.ing> |
||
|
|
71e9d6bb0b |
feat(release): candidate-branch beta builds and the release checklist (#11209)
> Follow-up to #11208 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote artifacts along canary → nightly → beta → stable, with the happy path being promotion of an existing build > - When one or two targeted fixes are needed before a beta or stable, the only options today are waiting for the next nightly or absorbing a whole day of unrelated master changes > - The channel model was designed with an escape hatch for exactly this: short-lived candidate branches carrying only cherry-picked fixes > - This pull request implements candidate-branch beta builds with full verification, documents the stable fix path through the soak-justification gate, and adds the release captain's checklist > - The benefit is that a surgical fix can ship forward without either delay or blast radius, with its provenance recorded ## Linked Issues or Issue Description Refs #11008 — completes the fix-path half of the channel model introduced there. **Subsystem affected** Release automation: `scripts/release.sh`, `.github/workflows/release.yml`, `doc/RELEASING.md`, new `doc/RELEASE-CHECKLIST.md`, tests. **Problem or motivation** Beta promotion only accepts commits that already shipped as a nightly, and stable promotion expects a soaked beta. There is no supported way to ship one or two cherry-picked fixes between lanes: an urgent fix must wait for the nightly cycle or pull in every unrelated master change from the day. The original channel design called for candidate branches to cover this, and they were deferred from the initial implementation. **Proposed solution** Candidate-branch beta builds: cut `candidate/beta-<target>` from a nightly's source commit, cherry-pick the fixes, and dispatch `channel: beta` with the new `candidate_branch` input. Selection enforces the naming convention, rejects heads that already shipped as a beta or predate the candidate tooling, and records the cherry-picked commits in the job summary. Because candidate heads never went through a canary or nightly, publication is gated on a full `release-verify` run (promoted nightlies keep skipping re-verification). The stable fix path (`candidate/release-<target>` as `source_ref`) works through the existing soak gate: the justification requirement is the deliberate, recorded trade-off for shipping unsoaked bits, and is now documented as such. ## What Changed - `scripts/release.sh`: `--from-candidate` flag (beta only) waives the shipped-a-nightly requirement while keeping the duplicate-beta guard - `.github/workflows/release.yml`: `candidate_branch` dispatch input; candidate mode in `select_beta` (naming validation, duplicate and tooling-era rejection, cherry-pick recording); new `verify_beta_candidate` job gating candidate publishes on full verification - `doc/RELEASING.md`: beta fix-path and stable fix-path sections - `doc/RELEASE-CHECKLIST.md` (new): the release captain's checklist for all four lanes as built - Tests: dry-run fixture coverage for `--from-candidate` (waives the nightly guard, keeps the duplicate guard, rejected outside beta) and wiring tests for candidate validation plus the verification gate ## Verification - `node --test` on the four affected suites: 42 pass in total (17 + 25 across the two runs), including the 5 new tests - `bash -n` on `release.sh`; YAML parse of the workflow - After merge: exercise the path end to end the first time a real cherry-picked beta is needed — dispatch with a `candidate/beta-*` branch and confirm the summary records the picks and verification runs ## Risks - Candidate builds bypass the smoke-tested-nightly provenance by design; the compensating controls are full verification before publish, the post-publish beta smoke, the human `npm-beta` gate, and recorded cherry-picks - The stable fix path rides the existing justification mechanism rather than adding a second bypass — one recorded escape hatch, not two ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2da6a248c3 |
fix(release): surface recovery commands when a lane tag push is rejected (#11208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's promotion lanes publish to npm, then push a lane tag and dispatch the Docker image build at that tag > - The first nightly of the beta-tooling merge published to npm and then died at the tag push: GITHUB_TOKEN may not create refs pointing at workflow-modifying commits from dispatch or scheduled runs > - The failure was a bare `remote rejected` with no guidance, leaving the release half-finished (npm live, no tag, no images) until an operator reverse-engineered the recovery > - This pull request makes every lane's tag push degrade into exact recovery instructions in the job summary > - The benefit is that a rare platform-permission rejection becomes a two-minute runbook operation instead of a forensic exercise ## Linked Issues or Issue Description Refs #11008 — the incident occurred promoting that change's own merge commit, the first workflow-modifying commit to flow through the lanes it introduced. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31445344811 published `2026.811.0-nightly.0` to npm, then failed pushing `nightly/v2026.811.0-nightly.0`: `refusing to allow a GitHub App to create or update workflow .github/workflows/release.yml without workflows permission`. The tagged commit modifies workflow files, and GITHUB_TOKEN may not create refs pointing at such commits from dispatch or scheduled runs (push-event runs are exempt, which is why the canary tag on the same commit succeeded). The job failed with no explanation and the Docker dispatch never ran. **Proposed solution** Wrap the nightly, beta, and stable tag pushes: on rejection, write the exact recovery commands into the job summary — create and push the tag with maintainer credentials, dispatch `docker.yml` at the tag, and for stable also run `create-github-release.sh` — then fail the job. Document the cause and recovery in the failure playbooks and pin the three recovery blocks with a wiring test. ## What Changed - `.github/workflows/release.yml`: recovery-summary wrappers on the nightly, beta, and stable tag-push steps - `doc/RELEASING.md`: failure-playbook entry for the workflows-permission rejection - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting all three lanes carry the recovery summary ## Verification - Wiring tests: 6 pass; YAML parse of the workflow - The recovery commands are exactly the ones used to resolve the real incident (tag push + `docker.yml` dispatch for `nightly/v2026.811.0-nightly.0`) ## Risks - Low. The happy path is unchanged (a successful push skips the wrapper); the failure path trades a bare error for actionable output and still fails the job, since the release state is genuinely incomplete ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6601014898 |
fix(release): reject promotion sources that predate their channel tooling (#11197)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem promotes builds along canary → nightly → beta → stable, and each publish job checks out the promotion's source commit and runs that tree's release tooling > - The first beta dispatch failed with `unexpected argument: beta`: the selected nightly's source predated the beta channel, so its `release.sh` did not know the argument > - The failure was clean (argument parsing, nothing published) but cryptic, and the same trap waits for any promotion of a source older than its target channel's tooling > - This pull request makes the selection jobs reject such sources with an actionable error and documents the property > - The benefit is that a bootstrapping or old-source promotion fails in seconds with instructions, instead of mid-publish with a parser error ## Linked Issues or Issue Description Refs #11008 — the guard hardens the beta promotion flow introduced there, after its first dispatch surfaced the gap described below. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31444045044 (first beta dispatch) failed in `publish_beta` with `unexpected argument: beta`. Promotions deliberately build from the pinned source commit, which means they also run that commit's `scripts/release.sh` — and a source that predates the target channel's introduction cannot publish it. Nothing guards this today; the error surfaces deep in the publish job with no explanation. **Proposed solution** Guard at selection time: `select_nightly` requires the source canary's `release.sh` to know the nightly channel, and `select_beta` requires the source nightly's `release.sh` to know the beta channel. Each guard literally matches the channel case arm and fails closed with a clear message naming the remedy (promote a newer source). Document the tooling-era property in `RELEASING.md` and pin the guards with a wiring test. ## What Changed - `.github/workflows/release.yml`: tooling-era guards in `select_nightly` and `select_beta` - `doc/RELEASING.md`: documents that promotions run the source commit's release tooling - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test pinning both guards ## Verification - Wiring tests: 5 pass - Guard expressions exercised against real commits: accepts the beta-capable merge commit of the beta-channel change, rejects a pre-beta commit - YAML parse of the workflow - After merge: the next beta dispatch selects a beta-capable nightly and passes the guard ## Risks - Low. Selection-time check only; the guards match the channel case arm literally and fail closed (with the same actionable message) if that line is ever reformatted ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f94f6003c6 |
fix(release-smoke): pin the smoke container to the lan bind preset (#11189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly on the release smoke suite, which boots the published artifact in a Docker container and drives real onboarding > - The gate kept failing even after the readiness budget fix (#11187), and the new container-log dump revealed the server was healthy but listening on 127.0.0.1 inside the container, unreachable through Docker's port mapping > - `onboard --yes` without an explicit `--bind` prefers trusted-local quickstart defaults: it writes a loopback bind into the instance config and ignores the deployment env vars the harness passes, and that config outranks `HOST` at runtime > - This pull request pins the smoke container to the `lan` bind preset and adds a wiring test for it > - The benefit is a working nightly gate, verified end to end against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `docker/Dockerfile.onboard-smoke`, `scripts/__tests__/release-verify-workflow.test.mjs`. **Problem or motivation** Nightly run 31428558684 failed in smoke with the server unreachable at the mapped port for the full 420 second budget. The container logs (captured thanks to #11187) show a fully booted server with `Bind loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the deployment env vars, but `onboard --yes` without `--bind` deliberately prefers trusted-local defaults, writes `bind: loopback` into the instance config, and the config outranks `HOST` at runtime. A loopback listener inside a container is invisible to the port mapping, so the health check can never pass. This behavior predates the current stable, so the harness was silently broken against every recent version — it only surfaced now because the nightly lane is the suite's first CI consumer. **Proposed solution** Pass `--bind lan` in the smoke container command (the flag is supported by `latest` and canary alike; it selects the all-interfaces preset and keeps the env-driven authenticated deployment), and pin the flag with a wiring test so it cannot regress silently. ## What Changed - `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard --yes --bind lan --data-dir ...`, with a comment explaining why the flag is load-bearing - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting the smoke Dockerfile pins a non-loopback bind preset ## Verification - Full local harness run against the real nightly candidate `2026.810.0-canary.1`: container healthy, bind banner shows `lan (0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap invite accepted, board session verified), `/api/health` returns `bootstrapStatus: ready` - `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4 pass - After merge: dispatch `release.yml` with `channel: nightly` to run the gate end to end in CI ## Risks - Low. The change only affects the smoke container. `--bind lan` inside a container exposes the port to the container network only; reachability from outside still goes through Docker's explicit port mapping ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI log forensics, upstream source tracing, local Docker reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
30f6999cbe |
fix(release-smoke): configurable readiness timeout and diagnostics for slow containers (#11187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane (#11006) gates every nightly publish on the release smoke suite, which boots the published artifact in a Docker container > - The suite's first CI execution failed at the health readiness check: the harness hard-codes a 90 second budget, but a CI container cold-installs paperclipai from npm and initializes embedded postgres with no warm caches > - When the timeout expired with the container still running, the harness printed no container logs, so the failure gave no diagnostics > - This pull request makes the readiness budget configurable, raises it for CI, and dumps container logs on timeout > - The benefit is that the nightly gate measures the artifact, not the runner's cold caches, and a red smoke run is diagnosable from its logs ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `scripts/docker-onboard-smoke.sh`, `.github/workflows/release-smoke.yml`. **Problem or motivation** Run 31426044332 (first forced nightly after #11006) failed in `smoke_nightly` with `server did not become ready at http://localhost:3232/api/health` after exactly 90 seconds. The harness's readiness window is hard-coded to 90 attempts at 1 second. Locally that works because the npm cache is warm; in CI the container downloads the full package set and embedded postgres first. The timeout path also printed no container logs when the container was still running, so there was no way to see how far boot had progressed. **Proposed solution** Make the readiness budget an environment variable (`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use), set it to 420 in the CI workflow, and dump the last 150 container log lines when the readiness check times out on a still-running container. ## What Changed - `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env var (default 90) replaces the hard-coded readiness budget; timeout with a still-running container now prints the tail of `docker logs` - `.github/workflows/release-smoke.yml`: sets `SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs ## Verification - `bash -n` on the harness and YAML parse of the workflow - The real proof is the next `channel: nightly` dispatch of `release.yml`, which re-runs this suite in CI with the new budget ## Risks - Low. The local default is unchanged; CI runs simply wait longer before declaring failure, and a genuinely broken artifact still fails (with logs now) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. Diagnosis from CI run logs; patch model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f9173782cd |
feat(release): add smoke-gated nightly channel and lane-separated Docker tags (#11006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem publishes the `paperclipai` npm package set and the Docker images on two lanes: canary on every master push, and stable on manual promotion > - There is no middle ground between those lanes. Users must track every merge or wait weeks for a stable. Docker `:latest` also tracks master, so Docker users have no stable image at all > - A calm prerelease lane needs to exist, and it must never ship a build that failed its checks > - This pull request adds the nightly channel: a scheduled job that selects the newest master commit with a green canary publish, runs the full release smoke suite against that exact published canary, and only then republishes it as the nightly. It also separates Docker tags by lane, so `:latest` finally means stable > - The benefit is that users can follow prereleases at a nightly cadence with a smoke-tested guarantee, and Docker users get real `:canary`, `:nightly`, and stable image tags ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** The project publishes only `canary` (every master push) and `latest` (manual stable). Users who want prereleases without per-merge churn have no option. Docker has a second problem: master builds overwrite `:latest`, and CI-published stables never produced Docker images, because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger in `docker.yml`. No stable-versioned image exists in ghcr today. **Proposed solution** Add a `nightly` channel. A scheduled job selects the newest canary-tagged master commit, smoke-tests that exact published canary, and republishes the same commit as `YYYY.MDD.P-nightly.N` under the `nightly` dist-tag. Separate Docker tags by lane (`:canary` for master, `:nightly` for nightly tags, `:latest` plus version tags for stable tags only), and have the release jobs dispatch `docker.yml` at the new tag so lane images actually build. **Alternatives considered** Moving the `nightly` dist-tag to the existing canary version without a republish. Rejected: the version string would say `canary` while the user is on nightly, which breaks at-a-glance lane identification in bug reports and `--version` output. ## What Changed - `scripts/release-lib.sh`: channel-parameterized `next_prerelease_version` and `prerelease_tag_name` helpers (canary helpers delegate to them), a `require_channel_tag_at_head` guard, and the no-provenance retry for Sigstore transparency-log duplicates now covers the `nightly` dist-tag as well as `canary` - `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry a `canary/v*` tag, publishes the full public package set as `YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source commit `nightly/vYYYY.MDD.P-nightly.N` - `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) — select candidate, smoke it via `release-smoke.yml`, publish on green under the existing `npm-canary` environment, push the tag, dispatch `docker.yml`. New `channel` dispatch input (default `stable`, so existing stable dispatches are unchanged) with `nightly_source_version` and `dry_run` support for forced runs. The stable path now also dispatches `docker.yml` at the new `v*` tag - `.github/workflows/docker.yml`: lane tag mapping for both image jobs — master pushes publish `:canary` and no longer move `:latest`; `nightly/v*` tags publish `:nightly`; only stable `v*` tags publish `:latest` and the versioned tags. New `workflow_dispatch` trigger for the release-job dispatches. Build-version stamping uses the exact nightly version on nightly tag builds - `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch choice list - `doc/CHANNELS.md` (new): user-facing guide to the channels - `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping table, and a nightly failure playbook - `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses `npm-canary` and needs no npm trusted-publisher changes - Tests: channel-parameterized version helper coverage in `scripts/release-registry-versions.test.mjs`, and nightly flow coverage (publish identity, notes not required, canary-tag guard) in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the release script suites: 68 pass, including 6 new tests. The only failure, `acpx-patch-packaging.test.mjs`, needs installed `node_modules` and fails identically on a pristine checkout of master in the same environment - `bash -n` on both shell scripts and YAML parse of all three workflows - Live fail-path check: `./scripts/release.sh nightly --print-version` from a master tip with no canary tag fails with `HEAD has no canary/v* tag` - Live success-path check: the same command from the `canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0` - Live selection check: the candidate-selection shell logic run against the real repository selects the commit of `canary/v2026.806.0-canary.7`, which matches the current npm `canary` dist-tag exactly - After merge: dispatch `release.yml` with `channel: nightly` and `dry_run: true` to preview, then a real forced run to validate end to end before the first scheduled run ## Risks - Docker `:latest` changes meaning from "latest master build" to "latest stable release". This is deliberate and will be announced. Users who want the old behavior pull `:canary`. Until the first stable release after this change, `:latest` stays at its current (master-built) image - The nightly is a rebuild of the same source commit, not the byte-identical canary artifact that was smoked. The lockfile pins dependencies, and the publish path's registry-visibility and clean-prefix install gates still run on the nightly artifacts - All npm publishing must stay inside `release.yml` because npm trusted publishing pins that workflow file per package. The nightly jobs were added to `release.yml` for exactly that reason; this constraint is now documented in `RELEASING.md` - The stable-lane Docker dispatch fails gracefully (a warning with manual instructions) when the source ref predates `docker.yml`'s `workflow_dispatch` trigger ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
03cfad7ceb |
feat(apps): connect Notion through MCP OAuth (#11009)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - The Apps gallery lists Notion, but the server required manually configured OAuth credentials. > - Notion's hosted MCP server supports OAuth discovery and dynamic client registration. > - Notion also requires HTTPS or a loopback HTTP redirect URI. > - This pull request adds a direct Notion MCP OAuth path with PKCE and reusable dynamic clients. > - It also adds the current Apps UI states for connect and reauthorization. > - The benefit is a secure Notion connection with no manual client credential setup. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps gallery, Apps connect route, OAuth token lifecycle, and managed MCP gateway. **Subsystem affected** `server/`, `packages/shared/`, `scripts/`, and `ui/`. **Current behavior** The Notion gallery cards are disabled. The server uses the classic Notion OAuth endpoints and requires operator-supplied client credentials. It does not register an OAuth client from provider metadata. Concurrent refreshes can also replay a rotating refresh token. **Proposed behavior** Enable the Notion Apps flow. Discover OAuth metadata from `https://mcp.notion.com/mcp`. Register and reuse a public RFC 7591 client with PKCE. Require HTTPS or loopback HTTP callbacks. Serialize refreshes, store each rotated refresh token before the new access token can be used, and show a reconnect state for `invalid_grant`. **Reason and benefit** Operators can connect the built-in Notion MCP app without creating or copying OAuth credentials. Paperclip keeps dynamic clients and rotating tokens in the company secret store. **Breaking changes** None. Explicit environment client credentials still take priority. Existing Slack and Linear OAuth endpoint hints remain unchanged. Other OAuth apps remain disabled unless they are allowlisted. **Additional context** PR #10910 is a related, broader Connections v3 wizard replacement. This PR is the focused current Apps flow. The MCP Tool Gateway and Connected Apps items in `ROADMAP.md` cover this planned capability. ## What Changed - Classify all 20 reviewed Notion MCP tools with provider-scoped read and write defaults. - Require approval for selected Notion mutations, including move, duplicate, and convert actions that generic verb matching missed. - Preserve company-scoped connection and catalog resolution for Notion profiles and policies. - Add RFC 7591 dynamic client registration with `token_endpoint_auth_method=none` and mandatory PKCE. - Store the dynamic client ID on the connection and store any returned client secret in the company secret store. - Reuse the registered client for later connects and keep explicit environment credentials as the first choice. - Discover protected-resource and authorization-server metadata from the Notion MCP endpoint. - Add `redirectConstraints: "https-or-loopback-http"` to the generated Notion app definition and shared contract. - Reject non-loopback plain HTTP callbacks before network access with a TLS setup error. - Serialize client registration and token refresh operations within the server process. - Store a rotated refresh token before publishing the refreshed access token. - Treat `invalid_grant` as terminal and move the connection to a clear reauthorization state. - Add focused coverage for registration reuse, callback constraints, refresh rotation, and terminal grants. - Enable the Notion Apps route and add connect, redirect, success, error, and reconnect UI states. - Keep non-allowlisted OAuth apps blocked and cover the UI policy with regression tests. ## Verification - The focused Notion policy integration test passed with embedded PostgreSQL. - The focused 20-tool classification test passed. - The server typecheck passed on the governance head. - `pnpm -r typecheck` passed on the rebased head. - `pnpm --filter @paperclipai/server typecheck` passed after the security follow-up. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t \u0027DCR|refresh tokens|invalid_grant|abandoned lease\u0027` passed 10 focused security tests. - `pnpm build` passed on the rebased head. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t 'OAuth|oauth'` passed 14 tests. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts` passed 5 tests. - The complete server group passed 3,686 tests with 4 skipped. - The complete UI group passed 3,656 tests. - The full local runner found one environment-only CLI failure because this agent runtime injects static AWS credentials into a test that expects `AWS_PROFILE` only. `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` passed all 8 tests. - The prior UI verification passed 55 focused tests, `pnpm check:token-gates`, the Storybook build, and review of six 1440 x 1000 screenshots. - OAuth request sequence: protected-resource metadata `GET https://mcp.notion.com/.well-known/oauth-protected-resource/mcp`; authorization metadata `GET https://mcp.notion.com/.well-known/oauth-authorization-server`; dynamic registration `POST https://mcp.notion.com/register`; authorization `GET https://mcp.notion.com/authorize`; token exchange and refresh `POST https://mcp.notion.com/token`; MCP traffic `POST https://mcp.notion.com/mcp`. - The live metadata and registration probe confirmed that Notion accepts HTTPS and loopback HTTP redirects. It rejects a plain HTTP private hostname. - A later QA task owns the full browser consent and managed gateway tool-list dry run against a configured HTTPS deployment. ## Risks - Notion can add tools. Unrecognized names use the generic classifier, and new or changed risky tools stay quarantined after connection activation. - A deployment that uses a private non-loopback hostname must configure HTTPS before it can connect Notion. - Dynamic registration creates a provider-side client. Paperclip reuses it because registration does not provide a standard delete operation. - Refresh coordination uses a database CAS lease across service instances. An unclean crash leaves an uncertain lease and requires reconnect instead of risking refresh-token replay. - The current Apps surface overlaps with PR #10910. Merge order can require a small conflict resolution if that PR lands first. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex on a GPT-5 runtime. The exact deployment ID and context window are not exposed. The runtime used reasoning, repository tools, code execution, and network tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
00a24d7e8f |
ci: split general-server tests into five shards with refreshed durations (#10925)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The PR workflow runs the server vitest suite across sharded runners because the suite is pinned to one worker. > - In successful PR run 30930345729 (2026-08-04), shard `server (3/4)` took 311 seconds of wall time and was the slowest check in the run. > - The suite has grown to about 946 seconds of serial vitest time, but the duration manifest was last sampled on 2026-08-01 at about 882 seconds. > - This pull request refreshes the per-suite duration manifest from that run's logs and splits the lane into five shards. > - The benefit is a shorter PR critical path: each shard carries about 196 seconds of suite time, level with the other lanes. ## Linked Issues or Issue Description Refs #10663 (previous split of this lane into four shards). Related: #10923 splits the separate serialized-suites lane into five shards. Both PRs touch `.github/workflows/pr.yml` in different matrix blocks; whichever merges second needs a trivial rebase. **What existing behavior does this improve?** The `general-server` vitest lane runs in four shards with a duration manifest sampled on 2026-08-01. **Current behavior** In PR run 30930345729, shard 3/4 ran for 311 seconds (273 seconds in the test step) and was the longest check in the run. The suite now totals about 946 seconds of serial vitest time. **Proposed behavior** Run the same suite set in five shards, balanced with a per-suite duration manifest refreshed from that run's shard logs (279 suites measured by diffing consecutive completion timestamps). **Reason and benefit** The refreshed LPT partition balances at about 196 seconds of suite time per shard (about 240 seconds per job), level with the other PR lanes. No test coverage is lost. **Breaking changes** None. The change only alters the CI partition size and the duration manifest. ## What Changed - Bump the `general-server` shard matrix in `.github/workflows/pr.yml` from four to five shards. - Refresh `scripts/general-server-shard-durations.json` from the 2026-08-04 run's shard logs. - Update `SHARD_COUNT` in `scripts/__tests__/run-vitest-stable-shard.test.mjs` to five. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass, including the complete non-overlapping partition proof and the duration-balance check. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass. - `node --test scripts/__tests__/e2e-shard.test.mjs` — 7/7 pass. - A 5-way dry-run partition covers all suites exactly once with equal projected weights. ## Risks - Low risk. The change only alters CI partition size and duration weights; the suite set is unchanged. - One more runner is used per PR run for this lane. - Stale duration weights degrade gracefully: suites missing from the manifest get the median weight. ## Model Used - Claude (Anthropic), Claude Code CLI, model ID `claude-fable-5`, extended thinking with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (workflow comments explain the new shard math) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b6e58019f2 |
ci: split serialized tests into five shards (#10923)
## Thinking Path > - Paperclip uses CI to keep control-plane changes safe and mergeable. > - The PR workflow splits serialized server tests across isolated runners. > - A recent successful run spent 305 seconds in serialized shard 2/4. > - That job was the slowest check in the run. > - The four shards reported about 739 seconds of Vitest suite time. > - This pull request adds a fifth serialized shard and keeps release verification aligned. > - The benefit is a shorter PR critical path with no loss of test coverage. ## Linked Issues or Issue Description **What existing behavior does this improve?** The PR and release verification workflows run serialized server tests in four shards. **Current behavior** Successful PR run 30876682788 spent 305 seconds in `Verify serialized server suites (2/4)`. The test step used 256 seconds and made this job the slowest check. **Proposed behavior** Run the same serialized suite set in five complete and non-overlapping shards. **Reason and benefit** The measured suites reported about 739 seconds of total Vitest time. Five runners reduce the expected average suite time from about 185 seconds to about 148 seconds before setup overhead. **Breaking changes** None. The change only alters CI partition size. ## What Changed - Split serialized server tests into five shards in the PR workflow. - Apply the same five-shard layout to release verification. - Add a partition test that proves complete and non-overlapping serialized coverage. - Update release workflow coverage tests for five shards. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` - `git diff --check` ## Risks - Low risk. CI uses one additional runner for the serialized lane. - Round-robin partition weights can still vary as suite timings change. > This change does not overlap with planned core work in `ROADMAP.md`. Related PR #10663 optimized the separate general-server lane. ## Model Used - OpenAI Codex, GPT-5, agentic coding with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Devin Foley <139239+devinfoley@users.noreply.github.com> |
||
|
|
c54936e2e9 |
fix(openclaw-gateway): use per-agent claimedApiKeyPath in wake text (#4668)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `openclaw-gateway` adapter wakes a remote agent over a WebSocket gateway. It sends a wake prompt. That prompt tells the agent which environment variables to set and which file holds its Paperclip API key. > - Each agent stores its claimed key in its own JSON file. The adapter already exposes a `claimedApiKeyPath` config field for this. The field is documented in `src/index.ts`. It also has an input in the agent settings UI. > - `buildWakeText` ignored that field. It hardcoded the shared default path into the wake prompt text. > - Every agent therefore read the same key file at wake time. Agents authenticated as the wrong identity. The first API call failed. > - This pull request passes `ctx.config.claimedApiKeyPath` into `buildWakeText`. It uses the existing `resolveClaimedApiKeyPath` helper. That helper falls back to the documented default. > - The benefit is that each agent reads its own claimed-key file. Each agent authenticates as itself. ## Linked Issues or Issue Description Fixes #10071 Fixes #4976 Fixes #3098 Fixes #8076 These four open issues report the same defect. Earlier duplicates are already closed: Refs #2561, Refs #2592, Refs #930. Related pull requests that address the same root problem (duplicate search): - #3396 — same core change, no tests - #3370 — heavier approach, injects `PAPERCLIP_CLAIMED_API_KEY_PATH` into the wake env and adds server onboarding defaults - #5970 — renames the config field to `paperclipApiKeyPath` - #8072 — same core change, bundled with an unrelated protocol-version change - #784 — adds shell quoting and preflight instructions - #3296 — bundled with an unrelated Claude hello-probe fix ## What Changed - `packages/adapters/openclaw-gateway/src/server/execute.ts` - `buildWakeText` now accepts `claimedApiKeyPath` as a parameter. It no longer hardcodes the path. - The `execute` call site passes `resolveClaimedApiKeyPath(ctx.config.claimedApiKeyPath)`. That helper returns the documented default `~/.openclaw/workspace/paperclip-claimed-api-key.json` when the agent sets no override. - `resolveClaimedApiKeyPath` is now exported so tests can call it. - `packages/adapters/openclaw-gateway/src/server/execute.test.ts` — adds `resolveClaimedApiKeyPath` cases: a configured value, an empty string, a whitespace-only string, `undefined`, `null`, and non-string input. - `packages/adapters/openclaw-gateway/vitest.config.ts` (new) — package-level vitest config. It matches the config used by sibling adapters such as `opencode-local`. - `vitest.config.ts` (root) — adds the adapter to the workspace project list. - `scripts/run-vitest-stable.mjs` — adds `@paperclipai/adapter-openclaw-gateway` to `nonServerProjects`. **Maintainer-added during rebase.** The CI test lanes do not run a bare `vitest`. They call `run-vitest-stable.mjs`, which invokes vitest with an explicit `--project` allowlist. Without this entry the CI lanes skip this package, and the root project-list entry alone has no effect on CI. ## Verification Run the package suite directly: ``` pnpm install --frozen-lockfile pnpm exec vitest run --project @paperclipai/adapter-openclaw-gateway ``` The suite covers `resolveSessionKey`, `buildAgentParams`, and the new `resolveClaimedApiKeyPath` cases. The first two already existed in this file but never executed in CI before this change. Typecheck the package: ``` pnpm --filter @paperclipai/adapter-openclaw-gateway typecheck ``` Behavioural check, which no automated test covers: 1. Set `claimedApiKeyPath` to a per-agent value such as `~/.openclaw/workspace/paperclip-keys/<agent>.json` in the agent's gateway adapter settings. 2. Trigger a wake for that agent. 3. Confirm the rendered wake text names that file. It must not name the shared default. Maintainer note: this branch was rebased onto current `master` by a maintainer. The original branch was two months stale. Only two conflicts occurred, both additive: the import line and the tail of `execute.test.ts`, and the project list in the root `vitest.config.ts`. The `execute.ts` change applied without conflict. CI and Greptile re-run against the rebased head. ## Risks - Low for existing deployments. `resolveClaimedApiKeyPath` preserves the default path exactly. Any agent that never set `claimedApiKeyPath` receives the same wake text as before. - The behaviour changes only for agents that already set a per-agent path. Those agents previously received the wrong instruction. They now receive the correct one. - No database, schema, or API surface changes. - CI now runs this package's test file for the first time. That file includes the pre-existing `resolveSessionKey` and `buildAgentParams` tests, which were never executed before. - Five other adapters (`cursor-cloud`, `cursor-local`, `gemini-local`, `grok-local`, `pi-local`) sit in the root project list but remain absent from the CI allowlist. This pull request does not change them. That gap is tracked separately. ## Model Used - Contributor's change: Anthropic Claude, model ID `claude-opus-4-7`, approximately 200K context, extended thinking. Used for triage, patch authoring, and the original description. - Rebase, the `run-vitest-stable.mjs` entry, and this description: Anthropic Claude, model ID `claude-opus-5`, tool use enabled. Run by a Paperclip maintainer. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details — the branch name carries an internal ticket id. A fork branch cannot be renamed without opening a new pull request, so this is left as-is. The internal reference has been removed from the description. - [ ] I have run tests locally and they pass — the contributor verified the pre-rebase branch. The rebased head is verified by CI on this pull request. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — `claimedApiKeyPath` is already documented in `src/index.ts` and exposed in the agent settings UI, so no documentation change is needed - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — pending the post-rebase run - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending re-review of the rebased head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Pieter (CTO) <pieter@openclaw.local> Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com> |
||
|
|
f0b06d2de9 |
feat(claude): environment-aware test-environment probe and claude-local CI coverage (#10833)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The claude_local adapter runs Claude Code on sandbox execution targets, and operators verify an agent's configuration with the test-environment probe before running it > - Real runs merge the selected environment's env vars (secret refs included) under the agent's adapter config env, but the probe built its config from the adapter config alone — so environment-level auth worked in runs while the Test button reported missing auth, and a dropped secret binding passed silently > - The claude env-test hints also did not recognize `CLAUDE_CODE_OAUTH_TOKEN` even though the CLI accepts it, and a hello probe that hit the subscription usage limit reported a hard failure although authentication worked > - Separately, the claude-local package test suites were absent from the CI project list, so two suites drifted broken without notice > - This pull request makes the probe resolve the same layered env as a real run, adds the missing auth hint, classifies usage-limit probe results as a warning, repairs the drifted suites, and turns the claude-local project on in CI > - The benefit is a Test button that tells the truth about environment-level configuration, and a test suite that actually gates the claude-local adapter ## Linked Issues or Issue Description No public issue exists; related open PRs: Refs #9488 (recognizes CLAUDE_CODE_OAUTH_TOKEN in environment checks — overlaps with the auth-hint portion of this PR via a differently named check; it does not cover the environment-envVars probe merge, the usage-limit classification, or the CI coverage), Refs #9933 (live credential validation in environment checks — complementary, no file-level conflict with the route change). The underlying problem, following the enhancement template: **Current behavior** The test-environment route builds the probe config from the agent's adapterConfig only. Real runs merge the selected environment's envVars under the agent env, so environment-level env vars (including auth such as `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN` bound as environment secrets) work in runs while "Test environment" cannot see them, and a missing secret binding passes silently. The claude env-test hints do not recognize `CLAUDE_CODE_OAUTH_TOKEN`. A hello probe that hits the subscription usage limit reports a hard `claude_hello_probe_failed`. The claude-local package test suites do not run in CI, and two of them are stale. **Proposed behavior** The probe resolves the selected environment's envVars (environment-consumer secret bindings included) and merges them under the agent config env with the run-path precedence; missing bindings surface as an explicit error check that fails the test. The env-test emits a `claude_oauth_token_configured` info check when that variable is set. Usage-limit probe results classify as a `claude_hello_probe_usage_limited` warning because auth works and only the usage window is spent. The claude-local suites run in CI. Docs state the resulting facts. **Reason and benefit** The Test button should tell the truth: it previously contradicted run behavior for environment-level configuration and hid broken secret bindings. Enabling the package suites in CI prevents further silent drift — two suites were already broken on master without anyone noticing. **Breaking changes** None. Runs are unchanged. The probe route only adds env layers and checks; setups without environment envVars behave exactly as before. ## What Changed - `server/src/routes/agents.ts`: the test-environment route resolves the selected environment's envVars (forbidden keys stripped, environment-consumer secret context) and merges them under the agent adapterConfig env, mirroring `resolveExecutionRunAdapterConfig` precedence. Missing secret bindings are skipped, reported as an `environment_env_binding_missing` error check, and fail the test — matching the `ConfigurationIncompleteFailure` a real dispatch would raise. - `packages/adapters/claude-local/src/server/test.ts`: new `claude_oauth_token_configured` info hint between the API-key warning and the subscription fallback; hello-probe classification gains a `claude_hello_probe_usage_limited` warning for provider-quota results (previously a hard `claude_hello_probe_failed`). - `scripts/run-vitest-stable.mjs`: add `@paperclipai/adapter-claude-local` to `nonServerProjects` so CI runs the package suites. - `packages/adapters/claude-local/src/server/execute.remote.test.ts`: assert both runtime asset syncs (skills and mcp-config); the suite predated the mcp-config asset. - `packages/adapters/claude-local/src/server/test.probe.test.ts`: usage-limit fixture now expects the usage-limited warning; new fixture covers the genuine transient path (529 overloaded); new tests cover the token hint and API-key precedence. - `server/src/__tests__/agent-test-environment-routes.test.ts`: new tests for the env merge (agent wins on conflict, forbidden key filtered), missing-binding reporting, and the no-execution-target fallback path. - `docs/adapters/claude-local.md`, `docs/adapters/overview.md`: state the auth-input facts (API key or oauth token wins over stored logins; snapshot-owns-auth applies when neither is configured) and describe the environment-aware Test behavior. ## Verification - `npx vitest run --project @paperclipai/adapter-claude-local` — 131 tests pass (both drifted suites repaired; they fail on master today). - `npx vitest run server/src/__tests__/agent-test-environment-routes.test.ts` — 7 tests pass. - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` — passes with the added project. - `pnpm typecheck` in `server/` and `packages/adapters/claude-local/` — clean. ## Risks - Low risk. The run path is untouched; the probe route change is additive and inert when the environment has no envVars. - The probe now performs environment-consumer secret resolution at test time; access is authorized per binding exactly as at run time, and the audit consumer is the environment (as before for adapter-config resolution). - Enabling the claude-local project in CI adds about 2 seconds of vitest wall time to the general workspaces group and could surface future regressions in that package — which is the point. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking enabled, agentic tool use via Claude Code (CLI). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
97590ff8c4 |
feat(dev): add pnpm dev:mobile and dev:both for prebuilt UI preview (#10718)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI is a React SPA served by the paperclip server; the standard local dev flow is `pnpm dev`, which runs vite in dev mode with HMR and an unbundled module graph > - The unbundled dev bundle is hundreds of MB of JS across many requests, which is fine on a local machine but unusable from a phone or tablet on slow/lossy links (airplane wifi, mobile data, distant tailnet peers) > - Contributors who want to iterate on the board from a mobile device today have no supported way to preview a small production-shaped bundle without stopping the dev server and running a one-off `vite preview` with manual proxy plumbing > - This pull request adds `pnpm dev:mobile` — build the UI and serve `ui/dist` via `vite preview` on port 3101, with `/api` proxied to the running dev server on 3100 — plus `pnpm dev:both` to run both flavors together > - The benefit is a supported second flavor of the dev server for phones/tablets that runs alongside the normal one, without touching the primary `pnpm dev` flow ## Linked Issues or Issue Description **Subsystem affected** ui/ — React + Vite board UI **Problem or motivation** The vite dev server serves an unbundled module graph, which is fine on localhost but unusable from a phone or tablet on a slow link. Contributors testing responsive behavior on mobile devices have no supported way to serve a small production-shaped SPA against the running dev API. Running `vite preview` directly does not work either — the server's board mutation guard checks that the browser's Origin matches the request Host, and a preview on a second port would fail every mutation. **Proposed solution** Add two root scripts: - `pnpm dev:mobile` — build `ui/dist` and serve it via `vite preview` on port 3101, with `/api` proxied to the API server on 3100. - `pnpm dev:both` — run `pnpm dev` and `pnpm dev:mobile` together in a single terminal with prefixed output and shared signal handling. The vite preview config binds `0.0.0.0`, sets `allowedHosts: true` so it accepts arbitrary hostnames (LAN, tailnet, ngrok, etc.), and the shared `/api` proxy forwards the client's original Host header as `x-forwarded-host`. The paperclip server's mutation guard already prefers `x-forwarded-host` over `host` when computing trusted origins, so the browser's Origin becomes trusted automatically. **Alternatives considered** - Bespoke node proxy script — works but duplicates what vite preview already does. - Loosen the mutation guard to accept arbitrary origins — reduces security for the primary server for the sake of a dev-only workflow. - Second server config that binds a second port from the paperclip server itself — much larger change and mixes runtime concerns with a dev-tooling convenience. ## What Changed - New `pnpm dev:mobile` script — build UI then run `vite preview` on port 3101. - New `pnpm dev:both` script — run `pnpm dev` and `pnpm dev:mobile` together via `scripts/dev-both.mjs`, which prefixes each child's output, propagates SIGINT/SIGTERM, and exits when either child exits. - `ui/vite.config.ts` — add a `preview` block (port 3101, host `0.0.0.0`, `allowedHosts: true`, shared `/api` proxy). - New `ui/src/lib/vite-api-proxy.ts` — extracts the `/api` proxy factory shared by dev and preview, and forwards the client Host as `x-forwarded-host` (plus `x-forwarded-proto`). - New unit test `ui/src/lib/vite-api-proxy.test.ts` covering the header-injection behavior and the pass-through when no Host is present. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/lib/vite-api-proxy.test.ts` — 3 tests pass. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `pnpm --filter @paperclipai/ui build` — clean. - Manual: ran `vite preview` against an echo listener and confirmed the request arrives with `x-forwarded-host` set to the client Host header and `x-forwarded-proto: http`. Then ran `pnpm dev:mobile` against the live dev server and verified board mutations (mark issue read, resolve recovery action, run routine) succeed from a second-port browser session that previously 403'd. ## Risks Low risk. Changes are limited to dev tooling — no runtime code paths, no server changes, no schema/migrations. The `apiProxy` refactor is a no-op behaviorally for the existing dev server (same target, same `ws: true`); the only new behavior is the two `x-forwarded-*` headers, and the server side already prefers those headers when trusting origins. `dev:mobile` and `dev:both` are additive; existing `pnpm dev` is untouched. ## Model Used Claude Opus 4.7 (1M context), extended thinking, tool use (bash, file edits). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
dcac49a4fd |
feat(workspaces): defer isolated setup until runtime start (#10653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Isolated workspaces give each task a safe and reproducible checkout. > - The existing setup cloned the development database before an agent needed to run the app. > - This made worktree creation slower and heavier for tasks that never start a service. > - Runtime services already use one server start path for heartbeat, operator, and startup recovery flows. > - This pull request moves heavy setup to that start path and keeps worktree creation lean. > - The benefit is faster isolated workspace creation with the same reliable runtime setup when a service starts. ## Linked Issues or Issue Description Related pull request: #10652 covers the initial deferred database-seeding slice. This pull request supersedes it with end-to-end runtime provisioning and safe cleanup. **What existing behavior does this improve?** This improves isolated worktree creation, runtime service startup, and isolated instance cleanup. **Subsystem affected** Cross-cutting: CLI worktree setup, server runtime orchestration, shared workspace contracts, and development scripts. **Current behavior** Paperclip seeds an isolated development database during worktree creation. It can also leave an isolated instance directory after workspace teardown. This work happens even when no runtime service starts. **Proposed behavior** Paperclip creates the worktree with a lean eager setup. It runs an idempotent runtime provision command before the first managed service spawn. Concurrent starts share one provision attempt. Teardown removes the isolated instance safely. **Reason and benefit** Many agent tasks only edit and test code. They do not need a running Paperclip instance. Deferring the database seed reduces workspace startup cost while preserving automatic setup for tasks that start the app. **Breaking changes** None. The new runtime provision command is optional. Existing workspace behavior is unchanged when it is absent. ## What Changed - Split Paperclip worktree setup into a lean eager script and an idempotent runtime provision script. - Added `runtimeProvisionCommand` to project, issue, realized workspace, and persisted workspace contracts. - Added a per-workspace provision mutex before local service spawn for heartbeat, operator, and startup recovery flows. - Added a persisted `provisioning` service state and the `workspace_runtime_provision` operation phase. - Kept provision time outside the service readiness timeout and made failed attempts visible and retryable. - Reclaimed isolated instance data during safe workspace teardown. - Serialized deferred database seeding across processes and bound teardown to the instance root captured in persisted workspace metadata. - Added tests for config flow, concurrency, retry, no-op behavior, readiness timing, scripts, CLI commands, and cleanup. - Documented the eager and runtime provisioning contracts. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` (server: 3,201 passed; UI: 3,345 passed; the CLI phase exposed one environment-sensitive AWS doctor assertion because the agent runtime injects static AWS credentials) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when non-secret provider config is present'` - Focused runtime tests cover serialized provisioning, retry after stderr failure, absent-command no-op behavior, operation logging, persisted state order, and readiness timeout exclusion. - Focused CLI and cleanup tests cover concurrent seed serialization, stale-lock fail-closed behavior, persisted instance ownership, and rewritten sibling pointers. ## Risks - A faulty runtime provision script blocks service startup. Paperclip records stderr, marks the service failed, and retries on the next start. - Concurrent service requests share an in-process provision attempt, while the seed command uses an atomic filesystem lock across processes. A stale lock fails closed and requires an operator to verify no seed is running before removing it. - Isolated instance cleanup is destructive. The cleanup service validates ownership and path containment before removal. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with agentic reasoning, tool use, and code execution. The service does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8540ce2973 |
ci: shard general-server tests 4 ways (#10663)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI must give contributors fast and stable feedback. > - The `general-server` Vitest lane runs many single-worker server suites. > - A recent completed PR run showed this lane as the slowest completed check. > - Three shards still left one runner with the largest share of work. > - This pull request splits that lane into four duration-balanced shards. > - The benefit is a shorter critical path for the same server test coverage. ## Linked Issues or Issue Description No public GitHub issue exists for this CI maintenance change. **Pre-submission checklist** - I confirmed this improves existing behavior. It does not add a new command, endpoint, or concept. - I searched open public issues and pull requests for related CI sharding work. **What existing behavior does this improve?** The pull request workflow's `general-server` Vitest lane. **Subsystem affected** Cross-cutting. This affects GitHub Actions CI and the Vitest shard duration manifest. **Current behavior** The `general-server` lane uses three shards. The server suites now total about 880 seconds of serial Vitest wall time. The slowest shard was about 313 seconds in the measured run. **Proposed behavior** The `general-server` lane uses four shards. Each shard receives about 220 seconds of predicted suite weight from the refreshed duration manifest. **Reason and benefit** The slowest PR check controls how soon a reviewer can trust the PR. Four balanced shards reduce the slowest `general-server` shard while keeping the same suite selection rules. **Breaking changes** None. This only changes CI partitioning and duration data for existing test suites. **Additional context** Related public searches found no exact open issue or pull request for this `general-server` sharding change. ## What Changed - Split the `general-server` CI matrix from three shards to four shards. - Refreshed `scripts/general-server-shard-durations.json` with wall-time weights from a recent completed PR run. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check origin/master...HEAD` - Dry-ran the four `general-server` shards locally during implementation. The partition covers 300 unique suites with about 220.56 seconds of predicted weight per shard. - Ran a local sensitive-data scan before push. It found only test filenames that contain words such as `secret` or `token`, not credential values. ## Risks Low risk. The main risk is that the duration manifest becomes stale as suite costs move. Missing suites fall back to the median weight, so the lane still runs if the manifest is incomplete. ## Model Used OpenAI Codex, GPT-5, with tool use and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
173d6d2a71 |
Reclaim isolated worktree instances during teardown (#10649)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip creates isolated instances for server-managed git worktrees. > - The worktree teardown path removes the git worktree but leaves its isolated instance directory behind. > - The leaked directory can retain an embedded PostgreSQL process and database files. > - Teardown must remove only the collision-resistant instance assigned to that exact worktree path. > - This pull request stops the verified embedded PostgreSQL process and removes the guarded instance directory. > - The benefit is complete worktree cleanup without risk to another, default, or live Paperclip instance. ## Linked Issues or Issue Description **What happened?** Closing a server-managed git worktree removed the git worktree and branch, but it left the isolated Paperclip instance directory behind. A live embedded PostgreSQL process could also keep running against that directory. **Expected behavior** Worktree teardown must stop the isolated embedded PostgreSQL process and remove only the instance assigned to that exact worktree. It must refuse mismatched instance IDs and all paths outside `PAPERCLIP_WORKTREES_DIR/instances/`. **Steps to reproduce** 1. Create a server-managed git worktree with a repo-local `.paperclip/.env` file. 2. Start its isolated embedded PostgreSQL instance. 3. Close the execution workspace. 4. Observe that the git worktree is removed but the isolated instance directory remains. **Paperclip version or commit** The bug reproduces on `master` before this change. **Deployment mode** Local development with a server-managed git worktree and embedded PostgreSQL. ## What Changed - Give server-managed worktrees collision-resistant instance IDs derived from their resolved absolute paths. - Capture the repo-local instance pointer before custom teardown commands can remove it. - Require the pointer's instance ID to match the exact worktree-derived ID. - Resolve and validate the instance path against the canonical managed worktree instance root. - Verify and stop the matching embedded PostgreSQL process before directory removal, including process-exit races. - Record successful and refused cleanup operations in the workspace operation log. - Add focused ownership, process-race, path-safety, and runtime integration tests. - Document automatic isolated-instance cleanup for server-managed worktrees. ## Verification - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-instance-cleanup.test.ts` — 9 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/workspace-runtime.test.ts -t "records teardown and cleanup operations when a recorder is provided"` — 1 test passed and 99 tests skipped. - `node scripts/__tests__/provision-worktree-self-heal.test.mjs` — 4 tests passed. - `bash -n scripts/provision-worktree.sh` — passed. - `pnpm --filter @paperclipai/server build` — passed. - `git diff --check` — passed. ## Risks The main risk is removal of the wrong instance directory. Provisioning assigns a path-derived ID with a SHA-256 suffix, and cleanup requires that exact ID in addition to a safe instance identifier, an absolute configured home, a strict child path, canonical path checks, and a second canonical path check immediately before removal. It refuses legacy or mismatched IDs, symlink escapes, and all paths outside the managed worktree instance root. Cleanup failures become visible warnings and do not delete an unverified path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The runtime does not expose the exact model snapshot or context-window size. The agent used reasoning, repository tools, GitHub tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <paperclip@paperclip.ing> |
||
|
|
8444e5735c |
ci: pass e2e shard specs without separator (#10640)
## Thinking Path > - Paperclip uses pull request CI to test changes before merge. > - The e2e PR lane runs Playwright specs in a shard matrix. > - Each shard builds a list of spec files for its matrix entry. > - The workflow passed that list after a literal `--` separator. > - Playwright did not receive the list as file filters. > - This pull request removes the separator and adds a guard test. > - The benefit is that each e2e shard runs only its assigned specs. ## Linked Issues or Issue Description Refs #10629. **What happened?** The e2e shard step used `pnpm run test:e2e -- $specs`. The shard spec list was not applied as Playwright file filters. **Expected behavior** Each e2e shard should pass only its selected specs to Playwright. **Steps to reproduce** 1. Inspect `.github/workflows/pr.yml` at the merge commit for #10629. 2. Find the `e2e_shards` command that invokes `pnpm run test:e2e`. 3. See the literal `--` before `$specs`. **Paperclip version or commit** `86767951` **Deployment mode** GitHub Actions PR CI. ## What Changed - Removed the literal `--` from the e2e shard `pnpm run test:e2e $specs` invocation. - Added a regression test that checks the workflow passes `$specs` without that separator. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks Low risk. This changes one CI command and one workflow guard test. The main risk is shell argument handling in the workflow, and the guard now covers the expected command shape. ## Model Used OpenAI GPT-5 through Codex. The run used shell and GitHub CLI tool access. The runtime did not expose a context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8676795188 |
ci: split e2e PR lane into three shards (#10629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow protects changes with a Playwright e2e lane. > - That lane already uses a weighted file partition so slow specs do not cluster by test count. > - Recent green PR runs showed the two e2e shard jobs were slower than the next slow required lane. > - The largest spec is indivisible, so a third shard lets that spec run alone and lets the rest split by duration. > - This pull request changes only the PR e2e shard matrix and the guard test. > - The benefit is a shorter expected PR critical path while the required `e2e` aggregate check name stays stable. ## Linked Issues or Issue Description Refs #9923 **What existing behavior does this improve?** The `pull_request` workflow Playwright e2e lane. **Subsystem affected** Cross-cutting: GitHub Actions CI and test scripts. **Current behavior** The PR workflow runs the weighted Playwright e2e partition across two jobs. Recent green runs showed those jobs as the slowest required checks. **Proposed behavior** The PR workflow runs the same e2e spec set across three weighted jobs. The aggregate required check stays named `e2e`. **Reason and benefit** The third shard lets the slow smoke-lab spec run alone while the rest of the catalog stays balanced. This should shorten the PR critical path. The win is bounded by fixed per-job setup time. **Breaking changes** None. The required aggregate check contract is preserved. ## What Changed - Change the PR e2e shard matrix from two entries to three entries. - Update the shard guard test to expect three shards. - Floor the balance bound at the largest single spec weight. - Assert that the workflow does not define more shard indexes than `SHARD_COUNT`. ## Verification - `node --test ./scripts/__tests__/e2e-shard.test.mjs` passes with 6 tests. - The recorded-weight partition is complete and non-overlapping: 168.0s, 116.5s, and 114.4s. - I checked `ROADMAP.md` and found no overlapping roadmap-level core feature. - I searched public GitHub PRs and issues for related e2e shard work. I found related PR #9923 and no open duplicate for this branch or change. ## Risks - This adds one extra GitHub Actions runner to the PR e2e lane. - The wall-clock win is bounded by fixed per-job setup. - Behavior risk is low because the aggregate required check remains named `e2e`. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent in this repository. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> |
||
|
|
c185e64b77 |
feat(ui): chat-style task view behind an experimental flag (#10606)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators spend most of their time on the issue detail page. They talk to the assigned agent there through comments. > - The current page reads as a ticket form. The thread sits below properties, the composer sits mid-page, and live agent activity renders as dense transcript logs. > - Talking to an agent is a conversation. A chat-first layout matches that mental model better than a ticket form. > - A layout change this large must not disrupt current users. It needs a safe opt-in path and full parity with the existing thread features. > - This pull request adds a chat-style task view behind a new "Chat-Style Tasks" experiment toggle. The flag is off by default and the existing page is unchanged when it is off. > - The benefit is a focused, readable conversation with the agent: live tool activity folds into compact summaries, the composer stays at the bottom, and properties, plan, and artifacts move into header tabs. ## Linked Issues or Issue Description Refs #49 (chat with agents is a much-wanted feature). Related PRs found in the dedup search: - #4489 — an earlier, closed attempt to promote the conversation to the primary surface on issue detail. This PR is a fresh, flag-gated take on the same goal. - #8837 — an open PR that proposes a two-column task layout. It restructures the same page but keeps the ticket paradigm; this PR is orthogonal because it is opt-in and chat-first. **Subsystem affected** UI (issue detail page). **Problem or motivation** The issue detail page presents agent conversations as a ticket: properties first, thread below, composer in the middle of the page, and raw transcript noise during live runs. Users who mainly converse with their agents must scroll past chrome to follow the conversation, and live activity is hard to read. **Proposed solution** An opt-in chat-style view of the issue detail page, gated by a new "Chat-Style Tasks" experiment toggle in Settings → Experimental. With the flag on, the thread fills the center pane, the composer docks to the bottom of the viewport, Properties / Plan / Artifacts become header tabs, live turns show a status pill with the current tool action and elapsed time, and settled turns collapse to a "Worked · N tools" summary that expands into per-tool rows. With the flag off, nothing changes. **Alternatives considered** Restyling the existing layout in place (rejected: too disruptive without an opt-out), and a separate chat page beside the issue page (rejected: splits the task's single source of truth). A per-request lab page (`/task-chat-lab`, dev-only) was kept for design iteration instead. **Roadmap alignment** ROADMAP.md "CEO Chat" wants lighter conversations that still resolve to real work objects. This PR keeps the core task-and-comments model — it only changes presentation, opt-in — so it does not duplicate that planned work. ## What Changed - New `enableTaskChatRedesign` instance setting, exposed as a "Chat-Style Tasks" experiment card in Settings → Experimental (shared feature catalog, validators, server instance-settings service, and UI settings page). - New `ui/src/components/task-chat/` component family: chat thread with turn grouping, agent reply bubbles, live status pill, collapsible turn summaries with per-tool rows, plan tab with a sticky CTA action bar, inline interaction cards, per-request mode chips, and a bottom-docked composer. - A shared tool taxonomy (`tool-taxonomy.ts`) maps tool names to verbs and icons; the status pill, tool rows, and the classic transcript view all use it. - A transcript adapter converts stored run logs into chat turns; it dedupes tool-call updates by `toolUseId` so tool counts match the expanded rows, and it keeps a tool row's first real name when later generic updates arrive. - Composer: posts on Cmd/Ctrl+Enter, supports image paste with object-URL thumbnail previews (revoked on clear/unmount), and uploads through the issue attachments route. - `IssueDetail.tsx`: with the flag on, pane tabs move to the header bar, the header is not sticky, and the chat fills the center; with the flag off, the previous layout renders unchanged. - Motion tokens for the new animations live in `ui/src/index.css` with a `motion-tokens.ts` catalog and a test that keeps the two in sync (the catalog now also covers the shared enter/exit/swap tokens that the decision/quicklook block declares). - A dev-only `/task-chat-lab` page with fixtures and a tweak panel for motion tuning. ## Verification - `pnpm typecheck` — clean across the workspace. - `pnpm check:token-gates` — 3/3 CLEAN. - `cd ui && pnpm vitest run` — 3,344 of 3,345 tests pass locally. The one failure is `IssueProperties.test.tsx` monitor-row time formatting, which is timezone-sensitive: it also fails on unmodified `origin/master` in a non-UTC timezone and passes with `TZ=UTC`. It is not related to this change. - `cd server && pnpm vitest run src/__tests__/instance-settings-service.test.ts` — 21/21 pass (covers the new setting). - Manual: start the dev server, open Settings → Experimental, enable "Chat-Style Tasks", and open any issue. The thread fills the page, the composer docks to the bottom, and Properties / Plan / Artifacts appear as header tabs. Assign an agent and comment to watch a live run: the status pill shows the current tool action with elapsed time, and the finished turn folds into a "Worked · N tools" summary. Disable the toggle and confirm the classic page is unchanged. - Visual snapshot baselines are intentionally not updated: per `doc/design/DECISION-SHEET.md`, "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks - The flag-off path goes through the same `IssueDetail.tsx` file, so a regression there would affect current users. Mitigation: the classic markup renders through the same components as before behind explicit flag conditionals, and the full UI suite passes. - The transcript adapter interprets stored run-log formats, including legacy entries without `toolUseId`. Malformed logs degrade to generic tool rows rather than crashing. - The new view changes no server behavior other than one additive instance setting; it is additive and default-off. Overall risk with the flag off is low. ## Model Used - Claude (Anthropic), model id `claude-fable-5` (Claude Fable 5), extended thinking enabled, agentic tool use (file editing, shell, test execution) via Claude Code / Claude Agent SDK. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b163c6c473 |
fix(server): preserve hot restart intent across path upgrade (#10593)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server can preserve eligible agent runs during a controlled hot restart. > - A path change moved restart state from the Paperclip home root to the instance root. > - A staged update can therefore make the old server and the new server read different intent files. > - The old server then drains live runs, while the new server can start without a shutdown snapshot. > - This pull request adds a correlated compatibility handoff and records the live preflight set. > - It also verifies the target process instance on Linux, macOS, and Windows. > - The benefit is complete and safe run classification across the path upgrade. ## Linked Issues or Issue Description No public GitHub issue covers this defect. **What happened?** A staged hot restart can run an older server that reads `hot-restart-intent.json` from the Paperclip home root and a new server that writes the file under the instance root. The old server misses the request and uses graceful drain. The new server later finds its marker without a shutdown snapshot. Before this change, that state could produce an empty loss list even when live runs existed before restart. **Expected behavior** The old server must receive the PID-targeted restart request at its legacy path. The new server must correlate the legacy shutdown snapshot with its instance-scoped request. Every run that was live during preflight must appear as adopted, finalized while down, or lost. A reused PID must not let a stale marker claim a different process instance. **Steps to reproduce** 1. Start a server version from before the instance-root marker change. 2. Keep one or more local-agent heartbeat runs active. 3. Stage a current build and request a hot restart from that build. 4. Observe that the old server reads only the home-root path while the staged build writes only the instance-root path. 5. Observe graceful drain and a new-server intent that has no shutdown snapshot. **Paperclip version or commit** The path transition entered `master` in #10045. The hot-restart adoption flow came from #9647. This fix targets current `master` and compatibility with the immediately preceding home-root behavior. **Deployment mode** Self-hosted server built from source with controlled service hot restarts. Related work: #9628 is the original broader hot-restart feature PR. #10556 addresses embedded PostgreSQL lifecycle behavior and does not address marker-path compatibility. ## What Changed - Write an authoritative instance-scoped intent and a correlated legacy home-root handoff marker. - Merge a legacy shutdown snapshot only when immutable request identity fields match. - Prevent a non-default instance from consuming an uncorrelated legacy-only marker. - Record preflight running heartbeat IDs and reconcile snapshot omissions from current database state. - Serialize marker claims, snapshot writes, stale recovery, and matching cleanup with recoverable per-path filesystem leases. - Read process start identity on Linux, macOS, and Windows to distinguish a reused PID from the original server. - Require identity for new restart requests and fail closed when a supported platform cannot provide it. - Classify older markers by comparing the replacement server boot time or operating-system process start time with the request time. - Close the preflight database client explicitly and use a root-safe SQL query. - Add focused unit, platform-branch, database-backed, and CI regression coverage. - Document the compatibility handoff, process identity probes, and instance-scoped report path. ## Verification - `pnpm exec vitest run server/src/services/hot-restart.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts -t "hot-restart|old-server legacy|preflight live|preflight run|spawn identity before hot restart"` — 24 tests passed and 90 tests were skipped across 2 files. - `pnpm exec vitest run server/src/services/hot-restart.test.ts` — 17 tests passed. - `pnpm exec vitest run server/src/__tests__/issue-watchdogs-routes.test.ts -t "restarts a stalled claimed run"` — 1 test passed and 10 tests were skipped. - `pnpm exec vitest run server/src/__tests__/agent-action-audit-routes.test.ts -t "allows an agent with issue:delegate"` — 1 test passed and 7 tests were skipped. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. - GitHub Actions — 26 of 26 checks passed at `55a79cb029be8b1dc89926d9d89ccd2181266d5c`. - Greptile — 5/5 at the same head with no unresolved current-head review threads. ## Risks - The legacy handoff path is shared across instances. Exclusive claims and per-path leases prevent overwrite and match-before-delete races. - Process identity uses platform commands as a fallback when the health endpoint has no identity. Linux reads `/proc`, macOS and BSD use `ps`, and Windows uses PowerShell. - A supported-platform identity probe failure aborts the restart. This fails closed instead of replacing an unknown live process. - Older intent files do not contain process identity. The server compares the replacement boot or process start time with the request time when those values are available. - A preflight database read can fail before the marker is written. The command fails closed instead of claiming a restart whose live-run set is unknown. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment model ID and context-window size were not exposed by this runtime. Reasoning, repository editing, shell execution, test execution, GitHub CLI, and Paperclip API capabilities were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9116686bd |
test(installer): cover cross-version update migrations (#10587)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Managed updates can change both the application payload and its database schema > - Unit tests cannot prove that an older live install upgrades through a real migration and remains recoverable > - The managed-install work in #10045 needs a repeatable cross-version system test > - This pull request adds an isolated end-to-end harness for update, migration, backup, service restart, and rollback behavior > - The benefit is a direct proof that managed upgrades preserve the existing database and service lifecycle across versions ## Linked Issues or Issue Description Refs #10045 This test is a focused follow-up to the managed install integration. Merge #10045 first so the tested install, update, service, backup, and rollback commands are available. ## What Changed - Added a cross-version managed-update E2E script. - Installed an older Git ref, initialized its embedded PostgreSQL database, and updated to a ref with one additional migration. - Verified the pre-update backup, payload switch, service recovery, migration result, database-cluster reuse, and rollback behavior. - Isolated Paperclip state under a dedicated test home and cleaned up the service and managed install on success or failure. - Added regression tests for shell syntax, required-ref validation, side-effect-free preflight failure, and complete failure cleanup. ## Verification - `node --test scripts/__tests__/e2e-update-migrations.test.mjs` - `bash -n scripts/e2e-update-migrations.sh` - GitHub latest-head CI: build, typecheck, release registry, canary dry-run, general tests, serialized suites, and both browser E2E shards passed. - Full harness execution needs an isolated macOS or Linux host with a real launchd or systemd user service. It is intentionally not run on a live Paperclip server host. ## Risks - The script manages a real user service and downloads two Git refs. Run it only on an isolated test host. - The test needs #10045 because `origin/master` does not yet contain the managed install lifecycle. - The script uses a dedicated `PAPERCLIP_HOME`, refuses a pre-existing shim or test home, and removes its service and install during cleanup. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The runtime did not expose a more specific deployment ID or context-window size. Reasoning, repository access, shell execution, and GitHub tooling were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc5a30805e |
feat(cli): add managed install, update, and service lifecycle (#10045)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need a predictable installation path that survives beyond an ephemeral `npx` process > - A durable installation needs an owned per-user payload store, stable command shim, safe shell integration, and supported service lifecycle > - Updates must preserve recoverability by backing up data, installing side-by-side, verifying the new payload, and retaining rollback state > - Bootstrap scripts and privileged service operations must fail closed across download, filesystem, ownership, and consent boundaries > - This pull request integrates managed install, update, rollback, service, uninstall, doctor, bootstrap-installer, and runtime-serving support into one workflow > - The benefit is a recoverable, inspectable, and documented installation lifecycle with explicit safety boundaries across Linux, macOS, containers, WSL, npm, npx, and source checkouts ## Linked Issues or Issue Description ### Problem Paperclip lacks a first-class durable installation and lifecycle workflow. Operators currently have to assemble npm/npx installation, PATH setup, background-service management, updates, rollback, diagnostics, and uninstall behavior themselves. That makes upgrades harder to recover, creates inconsistent behavior across platforms, and leaves shell/download/service trust boundaries without one documented implementation. ### Proposed Solution Add a managed per-user install store and stable shim, a verified shell bootstrap installer, service lifecycle commands, install-mode-aware update/rollback behavior, doctor checks, and documentation. Managed updates back up the database, install and smoke-test a side-by-side payload, atomically switch `current`, and retain prior payloads. The shell installer pins registry/download trust boundaries and requires explicit consent for non-interactive privileged actions. ### Alternatives Considered - Keep recommending `npx`: simple for evaluation, but ephemeral and unsuitable for stable services, atomic updates, or rollback. - Require global npm installation only: familiar, but cannot provide the owned side-by-side payload store and retained rollback semantics. - Split the capability across multiple PRs: rejected because install, update, service, uninstall, bootstrap, and serving behavior share contracts and security boundaries that need review together. ### Related Pull Requests - Supersedes #10042 and #10044 with one integrated final diff. - Incorporates and replaces the closed preparatory work in #10032 and #10034. ## What Changed - Added `paperclipai install`, `update`/`upgrade`, rollback, uninstall, service lifecycle, onboarding integration, and managed-install doctor checks. - Added a private managed payload store, verified manifest/marker ownership, exclusive mutation locks, atomic manifest/current/shim writes, retained previous payloads, and provenance validation. - Added npm and GitHub-ref install sources with exact target resolution, registry isolation, database backup, side-by-side verification, atomic activation, service restart coordination, and failure rollback. - Made managed-update backups report actionable service-start and `--no-backup` recovery guidance for unreachable databases, while clean never-onboarded instances skip an empty backup. - Added systemd user and launchd service definitions, status/health/log commands, single-instance coordination, stale-port recovery, and explicit sudo/lingering consent handling. - Added the `scripts/install.sh` bootstrap path with checked two-stage downloads, pinned public npm registry usage, platform checks, dry-run/non-interactive controls, and Docker fixtures. - Added embedded Postgres/native bootstrap integration, hot-restart/systemd-notify serving support, passive update notices, configuration contracts, README/CLI/install documentation, and focused regression tests. - Security re-review should explicitly re-verify: (1) `addManagedPathBlock`/`removeManagedPathBlock` reject symlinked or non-regular rc files, assert current-user ownership, preserve restrictive modes, and replace atomically; (2) managed shim replacement rejects unsafe parents, foreign-owned or multiply linked files, and uses checked atomic replacement; (3) the shell installer and sudo path preserve explicit consent and checked downloads; and (4) installed service/runtime serving remains bound to the validated managed shim and instance configuration. ## Verification - `bash -n scripts/install.sh scripts/clean-install-git.sh scripts/clean-install-npm.sh scripts/test-install-sh-docker.sh` - `pnpm exec vitest run cli/src/__tests__/install-store.test.ts cli/src/__tests__/install-command.test.ts cli/src/__tests__/managed-install-check.test.ts cli/src/__tests__/onboard-service.test.ts cli/src/__tests__/service-health-check.test.ts cli/src/__tests__/service-manager.test.ts cli/src/__tests__/update-command.test.ts cli/src/__tests__/update-notice.test.ts packages/db/src/embedded-postgres-native.test.ts` — 9 files, 66 tests passed - `pnpm --dir cli typecheck` - `pnpm --dir cli build` - Follow-up verification: `pnpm exec vitest run cli/src/__tests__/update-command.test.ts` (14/14), `pnpm --dir cli typecheck`, `pnpm --dir cli build`, and `pnpm --filter @paperclipai/server typecheck`. - `pnpm -r typecheck` - `pnpm build` - Full `pnpm test:run` exercised all suites; an injected static AWS credential changed one unrelated doctor expectation, which passed when those credentials were removed. A second run cleared that case and exposed stale pre-existing adapter-utils `dist` output; rebuilding `@paperclipai/adapter-utils` made the isolated test pass. The updated PR CI is the authoritative clean-workspace full-suite run. ## Risks - Installer/update code writes executable shims, symlinks, shell rc blocks, service definitions, and managed payloads; ownership, regular-file, symlink, hard-link, marker, and path-containment checks fail closed before destructive changes. - The bootstrap installer executes downloaded tooling; downloads are staged and checked before execution, npm traffic is pinned to the public registry, and non-interactive privileged behavior requires explicit consent. - Linux lingering may invoke `sudo`; the command is surfaced and confirmed before execution, and unsupported service managers fall back to foreground-run guidance. - Database migrations remain forward-only; payload rollback does not reverse migrations, so managed updates create a backup before activation unless explicitly disabled. - Service restart and runtime serving touch process/port ownership; lifecycle locks, health/version checks, and stable-shim service definitions reduce split-brain and stale-process risk. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agents using GPT-5.5 and GPT-5.6-sol, with reasoning, repository/API access, shell execution, and test tooling. The runtime did not expose a reliable context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
79eff0aea1 |
fix(scripts): self-heal isolated workspace provisioning when the base CLI is broken (#10574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run tasks in isolated execution workspaces that are provisioned as git worktrees by `scripts/provision-worktree.sh` > - The script runs the base workspace's CLI (`cli/src/index.ts` via the base `tsx` install) to seed each new worktree, and it only checked that those files exist > - pnpm links each package's `node_modules` into a hash-versioned virtual store; a lockfile change followed by a partial or filtered install prunes old hashed dirs without relinking every package, leaving dangling symlinks > - A CLI with dangling symlinks fails ESM resolution (`ERR_MODULE_NOT_FOUND`) at boot, so provisioning aborts with `setup_failed` — deterministically, on every retry, with no self-heal path > - This pull request makes provisioning health-check the CLI by actually booting it, repair the base install when the check fails, and degrade to the no-CLI fallback config instead of failing the run > - The benefit is that a class of permanent `setup_failed` loops becomes self-healing, and workspace provisioning survives a broken base CLI ## Linked Issues or Issue Description No public GitHub issue exists; the underlying bug is described here per `bug_report.yml`. Related PR: #10578 self-heals the sibling workspace-validation failure loop uncovered by the same incident diagnosis. **What happened?** Isolated-workspace runs failed at provision time with `setup_failed`. Every retry failed identically. One observed incident burned 4 runs across two adapters before the task was stranded. **Expected behavior** Provisioning either succeeds or degrades gracefully; a broken base CLI install repairs itself instead of permanently blocking all new worktrees. **Steps to reproduce** In the base workspace, cause a lockfile-affecting dependency bump plus a partial/filtered `pnpm install` so a package symlink (e.g. `cli/node_modules/drizzle-orm`) dangles into a pruned virtual-store dir. Start any isolated-workspace run. Provision fails with `ERR_MODULE_NOT_FOUND` and the run ends `setup_failed`; retries never recover. **Paperclip version or commit** master as of the branch point of this PR. **Deployment mode** Local trusted deployment with git-worktree isolated workspaces. ## What Changed - `base_cli_healthy` now boots the base CLI (`--help`) instead of only testing file existence, which exercises the top-level import graph. - New `repair_base_workspace_install`: when the health check fails, run a non-interactive `pnpm install --prod=false --force --frozen-lockfile` in the base workspace. `--force` guarantees relinking when pnpm's up-to-date heuristics would skip dangling symlinks; `--frozen-lockfile` keeps the repair from mutating the shared lockfile. - The repair install is serialized with `flock` on a lock file inside the resolved git dir (`git rev-parse --absolute-git-dir`), so locking also covers base workspaces that are linked worktrees, where `.git` is a file. - If every CLI candidate is unusable (including a base CLI the repair could not fix), provisioning falls back to the existing no-CLI fallback config writer (loudly, on stderr) instead of failing the run. A CLI that runs and fails `worktree init` still fails provisioning with its real exit code — that deliberate fail-closed policy is unchanged and covered by an existing server regression test. - Fixed a latent bug: `run_isolated_worktree_init` returned 0 unconditionally after the init subshell, so callers treated a failed init as success. Exit codes now propagate. ## Verification - Reproduced the incident state (dangling `cli/node_modules/drizzle-orm` symlink); the base CLI failed with the exact `ERR_MODULE_NOT_FOUND` seen in the incident run logs. - Ran the patched script against a fresh scratch worktree: health check failed → locked repair install ran (~26 s warm) → symlink relinked → `worktree init` completed → exit 0 with `.paperclip/config.json` and `.env` written. - Happy path (healthy base CLI): provisioning behavior unchanged, exit 0. - Verified `git rev-parse --absolute-git-dir` resolves a real directory for both a normal checkout and a linked worktree. - New hermetic tests: `node --test ./scripts/__tests__/provision-worktree-self-heal.test.mjs` (4 tests: healthy CLI used, broken CLI degrades, locked repair end-to-end with a fake pnpm, init failure propagates). Not yet wired into a CI workflow. - `server`: the existing `realizeExecutionWorkspace` fail-closed regression test ("fails instead of writing an unseeded fallback config when worktree init errors after CLI detection succeeds") passes against the new script. - `bash -n scripts/provision-worktree.sh` is clean. ## Risks - Low risk overall: the script only adds recovery paths; the happy path is unchanged. - The repair install runs in the shared base workspace. It is bounded by `--frozen-lockfile` (no lockfile mutation) and serialized by `flock`, but it can add ~30 s to the first provision after a base install breaks. - If the repair cannot fix the CLI and no other CLI candidate exists, runs now continue with an unseeded fallback config instead of failing; that is intentional, and the fallback path already existed. Genuine `worktree init` failures from a working CLI still fail the run. ## Model Used Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use (Claude Code harness). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd1a7f5290 |
Ensure app-home ownership before the privilege drop, not only on remap (#10530)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Docker image persists all instance state (project checkouts, worktrees, run logs, uploads) under `PAPERCLIP_HOME`, and deployments mount a volume there for durability > - The entrypoint starts as root and drops privileges to the `node` user, but it fixes `PAPERCLIP_HOME` ownership only when it remaps the user's UID/GID > - A freshly mounted volume arrives root-owned and shadows the image's build-time `chown`, so a default-UID boot drops privileges onto an unwritable home and the server crashes on its first `mkdir` > - This pull request makes the entrypoint probe the home's ownership and chown whenever it does not match the runtime user, before the privilege drop > - The benefit is that the image works out of the box on any platform-managed volume, with the common already-correct boot staying chown-free ## Linked Issues or Issue Description No public issue exists — describing the bug inline (per the bug report template). **What happened?** Running the image with a freshly created volume mounted at `/paperclip` (a Docker named volume, a Kubernetes PV, or any platform-managed volume) and the default `USER_UID`/`USER_GID` crashes on boot: `Error: EACCES: permission denied, mkdir '/paperclip/instances/default/logs'`. **Expected behavior** The container boots and initializes its instance tree on the mounted volume, exactly as it does when `/paperclip` is the image's own (build-time chowned) directory. **Steps to reproduce** 1. `docker volume create paperclip-data` 2. `docker run -v paperclip-data:/paperclip ghcr.io/paperclipai/paperclip:<any current tag>` 3. Observe the EACCES crash on the first `mkdir` under `/paperclip`. **Root cause** `scripts/docker-entrypoint.sh` chowns `/paperclip` only inside its UID/GID remap branch (`changed=1`). A fresh volume mount is root-owned and shadows the image's build-time `chown node:node /paperclip`; with the default 1000:1000 no remap happens, so no chown happens, and `gosu node` drops onto an unwritable home. **Paperclip version or commit:** reproduces on `master` and any published image. **Deployment mode:** any; observed on managed-cloud volume mounts and reproducible with plain Docker named volumes. **Installation method:** Docker image (`ghcr.io/paperclipai/paperclip`). **Related PRs (dedup search):** no open or merged PR touches the entrypoint ownership logic; the entrypoint's privilege-handling tests were added previously and this extends them. No duplicate found. ## What Changed - `scripts/docker-entrypoint.sh`: the remap-conditional `chown` is replaced by an ownership probe — after any UID/GID remap, the entrypoint stats `PAPERCLIP_HOME` (default `/paperclip`) and runs `chown -R node:node` only when the owner does not match the runtime user, before `exec gosu node`. Covers fresh root-owned mounts and trees written under a previous UID mapping; the already-correct boot performs no chown. The unprivileged (non-root start) branch is unchanged. - `server/src/__tests__/docker-entrypoint.test.ts`: `stat` stub added to the harness; new cases for the fresh root-owned mount with default UID/GID and for `PAPERCLIP_HOME`-relative probing; the remap case now models the post-remap ownership mismatch. ## Verification - `pnpm vitest run server/src/__tests__/docker-entrypoint.test.ts` — 7 passed (5 existing behaviors unchanged, 2 new). - Live on a managed deployment: a container that crash-looped with the EACCES above boots cleanly once the home is chowned before the drop (the same effect this entrypoint change produces; forced there by a UID remap as an interim workaround). ## Risks - Low. Behavior changes only for boots where `PAPERCLIP_HOME` exists with mismatched ownership — exactly the boots that crash today. `chown -R` on a large previously-mismatched tree adds one-time boot latency; correctly-owned homes skip it entirely. Kubernetes restricted / OpenShift non-root starts keep the existing exec-directly path untouched. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic; Claude Code CLI with extended thinking and tool use; tests executed locally via Vitest). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no duplicates; extends the existing entrypoint privilege tests) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b517b887ad | fix(acpx): decouple host proxy spawn cwd from in-sandbox remoteCwd (#10122) | ||
|
|
ad74fb5450 |
Add a feature catalog build artifact derived from the experimental settings schema (#10055)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Instances expose ~23 experimental feature settings, all declared in one shared zod schema and toggled per instance > - Deployment tooling and hosting control planes have no machine-readable list of those feature keys for a given release — the schema is only reachable from code that imports the package > - Any external system that references feature keys therefore does so as free text, and typos drift silently > - This pull request derives a versioned `feature-catalog.json` build artifact from the schema, with a compiler-checked metadata map so the schema stays the single source of truth > - The benefit is a stable contract external tooling can validate feature-key references against, with zero runtime behavior change ## Linked Issues or Issue Description No public issue exists; `feature_request` template fields: **Problem or motivation:** External deployment tooling cannot enumerate or validate an instance's feature keys per release; free-text references fail silently when keys are renamed or removed. **Proposed solution:** A metadata map keyed by the settings schema's own keys (compiler flags drift) plus a build step emitting `feature-catalog.json` (keys, tiers, defaults, `catalogVersion`) as a release artifact. **Alternatives considered:** A hand-maintained catalog file (drifts from the schema); serving the schema from a runtime API (requires a running instance at validation time — a build artifact works offline and pins to a release). **Roadmap alignment:** Supports the in-progress "Cloud deployments" milestone in `ROADMAP.md`. ## What Changed Adds a metadata map (title, description, tier, cloud/self-hosted defaults) keyed by the keys of `instanceExperimentalSettingsSchema`, so the schema stays the single source of truth and the compiler flags any drift. A new build step (`build:feature-catalog --version <v>`) emits `feature-catalog.json` — all 23 feature keys, their tiers, and a `catalogVersion` — as a release artifact that managed-hosting control planes can validate feature-flag writes against. No runtime behavior changes. - New `packages/shared/src/feature-catalog.ts`: per-flag metadata map keyed by a type derived from the settings schema (adding/removing/renaming a flag without updating the map is a compile error), plus `featureCatalogArtifactSchema` and `buildFeatureCatalogArtifact`/`renderFeatureCatalogArtifact` for the artifact - New `scripts/generate-feature-catalog.ts` wired as `pnpm build:feature-catalog --version <v>` - `scripts/create-github-release.sh` generates the artifact and uploads it as a GitHub Release asset (with a dry-run preview line) - Tests in `packages/shared/src/feature-catalog.test.ts` ## Verification - `vitest run packages/shared/src/feature-catalog.test.ts` — 9 tests: schema-key coverage, drift detection, artifact shape - `pnpm --filter @paperclipai/shared typecheck` - Artifact generation run end-to-end: `pnpm build:feature-catalog --version 0.0.0-test` emits 23 keys with `catalogVersion` ## Risks Low risk — no runtime behavior changes; the change is metadata, a build script, and a release-artifact emission step only. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use; independently peer-reviewed by a second AI agent before push ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
204c416478 |
fix(release): publish bundled packages with trusted npm staging (#10047)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip publishes its CLI, server, adapters, and shared packages through automated canary and stable release workflows. > - `@paperclipai/adapter-utils` bundles the patched `acpx` runtime, so it must use npm 11 for OIDC trusted publishing. > - The prior staging directory contained pnpm's `.pnpm` symlink forest, which crashes npm 11's directory-pack step on GitHub runners and produces a consumer-broken bundled dependency tree. > - This pull request rebuilds staged production dependencies as a physical npm tree, reapplies repository patches, and publishes that clean directory directly with npm 11 trusted publishing. > - The benefit is a release path that retains GitHub Actions OIDC trusted publishing while shipping a working patched acpx runtime to consumers. ## Linked Issues or Issue Description Refs: #9980, #10030, #10041 No public GitHub issue exists for this release failure. ### What happened? Canary and stable publishing began routing `@paperclipai/adapter-utils` through npm after it declared `bundleDependencies: ["acpx"]`. Publishing the pnpm-deployed directory with npm 11 crashes during npm's directory-pack phase on GitHub runners with `Exit handler never called!`. The same staged shape also produces a broken consumer artifact because acpx cannot resolve transitive runtime dependencies after installation. ### Expected behavior Bundled packages publish directly from a self-contained staging directory through npm 11 OIDC trusted publishing, and consumers receive a working patched acpx runtime with its transitive dependencies. ### Steps to reproduce 1. Stage `packages/adapter-utils` using the old `pnpm deploy`-only shape. 2. Publish that directory with npm 11 on a GitHub runner. 3. npm crashes before registry/OIDC activity while walking the `.pnpm` symlink forest. 4. Install an artifact packed from that old shape into a fresh npm project and run acpx; its runtime dependency resolution fails. ### Deployment mode GitHub Actions canary/stable release workflow. ### Relevant logs or output ```text npm error Exit handler never called! ``` ## What Changed - After `pnpm deploy`, remove the staged pnpm `node_modules` tree and run `npm install --omit=dev --ignore-scripts --no-audit --no-fund` to create a physical hoisted production tree. - Apply every root `pnpm.patchedDependencies` patch whose package is declared in the staged package's bundled dependencies, failing staging if any patch cannot apply. - Assert the staged acpx runtime contains the required `onAgentStderr` patch marker. - Publish the clean staging directory directly with pinned npm 11.18.0, retaining GitHub Actions OIDC trusted publishing, verbose diagnostics, and the duplicate-transparency-log retry without provenance. - Keep pinned npm 10.9.7 packing only for local/dry-run payload verification; registry publishing does not use a tarball argument. - Add focused coverage for npm-tree staging, patch application, direct directory publish arguments, and bundled tlog retries. ## Verification - `bash -n scripts/release-lib.sh scripts/release.sh` - `node --test scripts/release-lib.test.mjs scripts/acpx-patch-packaging.test.mjs` — 12/12 passed. - `pnpm test:release-registry` — 68/68 passed. - Real staging smoke: `node scripts/prepare-bundled-package.mjs packages/adapter-utils <stage>` produced a real `node_modules/acpx` directory, no `.pnpm` directory, and the `onAgentStderr` patch marker. - Real npm 11 directory-publish smoke: `npx --yes npm@11.18.0 publish --dry-run --tag canary --access public --loglevel verbose` packed 26 bundled dependencies and reached the expected existing-version registry rejection without `Exit handler never called!`. - The merge-triggered `publish_canary` workflow remains the live OIDC trusted-publishing verification. ## Risks - The live GitHub Actions trusted-publishing path can only be fully proven by the merge-triggered canary run; npm debug-log upload remains available if it fails. - Bundling acpx continues to freeze platform-specific transitive artifacts such as esbuild binaries from the Linux release runner. This is a pre-existing consequence of the bundling decision in #9980 and is not expanded here. - Rebuilding dependencies with npm depends on the exact bundled dependency versions in the staged manifest; staging fails hard if repository patches no longer apply. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.6-sol`, high reasoning mode, with repository tool use and code execution. The harness did not expose a model context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |