mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
ffff1fe6e38b4f457cd1eb091cb13dfc56958ea8
32
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ffff1fe6e3 |
feat(runner): define package API and verification boundary (#12129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package now has protocol, transport, provider, catalog, and authorization foundations. > - Its first upstream package boundary should expose only the implemented runtime and test-helper surfaces. > - Rust correctness belongs in the repository existing build verification, without introducing a parallel release process. > - Direct package creation must build the files declared by the package manifest. > - This pull request defines the minimal package API and verifies the optimized runner binaries in the existing PR and release Build jobs. > - The benefit is a production-ready runner package boundary with minimal build-process change. ## Linked Issues or Issue Description Refs #11962 This pull request replaces one bounded part of the archived large runner change. It follows the package-local authorization change in #12126. ## What Changed - Export only `@paperclipai/paperclip-runner` and `@paperclipai/paperclip-runner/testing`. - Keep Node-only fixture loading and semantic conformance helpers out of the runtime root. - Add a provider-neutral semantic conformance kit with stable JSON comparison and fail-closed input checks. - Keep deferred SDK, eval, browser, React, lab, and command surfaces private. - Pin the runner Rust toolchain to 1.97.1 with the minimal profile and `rustfmt`. - Run the Rust workspace tests in release mode. - Launch the optimized `paperclip-runnerd` and fake-harness binaries in process-level integration coverage. - Add one `pnpm --filter @paperclipai/paperclip-runner check:all` step to each existing PR and release Build job. - Make the existing server `prepack` lifecycle run its existing build after it prepares UI assets. - Document that no production adapter starts runnerd yet. This revision adds no standalone GitHub Actions job. It adds no server runner dependency or runner vendoring. It adds no Docker bootstrap or clean-consumer harness. It does not change `pnpm-lock.yaml`. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` - 66 TypeScript tests - 8 protocol contract tests - 56 Rust unit and integration tests - Release-mode integration coverage launches the optimized runnerd and fake-harness binaries. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-package-build-script.test.ts` (2 tests) - Clean `pnpm pack` from `server/` rebuilt the server and produced both `package/dist/index.js` and `package/dist/index.d.ts`. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` (8 tests) - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `git diff --check` - No `pnpm-lock.yaml` diff. - The diff changes 12 files. ## Risks The runner adds Rust work to the existing Build jobs. These jobs can take longer on a cold cache. The pinned toolchain makes contributor and CI behavior reproducible. Cargo tests use `--release` to verify optimized executables. The server prepack lifecycle now performs the build that its published entry points require. This can make direct server packing slower. This pull request does not wire runnerd into the server. It does not select runnerd for any adapter. Existing application execution and finalization paths remain unchanged. ## Model Used OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dc5b070709 |
fix(runtime): guard empty Bash 3.2 array expansion (#11891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in isolated worktrees with a separate Paperclip runtime. > - Runtime provisioning uses a Bash script on macOS hosts. > - macOS ships Bash 3.2, where an empty array expansion fails under `set -u`. > - The source-config argument array is empty when the base workspace already has a config. > - This pull request guards that expansion and tests the normal base-config path on Bash 3.2. > - The benefit is that managed worktree provisioning no longer fails before database seeding. ## Linked Issues or Issue Description No public GitHub issue exists for this problem. PR #11752 added the conditional source-config argument that exposed the failure. **What happened?** `scripts/provision-worktree-runtime.sh` expands an empty `source_config_args` array while `set -u` is active. Bash 3.2 reports `source_config_args[@]: unbound variable` and stops provisioning when the registered base workspace already has `.paperclip/config.json`. **Expected behavior** Runtime provisioning must call `worktree ensure-seeded` without a source override when the base workspace config exists. It must work with the Bash 3.2 version that macOS supplies. **Steps to reproduce** 1. Use macOS system Bash 3.2. 2. Create a base workspace with `.paperclip/config.json`. 3. Run `scripts/provision-worktree-runtime.sh` with `set -u` active in the script. 4. Observe the unbound-variable error before `worktree ensure-seeded` runs. **Paperclip version or commit** Reproduced on `origin/master` before this change. **Deployment mode** Local managed worktree runtime on macOS. ## What Changed - Guard all three optional source-config array expansions with Bash 3.2-compatible parameter expansion. - Add a regression test that uses the base-config path and verifies that no `--from-config` argument is sent. - Document the Bash 3.2 compatibility requirement in the runtime script. ## Verification - `/bin/bash -n scripts/provision-worktree-runtime.sh` - `node --test --test-name-pattern='runtime provisioning invokes ensure-seeded once|runtime provisioning omits the source override|runtime provisioning guards every optional source-config expansion' scripts/__tests__/provision-worktree-self-heal.test.mjs` - `git diff --check` ## Risks Low risk. The change only affects expansion of an optional two-element CLI argument array. The regression tests cover both the empty and non-empty paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model ID `gpt-5`. The runtime did not expose the context-window size. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9d1f740f0 |
fix(workspaces): seed managed worktrees when the base checkout has no config (#11752)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents do that work in isolated git worktrees, and a managed
worktree runs its own Paperclip instance with a cloned database
> - That clone needs a seed source, and the source must come from
server-owned registration, never from state the workspace itself can
rewrite
> - The seed-source resolver requires the registered base project
workspace to hold its own `.paperclip/config.json`
> - A managed project workspace is a plain `git clone`, and no code
writes that file into it
> - Every isolated worktree provision, deferred seed, and workspace
repair therefore fails on a managed checkout
> - This pull request lets a named source supply the config when the
base checkout has none
> - The benefit is that managed worktrees provision again, and the seed
source stays server-owned
## Linked Issues or Issue Description
No public GitHub issue exists for this problem. It is described below.
**What happened?**
Agent runs that need an isolated worktree fail during provisioning. The
provision command exits with this error (paths redacted):
```
Execution workspace provision command "bash ./scripts/provision-worktree.sh" failed:
Registered base project workspace has no canonical Paperclip config:
<instance-home>/instances/default/projects/<company-id>/<project-id>/<repo>/.paperclip/config.json
```
`resolveRegisteredWorktreeSeedSource` sets `registeredConfigPath` to
`<baseCwd>/.paperclip/config.json` whenever the caller names a
registered base workspace. It then requires that file to exist.
`scripts/provision-worktree.sh` applies the same rule.
A managed project workspace never has that file.
`materializeManagedProjectWorkspace` creates it with `git clone` and a
rename, so the checkout holds repository content only. The control plane
keeps its config at `<home>/instances/<id>/config.json` instead.
The failure reaches three paths: worktree provisioning, deferred seeding
through `worktree ensure-seeded`, and workspace repair.
The behavior changed in #11671. That pull request replaced a fallback
chain with a single hard requirement. Fixture code in
`scripts/__tests__/provision-worktree-self-heal.test.mjs` writes a
config into the fake base workspace, so tests kept passing.
**Expected behavior**
A managed worktree provisions and seeds from the registered source. The
seed manifest still never selects that source.
**Steps to reproduce**
1. Register the Paperclip repository as a project with a `repoUrl`, so
the server materializes a managed checkout.
2. Assign an issue to an agent whose workspace strategy is
`git_worktree`.
3. Watch the workspace operation log for the provision command.
4. The command exits non-zero with the error above.
**Paperclip version or commit**
Reproduced on `master` at
|
||
|
|
bd059a073d |
fix(workspaces): make managed runtimes reliable across restarts (#11740)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces need isolated databases, ports, and runtime services > - Concurrent workspaces could reuse ports or lose service ownership after a restart > - A markerless worktree also needed seed recovery, but normal markerless instances still needed to boot > - This pull request makes seed, port, and service ownership state explicit and recoverable > - It also checks live process and listener identity before it reclaims shared resources > - The benefit is reliable workspace startup, restart, adoption, and concurrent provisioning ## Linked Issues or Issue Description **What happened?** Managed workspaces could lose runtime service ownership after a control-plane restart. Concurrent worktrees could also reuse a port when their parent paths differed. A seed recovery change made every markerless instance resolve a worktree seed source, so normal instances without a source could not start. **Expected behavior** Paperclip must preserve healthy managed services across restarts. It must reserve unique ports across worktree parents. It must provision a registered markerless worktree, but it must skip seed work for a normal markerless instance. **Steps to reproduce** 1. Start two managed worktrees under different parent paths at the same time. 2. Restart the control plane while a managed service stays alive. 3. Start Paperclip with a config that has no seed markers and no registered worktree source. 4. Observe duplicate port selection, lost service adoption, or a seed-source startup error. **Paperclip version or commit** Current `master` plus the workspace runtime reliability changes in this pull request. **Deployment mode** Local development with managed execution workspaces and embedded Postgres. ## What Changed - Added a shared port registry with lease heartbeats, process identity checks, and live listener probes. - Reserved worktree ports across custom parent paths and repaired duplicate legacy assignments. - Preserved and adopted healthy managed services across control-plane restarts. - Reconciled guest bind modes and verified listener ownership before termination or reuse. - Provisioned registered markerless worktree databases and kept normal markerless instance startup as a no-op. - Added CLI, shared, server, and shell regression tests for seed, port, listener, restart, and adoption behavior. - Updated the worktree development documentation. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts --reporter=verbose` — 63 tests passed. - `pnpm exec vitest run packages/shared/src/worktree-port-registry.test.ts --reporter=verbose` — 5 tests passed. - Focused runtime Vitest set — 199 tests passed across 37 suites. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 10 tests passed. - `git diff --check` passed. ## Risks - Port reservation now depends on lease and process identity data. The fallback listener probe prevents early reclamation when process metadata is incomplete. - Runtime adoption is stricter about bind and owner identity. The tests cover healthy adoption, stale records, PID reuse, and unrelated listeners. - Markerless seed detection now separates registered worktrees from normal instances. The tests cover both paths. - There are no database schema migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
49217aadf0 |
refactor: balance serialized server shards by recorded suite duration (#11528)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ec984502a |
fix(release): stop smoke_beta silently skipping on promote-mode betas (#11582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system re-smokes every published beta as post-publish verification (`smoke_beta`) > - The candidate-branch beta lane (#11209) added `verify_beta_candidate` to `publish_beta`'s needs; that job is skipped on every normal promote-mode beta > - `smoke_beta`'s condition has no status-check function, so GitHub attaches an implicit `success()` that evaluates the needs chain transitively — a skipped ancestor makes it false > - This pull request makes the condition explicit so promote-mode betas smoke again, and pins the shape in the workflow wiring test > - The benefit is that the post-publish beta gate actually runs instead of silently skipping ## Linked Issues or Issue Description **What happened?** Beta `2026.818.0-beta.0` (run 32082007439) published successfully, but its post-publish `smoke_beta` job was skipped. No configuration or input asked for that: the run was a plain `channel: beta` dispatch with `dry_run` at its default `false`, and the same expression `!inputs.dry_run` evaluated true inside `publish_beta`'s own steps (the Docker dispatch step ran). **Expected behavior** Every non-dry-run beta publish is followed by the release smoke suite against the exact published version, as documented in `doc/RELEASING.md` and `doc/RELEASE-CHECKLIST.md`. **Steps to reproduce** Dispatch `release.yml` with `channel: beta` promoting a nightly (promote mode). `verify_beta_candidate` is skipped by design; `publish_beta` runs through its explicit `!cancelled()` condition; `smoke_beta` then skips because its implicit `success()` sees the skipped ancestor in the transitive needs chain (actions/runner#2205 semantics). The beta published on 2026-08-11 predated #11209, so this never surfaced before. **Paperclip version or commit** master at `43ab441f0` (workflow file, current head). Related (not duplicates): #11209 introduced the candidate lane whose skipped job triggers this; #11208 covers the adjacent tag-push failure playbooks. ## What Changed - `smoke_beta`'s condition becomes `!cancelled() && needs.publish_beta.result == 'success' && !inputs.dry_run` — an explicit status-check function suppresses the implicit `success()`, and the result check keeps the dependency on a successful publish. - A comment above the job records why the explicit form is load-bearing. - `scripts/__tests__/release-verify-workflow.test.mjs` pins the new shape so the implicit form cannot silently return. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 8 pass, including the new assertion. - `release.yml` re-parsed as YAML. - The exact skip is visible on run 32082007439 (`smoke_beta: skipped` after `publish_beta: success`); the coverage gap for that beta was closed manually by dispatching `release-smoke.yml` with `paperclip_version: beta` (run 32084880767). - Not exercised end-to-end: the corrected condition needs the next real promote-mode beta to demonstrate; the expression change is minimal and the semantics are the documented actions/runner behavior. ## Risks - Low risk: condition-only change on one job plus a test. Dry runs still skip the smoke (`!inputs.dry_run` retained). Candidate-mode betas, where `verify_beta_candidate` actually runs, behave as before. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
71e9d6bb0b |
feat(release): candidate-branch beta builds and the release checklist (#11209)
> Follow-up to #11208 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote artifacts along canary → nightly → beta → stable, with the happy path being promotion of an existing build > - When one or two targeted fixes are needed before a beta or stable, the only options today are waiting for the next nightly or absorbing a whole day of unrelated master changes > - The channel model was designed with an escape hatch for exactly this: short-lived candidate branches carrying only cherry-picked fixes > - This pull request implements candidate-branch beta builds with full verification, documents the stable fix path through the soak-justification gate, and adds the release captain's checklist > - The benefit is that a surgical fix can ship forward without either delay or blast radius, with its provenance recorded ## Linked Issues or Issue Description Refs #11008 — completes the fix-path half of the channel model introduced there. **Subsystem affected** Release automation: `scripts/release.sh`, `.github/workflows/release.yml`, `doc/RELEASING.md`, new `doc/RELEASE-CHECKLIST.md`, tests. **Problem or motivation** Beta promotion only accepts commits that already shipped as a nightly, and stable promotion expects a soaked beta. There is no supported way to ship one or two cherry-picked fixes between lanes: an urgent fix must wait for the nightly cycle or pull in every unrelated master change from the day. The original channel design called for candidate branches to cover this, and they were deferred from the initial implementation. **Proposed solution** Candidate-branch beta builds: cut `candidate/beta-<target>` from a nightly's source commit, cherry-pick the fixes, and dispatch `channel: beta` with the new `candidate_branch` input. Selection enforces the naming convention, rejects heads that already shipped as a beta or predate the candidate tooling, and records the cherry-picked commits in the job summary. Because candidate heads never went through a canary or nightly, publication is gated on a full `release-verify` run (promoted nightlies keep skipping re-verification). The stable fix path (`candidate/release-<target>` as `source_ref`) works through the existing soak gate: the justification requirement is the deliberate, recorded trade-off for shipping unsoaked bits, and is now documented as such. ## What Changed - `scripts/release.sh`: `--from-candidate` flag (beta only) waives the shipped-a-nightly requirement while keeping the duplicate-beta guard - `.github/workflows/release.yml`: `candidate_branch` dispatch input; candidate mode in `select_beta` (naming validation, duplicate and tooling-era rejection, cherry-pick recording); new `verify_beta_candidate` job gating candidate publishes on full verification - `doc/RELEASING.md`: beta fix-path and stable fix-path sections - `doc/RELEASE-CHECKLIST.md` (new): the release captain's checklist for all four lanes as built - Tests: dry-run fixture coverage for `--from-candidate` (waives the nightly guard, keeps the duplicate guard, rejected outside beta) and wiring tests for candidate validation plus the verification gate ## Verification - `node --test` on the four affected suites: 42 pass in total (17 + 25 across the two runs), including the 5 new tests - `bash -n` on `release.sh`; YAML parse of the workflow - After merge: exercise the path end to end the first time a real cherry-picked beta is needed — dispatch with a `candidate/beta-*` branch and confirm the summary records the picks and verification runs ## Risks - Candidate builds bypass the smoke-tested-nightly provenance by design; the compensating controls are full verification before publish, the post-publish beta smoke, the human `npm-beta` gate, and recorded cherry-picks - The stable fix path rides the existing justification mechanism rather than adding a second bypass — one recorded escape hatch, not two ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2da6a248c3 |
fix(release): surface recovery commands when a lane tag push is rejected (#11208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's promotion lanes publish to npm, then push a lane tag and dispatch the Docker image build at that tag > - The first nightly of the beta-tooling merge published to npm and then died at the tag push: GITHUB_TOKEN may not create refs pointing at workflow-modifying commits from dispatch or scheduled runs > - The failure was a bare `remote rejected` with no guidance, leaving the release half-finished (npm live, no tag, no images) until an operator reverse-engineered the recovery > - This pull request makes every lane's tag push degrade into exact recovery instructions in the job summary > - The benefit is that a rare platform-permission rejection becomes a two-minute runbook operation instead of a forensic exercise ## Linked Issues or Issue Description Refs #11008 — the incident occurred promoting that change's own merge commit, the first workflow-modifying commit to flow through the lanes it introduced. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31445344811 published `2026.811.0-nightly.0` to npm, then failed pushing `nightly/v2026.811.0-nightly.0`: `refusing to allow a GitHub App to create or update workflow .github/workflows/release.yml without workflows permission`. The tagged commit modifies workflow files, and GITHUB_TOKEN may not create refs pointing at such commits from dispatch or scheduled runs (push-event runs are exempt, which is why the canary tag on the same commit succeeded). The job failed with no explanation and the Docker dispatch never ran. **Proposed solution** Wrap the nightly, beta, and stable tag pushes: on rejection, write the exact recovery commands into the job summary — create and push the tag with maintainer credentials, dispatch `docker.yml` at the tag, and for stable also run `create-github-release.sh` — then fail the job. Document the cause and recovery in the failure playbooks and pin the three recovery blocks with a wiring test. ## What Changed - `.github/workflows/release.yml`: recovery-summary wrappers on the nightly, beta, and stable tag-push steps - `doc/RELEASING.md`: failure-playbook entry for the workflows-permission rejection - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting all three lanes carry the recovery summary ## Verification - Wiring tests: 6 pass; YAML parse of the workflow - The recovery commands are exactly the ones used to resolve the real incident (tag push + `docker.yml` dispatch for `nightly/v2026.811.0-nightly.0`) ## Risks - Low. The happy path is unchanged (a successful push skips the wrapper); the failure path trades a bare error for actionable output and still fails the job, since the release state is genuinely incomplete ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6601014898 |
fix(release): reject promotion sources that predate their channel tooling (#11197)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem promotes builds along canary → nightly → beta → stable, and each publish job checks out the promotion's source commit and runs that tree's release tooling > - The first beta dispatch failed with `unexpected argument: beta`: the selected nightly's source predated the beta channel, so its `release.sh` did not know the argument > - The failure was clean (argument parsing, nothing published) but cryptic, and the same trap waits for any promotion of a source older than its target channel's tooling > - This pull request makes the selection jobs reject such sources with an actionable error and documents the property > - The benefit is that a bootstrapping or old-source promotion fails in seconds with instructions, instead of mid-publish with a parser error ## Linked Issues or Issue Description Refs #11008 — the guard hardens the beta promotion flow introduced there, after its first dispatch surfaced the gap described below. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31444045044 (first beta dispatch) failed in `publish_beta` with `unexpected argument: beta`. Promotions deliberately build from the pinned source commit, which means they also run that commit's `scripts/release.sh` — and a source that predates the target channel's introduction cannot publish it. Nothing guards this today; the error surfaces deep in the publish job with no explanation. **Proposed solution** Guard at selection time: `select_nightly` requires the source canary's `release.sh` to know the nightly channel, and `select_beta` requires the source nightly's `release.sh` to know the beta channel. Each guard literally matches the channel case arm and fails closed with a clear message naming the remedy (promote a newer source). Document the tooling-era property in `RELEASING.md` and pin the guards with a wiring test. ## What Changed - `.github/workflows/release.yml`: tooling-era guards in `select_nightly` and `select_beta` - `doc/RELEASING.md`: documents that promotions run the source commit's release tooling - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test pinning both guards ## Verification - Wiring tests: 5 pass - Guard expressions exercised against real commits: accepts the beta-capable merge commit of the beta-channel change, rejects a pre-beta commit - YAML parse of the workflow - After merge: the next beta dispatch selects a beta-capable nightly and passes the guard ## Risks - Low. Selection-time check only; the guards match the channel case arm literally and fail closed (with the same actionable message) if that line is ever reformatted ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f94f6003c6 |
fix(release-smoke): pin the smoke container to the lan bind preset (#11189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly on the release smoke suite, which boots the published artifact in a Docker container and drives real onboarding > - The gate kept failing even after the readiness budget fix (#11187), and the new container-log dump revealed the server was healthy but listening on 127.0.0.1 inside the container, unreachable through Docker's port mapping > - `onboard --yes` without an explicit `--bind` prefers trusted-local quickstart defaults: it writes a loopback bind into the instance config and ignores the deployment env vars the harness passes, and that config outranks `HOST` at runtime > - This pull request pins the smoke container to the `lan` bind preset and adds a wiring test for it > - The benefit is a working nightly gate, verified end to end against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `docker/Dockerfile.onboard-smoke`, `scripts/__tests__/release-verify-workflow.test.mjs`. **Problem or motivation** Nightly run 31428558684 failed in smoke with the server unreachable at the mapped port for the full 420 second budget. The container logs (captured thanks to #11187) show a fully booted server with `Bind loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the deployment env vars, but `onboard --yes` without `--bind` deliberately prefers trusted-local defaults, writes `bind: loopback` into the instance config, and the config outranks `HOST` at runtime. A loopback listener inside a container is invisible to the port mapping, so the health check can never pass. This behavior predates the current stable, so the harness was silently broken against every recent version — it only surfaced now because the nightly lane is the suite's first CI consumer. **Proposed solution** Pass `--bind lan` in the smoke container command (the flag is supported by `latest` and canary alike; it selects the all-interfaces preset and keeps the env-driven authenticated deployment), and pin the flag with a wiring test so it cannot regress silently. ## What Changed - `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard --yes --bind lan --data-dir ...`, with a comment explaining why the flag is load-bearing - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting the smoke Dockerfile pins a non-loopback bind preset ## Verification - Full local harness run against the real nightly candidate `2026.810.0-canary.1`: container healthy, bind banner shows `lan (0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap invite accepted, board session verified), `/api/health` returns `bootstrapStatus: ready` - `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4 pass - After merge: dispatch `release.yml` with `channel: nightly` to run the gate end to end in CI ## Risks - Low. The change only affects the smoke container. `--bind lan` inside a container exposes the port to the container network only; reachability from outside still goes through Docker's explicit port mapping ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI log forensics, upstream source tracing, local Docker reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
30f6999cbe |
fix(release-smoke): configurable readiness timeout and diagnostics for slow containers (#11187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane (#11006) gates every nightly publish on the release smoke suite, which boots the published artifact in a Docker container > - The suite's first CI execution failed at the health readiness check: the harness hard-codes a 90 second budget, but a CI container cold-installs paperclipai from npm and initializes embedded postgres with no warm caches > - When the timeout expired with the container still running, the harness printed no container logs, so the failure gave no diagnostics > - This pull request makes the readiness budget configurable, raises it for CI, and dumps container logs on timeout > - The benefit is that the nightly gate measures the artifact, not the runner's cold caches, and a red smoke run is diagnosable from its logs ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `scripts/docker-onboard-smoke.sh`, `.github/workflows/release-smoke.yml`. **Problem or motivation** Run 31426044332 (first forced nightly after #11006) failed in `smoke_nightly` with `server did not become ready at http://localhost:3232/api/health` after exactly 90 seconds. The harness's readiness window is hard-coded to 90 attempts at 1 second. Locally that works because the npm cache is warm; in CI the container downloads the full package set and embedded postgres first. The timeout path also printed no container logs when the container was still running, so there was no way to see how far boot had progressed. **Proposed solution** Make the readiness budget an environment variable (`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use), set it to 420 in the CI workflow, and dump the last 150 container log lines when the readiness check times out on a still-running container. ## What Changed - `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env var (default 90) replaces the hard-coded readiness budget; timeout with a still-running container now prints the tail of `docker logs` - `.github/workflows/release-smoke.yml`: sets `SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs ## Verification - `bash -n` on the harness and YAML parse of the workflow - The real proof is the next `channel: nightly` dispatch of `release.yml`, which re-runs this suite in CI with the new budget ## Risks - Low. The local default is unchanged; CI runs simply wait longer before declaring failure, and a genuinely broken artifact still fails (with logs now) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. Diagnosis from CI run logs; patch model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f9173782cd |
feat(release): add smoke-gated nightly channel and lane-separated Docker tags (#11006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem publishes the `paperclipai` npm package set and the Docker images on two lanes: canary on every master push, and stable on manual promotion > - There is no middle ground between those lanes. Users must track every merge or wait weeks for a stable. Docker `:latest` also tracks master, so Docker users have no stable image at all > - A calm prerelease lane needs to exist, and it must never ship a build that failed its checks > - This pull request adds the nightly channel: a scheduled job that selects the newest master commit with a green canary publish, runs the full release smoke suite against that exact published canary, and only then republishes it as the nightly. It also separates Docker tags by lane, so `:latest` finally means stable > - The benefit is that users can follow prereleases at a nightly cadence with a smoke-tested guarantee, and Docker users get real `:canary`, `:nightly`, and stable image tags ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** The project publishes only `canary` (every master push) and `latest` (manual stable). Users who want prereleases without per-merge churn have no option. Docker has a second problem: master builds overwrite `:latest`, and CI-published stables never produced Docker images, because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger in `docker.yml`. No stable-versioned image exists in ghcr today. **Proposed solution** Add a `nightly` channel. A scheduled job selects the newest canary-tagged master commit, smoke-tests that exact published canary, and republishes the same commit as `YYYY.MDD.P-nightly.N` under the `nightly` dist-tag. Separate Docker tags by lane (`:canary` for master, `:nightly` for nightly tags, `:latest` plus version tags for stable tags only), and have the release jobs dispatch `docker.yml` at the new tag so lane images actually build. **Alternatives considered** Moving the `nightly` dist-tag to the existing canary version without a republish. Rejected: the version string would say `canary` while the user is on nightly, which breaks at-a-glance lane identification in bug reports and `--version` output. ## What Changed - `scripts/release-lib.sh`: channel-parameterized `next_prerelease_version` and `prerelease_tag_name` helpers (canary helpers delegate to them), a `require_channel_tag_at_head` guard, and the no-provenance retry for Sigstore transparency-log duplicates now covers the `nightly` dist-tag as well as `canary` - `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry a `canary/v*` tag, publishes the full public package set as `YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source commit `nightly/vYYYY.MDD.P-nightly.N` - `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) — select candidate, smoke it via `release-smoke.yml`, publish on green under the existing `npm-canary` environment, push the tag, dispatch `docker.yml`. New `channel` dispatch input (default `stable`, so existing stable dispatches are unchanged) with `nightly_source_version` and `dry_run` support for forced runs. The stable path now also dispatches `docker.yml` at the new `v*` tag - `.github/workflows/docker.yml`: lane tag mapping for both image jobs — master pushes publish `:canary` and no longer move `:latest`; `nightly/v*` tags publish `:nightly`; only stable `v*` tags publish `:latest` and the versioned tags. New `workflow_dispatch` trigger for the release-job dispatches. Build-version stamping uses the exact nightly version on nightly tag builds - `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch choice list - `doc/CHANNELS.md` (new): user-facing guide to the channels - `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping table, and a nightly failure playbook - `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses `npm-canary` and needs no npm trusted-publisher changes - Tests: channel-parameterized version helper coverage in `scripts/release-registry-versions.test.mjs`, and nightly flow coverage (publish identity, notes not required, canary-tag guard) in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the release script suites: 68 pass, including 6 new tests. The only failure, `acpx-patch-packaging.test.mjs`, needs installed `node_modules` and fails identically on a pristine checkout of master in the same environment - `bash -n` on both shell scripts and YAML parse of all three workflows - Live fail-path check: `./scripts/release.sh nightly --print-version` from a master tip with no canary tag fails with `HEAD has no canary/v* tag` - Live success-path check: the same command from the `canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0` - Live selection check: the candidate-selection shell logic run against the real repository selects the commit of `canary/v2026.806.0-canary.7`, which matches the current npm `canary` dist-tag exactly - After merge: dispatch `release.yml` with `channel: nightly` and `dry_run: true` to preview, then a real forced run to validate end to end before the first scheduled run ## Risks - Docker `:latest` changes meaning from "latest master build" to "latest stable release". This is deliberate and will be announced. Users who want the old behavior pull `:canary`. Until the first stable release after this change, `:latest` stays at its current (master-built) image - The nightly is a rebuild of the same source commit, not the byte-identical canary artifact that was smoked. The lockfile pins dependencies, and the publish path's registry-visibility and clean-prefix install gates still run on the nightly artifacts - All npm publishing must stay inside `release.yml` because npm trusted publishing pins that workflow file per package. The nightly jobs were added to `release.yml` for exactly that reason; this constraint is now documented in `RELEASING.md` - The stable-lane Docker dispatch fails gracefully (a warning with manual instructions) when the source ref predates `docker.yml`'s `workflow_dispatch` trigger ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
00a24d7e8f |
ci: split general-server tests into five shards with refreshed durations (#10925)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The PR workflow runs the server vitest suite across sharded runners because the suite is pinned to one worker. > - In successful PR run 30930345729 (2026-08-04), shard `server (3/4)` took 311 seconds of wall time and was the slowest check in the run. > - The suite has grown to about 946 seconds of serial vitest time, but the duration manifest was last sampled on 2026-08-01 at about 882 seconds. > - This pull request refreshes the per-suite duration manifest from that run's logs and splits the lane into five shards. > - The benefit is a shorter PR critical path: each shard carries about 196 seconds of suite time, level with the other lanes. ## Linked Issues or Issue Description Refs #10663 (previous split of this lane into four shards). Related: #10923 splits the separate serialized-suites lane into five shards. Both PRs touch `.github/workflows/pr.yml` in different matrix blocks; whichever merges second needs a trivial rebase. **What existing behavior does this improve?** The `general-server` vitest lane runs in four shards with a duration manifest sampled on 2026-08-01. **Current behavior** In PR run 30930345729, shard 3/4 ran for 311 seconds (273 seconds in the test step) and was the longest check in the run. The suite now totals about 946 seconds of serial vitest time. **Proposed behavior** Run the same suite set in five shards, balanced with a per-suite duration manifest refreshed from that run's shard logs (279 suites measured by diffing consecutive completion timestamps). **Reason and benefit** The refreshed LPT partition balances at about 196 seconds of suite time per shard (about 240 seconds per job), level with the other PR lanes. No test coverage is lost. **Breaking changes** None. The change only alters the CI partition size and the duration manifest. ## What Changed - Bump the `general-server` shard matrix in `.github/workflows/pr.yml` from four to five shards. - Refresh `scripts/general-server-shard-durations.json` from the 2026-08-04 run's shard logs. - Update `SHARD_COUNT` in `scripts/__tests__/run-vitest-stable-shard.test.mjs` to five. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass, including the complete non-overlapping partition proof and the duration-balance check. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass. - `node --test scripts/__tests__/e2e-shard.test.mjs` — 7/7 pass. - A 5-way dry-run partition covers all suites exactly once with equal projected weights. ## Risks - Low risk. The change only alters CI partition size and duration weights; the suite set is unchanged. - One more runner is used per PR run for this lane. - Stale duration weights degrade gracefully: suites missing from the manifest get the median weight. ## Model Used - Claude (Anthropic), Claude Code CLI, model ID `claude-fable-5`, extended thinking with tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (workflow comments explain the new shard math) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude <claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b6e58019f2 |
ci: split serialized tests into five shards (#10923)
## Thinking Path > - Paperclip uses CI to keep control-plane changes safe and mergeable. > - The PR workflow splits serialized server tests across isolated runners. > - A recent successful run spent 305 seconds in serialized shard 2/4. > - That job was the slowest check in the run. > - The four shards reported about 739 seconds of Vitest suite time. > - This pull request adds a fifth serialized shard and keeps release verification aligned. > - The benefit is a shorter PR critical path with no loss of test coverage. ## Linked Issues or Issue Description **What existing behavior does this improve?** The PR and release verification workflows run serialized server tests in four shards. **Current behavior** Successful PR run 30876682788 spent 305 seconds in `Verify serialized server suites (2/4)`. The test step used 256 seconds and made this job the slowest check. **Proposed behavior** Run the same serialized suite set in five complete and non-overlapping shards. **Reason and benefit** The measured suites reported about 739 seconds of total Vitest time. Five runners reduce the expected average suite time from about 185 seconds to about 148 seconds before setup overhead. **Breaking changes** None. The change only alters CI partition size. ## What Changed - Split serialized server tests into five shards in the PR workflow. - Apply the same five-shard layout to release verification. - Add a partition test that proves complete and non-overlapping serialized coverage. - Update release workflow coverage tests for five shards. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` - `git diff --check` ## Risks - Low risk. CI uses one additional runner for the serialized lane. - Round-robin partition weights can still vary as suite timings change. > This change does not overlap with planned core work in `ROADMAP.md`. Related PR #10663 optimized the separate general-server lane. ## Model Used - OpenAI Codex, GPT-5, agentic coding with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Devin Foley <139239+devinfoley@users.noreply.github.com> |
||
|
|
dcac49a4fd |
feat(workspaces): defer isolated setup until runtime start (#10653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Isolated workspaces give each task a safe and reproducible checkout. > - The existing setup cloned the development database before an agent needed to run the app. > - This made worktree creation slower and heavier for tasks that never start a service. > - Runtime services already use one server start path for heartbeat, operator, and startup recovery flows. > - This pull request moves heavy setup to that start path and keeps worktree creation lean. > - The benefit is faster isolated workspace creation with the same reliable runtime setup when a service starts. ## Linked Issues or Issue Description Related pull request: #10652 covers the initial deferred database-seeding slice. This pull request supersedes it with end-to-end runtime provisioning and safe cleanup. **What existing behavior does this improve?** This improves isolated worktree creation, runtime service startup, and isolated instance cleanup. **Subsystem affected** Cross-cutting: CLI worktree setup, server runtime orchestration, shared workspace contracts, and development scripts. **Current behavior** Paperclip seeds an isolated development database during worktree creation. It can also leave an isolated instance directory after workspace teardown. This work happens even when no runtime service starts. **Proposed behavior** Paperclip creates the worktree with a lean eager setup. It runs an idempotent runtime provision command before the first managed service spawn. Concurrent starts share one provision attempt. Teardown removes the isolated instance safely. **Reason and benefit** Many agent tasks only edit and test code. They do not need a running Paperclip instance. Deferring the database seed reduces workspace startup cost while preserving automatic setup for tasks that start the app. **Breaking changes** None. The new runtime provision command is optional. Existing workspace behavior is unchanged when it is absent. ## What Changed - Split Paperclip worktree setup into a lean eager script and an idempotent runtime provision script. - Added `runtimeProvisionCommand` to project, issue, realized workspace, and persisted workspace contracts. - Added a per-workspace provision mutex before local service spawn for heartbeat, operator, and startup recovery flows. - Added a persisted `provisioning` service state and the `workspace_runtime_provision` operation phase. - Kept provision time outside the service readiness timeout and made failed attempts visible and retryable. - Reclaimed isolated instance data during safe workspace teardown. - Serialized deferred database seeding across processes and bound teardown to the instance root captured in persisted workspace metadata. - Added tests for config flow, concurrency, retry, no-op behavior, readiness timing, scripts, CLI commands, and cleanup. - Documented the eager and runtime provisioning contracts. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run` (server: 3,201 passed; UI: 3,345 passed; the CLI phase exposed one environment-sensitive AWS doctor assertion because the agent runtime injects static AWS credentials) - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts -t 'passes AWS doctor checks when non-secret provider config is present'` - Focused runtime tests cover serialized provisioning, retry after stderr failure, absent-command no-op behavior, operation logging, persisted state order, and readiness timeout exclusion. - Focused CLI and cleanup tests cover concurrent seed serialization, stale-lock fail-closed behavior, persisted instance ownership, and rewritten sibling pointers. ## Risks - A faulty runtime provision script blocks service startup. Paperclip records stderr, marks the service failed, and retries on the next start. - Concurrent service requests share an in-process provision attempt, while the seed command uses an atomic filesystem lock across processes. A stale lock fails closed and requires an operator to verify no seed is running before removing it. - Isolated instance cleanup is destructive. The cleanup service validates ownership and path containment before removal. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5.6-sol`, with agentic reasoning, tool use, and code execution. The service does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8540ce2973 |
ci: shard general-server tests 4 ways (#10663)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI must give contributors fast and stable feedback. > - The `general-server` Vitest lane runs many single-worker server suites. > - A recent completed PR run showed this lane as the slowest completed check. > - Three shards still left one runner with the largest share of work. > - This pull request splits that lane into four duration-balanced shards. > - The benefit is a shorter critical path for the same server test coverage. ## Linked Issues or Issue Description No public GitHub issue exists for this CI maintenance change. **Pre-submission checklist** - I confirmed this improves existing behavior. It does not add a new command, endpoint, or concept. - I searched open public issues and pull requests for related CI sharding work. **What existing behavior does this improve?** The pull request workflow's `general-server` Vitest lane. **Subsystem affected** Cross-cutting. This affects GitHub Actions CI and the Vitest shard duration manifest. **Current behavior** The `general-server` lane uses three shards. The server suites now total about 880 seconds of serial Vitest wall time. The slowest shard was about 313 seconds in the measured run. **Proposed behavior** The `general-server` lane uses four shards. Each shard receives about 220 seconds of predicted suite weight from the refreshed duration manifest. **Reason and benefit** The slowest PR check controls how soon a reviewer can trust the PR. Four balanced shards reduce the slowest `general-server` shard while keeping the same suite selection rules. **Breaking changes** None. This only changes CI partitioning and duration data for existing test suites. **Additional context** Related public searches found no exact open issue or pull request for this `general-server` sharding change. ## What Changed - Split the `general-server` CI matrix from three shards to four shards. - Refreshed `scripts/general-server-shard-durations.json` with wall-time weights from a recent completed PR run. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check origin/master...HEAD` - Dry-ran the four `general-server` shards locally during implementation. The partition covers 300 unique suites with about 220.56 seconds of predicted weight per shard. - Ran a local sensitive-data scan before push. It found only test filenames that contain words such as `secret` or `token`, not credential values. ## Risks Low risk. The main risk is that the duration manifest becomes stale as suite costs move. Missing suites fall back to the median weight, so the lane still runs if the manifest is incomplete. ## Model Used OpenAI Codex, GPT-5, with tool use and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8444e5735c |
ci: pass e2e shard specs without separator (#10640)
## Thinking Path > - Paperclip uses pull request CI to test changes before merge. > - The e2e PR lane runs Playwright specs in a shard matrix. > - Each shard builds a list of spec files for its matrix entry. > - The workflow passed that list after a literal `--` separator. > - Playwright did not receive the list as file filters. > - This pull request removes the separator and adds a guard test. > - The benefit is that each e2e shard runs only its assigned specs. ## Linked Issues or Issue Description Refs #10629. **What happened?** The e2e shard step used `pnpm run test:e2e -- $specs`. The shard spec list was not applied as Playwright file filters. **Expected behavior** Each e2e shard should pass only its selected specs to Playwright. **Steps to reproduce** 1. Inspect `.github/workflows/pr.yml` at the merge commit for #10629. 2. Find the `e2e_shards` command that invokes `pnpm run test:e2e`. 3. See the literal `--` before `$specs`. **Paperclip version or commit** `86767951` **Deployment mode** GitHub Actions PR CI. ## What Changed - Removed the literal `--` from the e2e shard `pnpm run test:e2e $specs` invocation. - Added a regression test that checks the workflow passes `$specs` without that separator. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks Low risk. This changes one CI command and one workflow guard test. The main risk is shell argument handling in the workflow, and the guard now covers the expected command shape. ## Model Used OpenAI GPT-5 through Codex. The run used shell and GitHub CLI tool access. The runtime did not expose a context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8676795188 |
ci: split e2e PR lane into three shards (#10629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow protects changes with a Playwright e2e lane. > - That lane already uses a weighted file partition so slow specs do not cluster by test count. > - Recent green PR runs showed the two e2e shard jobs were slower than the next slow required lane. > - The largest spec is indivisible, so a third shard lets that spec run alone and lets the rest split by duration. > - This pull request changes only the PR e2e shard matrix and the guard test. > - The benefit is a shorter expected PR critical path while the required `e2e` aggregate check name stays stable. ## Linked Issues or Issue Description Refs #9923 **What existing behavior does this improve?** The `pull_request` workflow Playwright e2e lane. **Subsystem affected** Cross-cutting: GitHub Actions CI and test scripts. **Current behavior** The PR workflow runs the weighted Playwright e2e partition across two jobs. Recent green runs showed those jobs as the slowest required checks. **Proposed behavior** The PR workflow runs the same e2e spec set across three weighted jobs. The aggregate required check stays named `e2e`. **Reason and benefit** The third shard lets the slow smoke-lab spec run alone while the rest of the catalog stays balanced. This should shorten the PR critical path. The win is bounded by fixed per-job setup time. **Breaking changes** None. The required aggregate check contract is preserved. ## What Changed - Change the PR e2e shard matrix from two entries to three entries. - Update the shard guard test to expect three shards. - Floor the balance bound at the largest single spec weight. - Assert that the workflow does not define more shard indexes than `SHARD_COUNT`. ## Verification - `node --test ./scripts/__tests__/e2e-shard.test.mjs` passes with 6 tests. - The recorded-weight partition is complete and non-overlapping: 168.0s, 116.5s, and 114.4s. - I checked `ROADMAP.md` and found no overlapping roadmap-level core feature. - I searched public GitHub PRs and issues for related e2e shard work. I found related PR #9923 and no open duplicate for this branch or change. ## Risks - This adds one extra GitHub Actions runner to the PR e2e lane. - The wall-clock win is bounded by fixed per-job setup. - Behavior risk is low because the aggregate required check remains named `e2e`. ## Model Used OpenAI Codex, GPT-5, tool-enabled coding agent in this repository. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Cody <noreply@paperclip.ing> |
||
|
|
c9116686bd |
test(installer): cover cross-version update migrations (#10587)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Managed updates can change both the application payload and its database schema > - Unit tests cannot prove that an older live install upgrades through a real migration and remains recoverable > - The managed-install work in #10045 needs a repeatable cross-version system test > - This pull request adds an isolated end-to-end harness for update, migration, backup, service restart, and rollback behavior > - The benefit is a direct proof that managed upgrades preserve the existing database and service lifecycle across versions ## Linked Issues or Issue Description Refs #10045 This test is a focused follow-up to the managed install integration. Merge #10045 first so the tested install, update, service, backup, and rollback commands are available. ## What Changed - Added a cross-version managed-update E2E script. - Installed an older Git ref, initialized its embedded PostgreSQL database, and updated to a ref with one additional migration. - Verified the pre-update backup, payload switch, service recovery, migration result, database-cluster reuse, and rollback behavior. - Isolated Paperclip state under a dedicated test home and cleaned up the service and managed install on success or failure. - Added regression tests for shell syntax, required-ref validation, side-effect-free preflight failure, and complete failure cleanup. ## Verification - `node --test scripts/__tests__/e2e-update-migrations.test.mjs` - `bash -n scripts/e2e-update-migrations.sh` - GitHub latest-head CI: build, typecheck, release registry, canary dry-run, general tests, serialized suites, and both browser E2E shards passed. - Full harness execution needs an isolated macOS or Linux host with a real launchd or systemd user service. It is intentionally not run on a live Paperclip server host. ## Risks - The script manages a real user service and downloads two Git refs. Run it only on an isolated test host. - The test needs #10045 because `origin/master` does not yet contain the managed install lifecycle. - The script uses a dedicated `PAPERCLIP_HOME`, refuses a pre-existing shim or test home, and removes its service and install during cleanup. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The runtime did not expose a more specific deployment ID or context-window size. Reasoning, repository access, shell execution, and GitHub tooling were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
79eff0aea1 |
fix(scripts): self-heal isolated workspace provisioning when the base CLI is broken (#10574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run tasks in isolated execution workspaces that are provisioned as git worktrees by `scripts/provision-worktree.sh` > - The script runs the base workspace's CLI (`cli/src/index.ts` via the base `tsx` install) to seed each new worktree, and it only checked that those files exist > - pnpm links each package's `node_modules` into a hash-versioned virtual store; a lockfile change followed by a partial or filtered install prunes old hashed dirs without relinking every package, leaving dangling symlinks > - A CLI with dangling symlinks fails ESM resolution (`ERR_MODULE_NOT_FOUND`) at boot, so provisioning aborts with `setup_failed` — deterministically, on every retry, with no self-heal path > - This pull request makes provisioning health-check the CLI by actually booting it, repair the base install when the check fails, and degrade to the no-CLI fallback config instead of failing the run > - The benefit is that a class of permanent `setup_failed` loops becomes self-healing, and workspace provisioning survives a broken base CLI ## Linked Issues or Issue Description No public GitHub issue exists; the underlying bug is described here per `bug_report.yml`. Related PR: #10578 self-heals the sibling workspace-validation failure loop uncovered by the same incident diagnosis. **What happened?** Isolated-workspace runs failed at provision time with `setup_failed`. Every retry failed identically. One observed incident burned 4 runs across two adapters before the task was stranded. **Expected behavior** Provisioning either succeeds or degrades gracefully; a broken base CLI install repairs itself instead of permanently blocking all new worktrees. **Steps to reproduce** In the base workspace, cause a lockfile-affecting dependency bump plus a partial/filtered `pnpm install` so a package symlink (e.g. `cli/node_modules/drizzle-orm`) dangles into a pruned virtual-store dir. Start any isolated-workspace run. Provision fails with `ERR_MODULE_NOT_FOUND` and the run ends `setup_failed`; retries never recover. **Paperclip version or commit** master as of the branch point of this PR. **Deployment mode** Local trusted deployment with git-worktree isolated workspaces. ## What Changed - `base_cli_healthy` now boots the base CLI (`--help`) instead of only testing file existence, which exercises the top-level import graph. - New `repair_base_workspace_install`: when the health check fails, run a non-interactive `pnpm install --prod=false --force --frozen-lockfile` in the base workspace. `--force` guarantees relinking when pnpm's up-to-date heuristics would skip dangling symlinks; `--frozen-lockfile` keeps the repair from mutating the shared lockfile. - The repair install is serialized with `flock` on a lock file inside the resolved git dir (`git rev-parse --absolute-git-dir`), so locking also covers base workspaces that are linked worktrees, where `.git` is a file. - If every CLI candidate is unusable (including a base CLI the repair could not fix), provisioning falls back to the existing no-CLI fallback config writer (loudly, on stderr) instead of failing the run. A CLI that runs and fails `worktree init` still fails provisioning with its real exit code — that deliberate fail-closed policy is unchanged and covered by an existing server regression test. - Fixed a latent bug: `run_isolated_worktree_init` returned 0 unconditionally after the init subshell, so callers treated a failed init as success. Exit codes now propagate. ## Verification - Reproduced the incident state (dangling `cli/node_modules/drizzle-orm` symlink); the base CLI failed with the exact `ERR_MODULE_NOT_FOUND` seen in the incident run logs. - Ran the patched script against a fresh scratch worktree: health check failed → locked repair install ran (~26 s warm) → symlink relinked → `worktree init` completed → exit 0 with `.paperclip/config.json` and `.env` written. - Happy path (healthy base CLI): provisioning behavior unchanged, exit 0. - Verified `git rev-parse --absolute-git-dir` resolves a real directory for both a normal checkout and a linked worktree. - New hermetic tests: `node --test ./scripts/__tests__/provision-worktree-self-heal.test.mjs` (4 tests: healthy CLI used, broken CLI degrades, locked repair end-to-end with a fake pnpm, init failure propagates). Not yet wired into a CI workflow. - `server`: the existing `realizeExecutionWorkspace` fail-closed regression test ("fails instead of writing an unseeded fallback config when worktree init errors after CLI detection succeeds") passes against the new script. - `bash -n scripts/provision-worktree.sh` is clean. ## Risks - Low risk overall: the script only adds recovery paths; the happy path is unchanged. - The repair install runs in the shared base workspace. It is bounded by `--frozen-lockfile` (no lockfile mutation) and serialized by `flock`, but it can add ~30 s to the first provision after a base install breaks. - If the repair cannot fix the CLI and no other CLI candidate exists, runs now continue with an unseeded fallback config instead of failing; that is intentional, and the fallback path already existed. Genuine `worktree init` failures from a working CLI still fail the run. ## Model Used Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use (Claude Code harness). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1944c86153 |
fix(ci): preserve required e2e check for sharded runs (#9923)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The pull request workflow is the main merge gate for changes to that app. > - The Playwright e2e lane is expensive because every spec shares one isolated server and runs serially. > - Splitting that lane across runners shortens the critical path, but the public required-check contract still needs a check named exactly `e2e`. > - This pull request shards the real e2e work while preserving a fast aggregate `e2e` job for branch protection. > - The benefit is a faster PR workflow without making otherwise-good PRs unmergeable because a legacy required check disappeared. ## Linked Issues or Issue Description No public GitHub issue exists for this CI follow-up. Related prior CI work: - Refs #8360 - Refs #9168 - Refs #9516 Bug report: ### What happened? Sharding the PR e2e lane directly at the workflow job level changes the emitted check names to shard-specific names, while existing branch protection expects a check named exactly `e2e`. ### Expected behavior The PR workflow should be able to run e2e specs across multiple runners while still emitting a stable aggregate check named `e2e`. ### Steps to reproduce 1. Open a PR against `master`. 2. Run the PR workflow with the e2e lane split only as a matrix job. 3. Observe that the shard checks complete, but a required check named exactly `e2e` never appears. ### Paperclip version or commit Current `master`. ### Deployment mode GitHub Actions pull request workflow. ## What Changed - Added `scripts/e2e-shard.mjs`, which partitions default Playwright e2e specs by recorded per-spec duration. - Added `scripts/e2e-shard-durations.json` with measured e2e spec durations so the slow smoke-lab spec does not dominate one runner. - Split the PR workflow e2e lane into two `e2e_shards` matrix jobs and added a fast aggregate job named exactly `e2e`. - Added `scripts/__tests__/e2e-shard.test.mjs` to lock the shard partition, ignored-spec sync, manifest coverage, and aggregate required-check contract. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check upstream/master..HEAD` - Searched GitHub for duplicate or related e2e-shard / required-check PRs and issues before opening this PR; no direct duplicate was found. ## Risks Low risk. The main risk is that the duration manifest can drift as specs are added or runtimes change; missing specs fall back to the median known duration, and the focused shard test catches empty, overlapping, or badly imbalanced partitions. ## Model Used OpenAI GPT-5 via Codex CLI coding agent, with shell/tool execution and repository inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6a36ae47fe |
Fix stable release dry-run notes gate (#9334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release subsystem uses GitHub Actions dry-run previews so maintainers can validate stable and canary release behavior before publishing. > - Stable release publishing needs a release notes gate so real `latest` publishes never happen without authored notes. > - The same gate was also blocking stable dry-run previews, which made release QA fail before publish-sensitive work could be previewed. > - This pull request narrows the notes requirement to real non-dry-run stable publishes. > - The benefit is that stable dry-run dispatches can preview the release without same-day notes, while real stable publishes remain protected. ## Linked Issues or Issue Description No public GitHub issue exists for this release-QA blocker, so the bug is described inline below. ### What happened? Stable dry-run release previews fail when the same-day release notes file is absent. The release script runs the stable notes-file gate before release preview work even when `--dry-run` is set. ### Expected behavior `./scripts/release.sh stable --dry-run` should preview the stable release without requiring `releases/vYYYY.MDD.P.md`. Real non-dry-run stable publishes must still fail before build/publish work starts when the notes file is missing. ### Steps to reproduce 1. Check out current `master` before this fix. 2. Ensure the computed same-day stable release notes file does not exist under `releases/`. 3. Run `./scripts/release.sh stable --skip-verify --dry-run`. 4. Observe that the script exits with `stable release notes file is required` instead of reaching the release preview plan. ### Paperclip version or commit Reproduced on `master` at `9a1d4b7983dfd50e8eb40ee9770e44999d405f60`. ### Deployment mode Built from source / GitHub Actions release workflow. ## What Changed - Narrowed the stable release notes gate to `channel=stable` and `dry_run=false`. - Updated the release script usage note to say the notes file is required for non-dry-run stable releases. - Added a targeted Node test covering dry-run allowed behavior and non-dry-run blocked behavior. - Stubbed release fixture registry-version checks so the test isolates the notes gate without hitting npm. ## Verification - `node --test scripts/__tests__/release-dry-run-notes.test.mjs` passed with 2/2 subtests. - `bash -n scripts/release.sh` exited 0. ## Risks Low risk. The behavior change only relaxes the notes-file gate for stable dry-runs. The new test verifies real non-dry-run stable publish still fails before build/publish work starts when notes are missing. ## Model Used OpenAI Codex, GPT-5-based coding agent, with shell/tool execution in a local repository workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce7dedf33d |
perf(ci): balance general-server test shards by recorded suite duration (#9516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its PR CI runs the general-server vitest lane pinned to `maxWorkers=1` and sharded across 3 runners (introduced in #8360) > - Suites were assigned to shards round-robin by sorted file index, so shard test time was unbalanced: a recent PR run split 73s / 153s / 115s, and the heaviest shard made "General tests (server 2/3)" the slowest check in the whole workflow at 314s wall > - The slowest shard sets the lane's wall time, so unbalanced partitions waste the other two runners and stretch the PR critical path > - This pull request replaces the round-robin assignment with a deterministic longest-processing-time partition weighted by a checked-in per-suite duration manifest > - The benefit is near-even shard weights (projected 113s / 113s / 113s with the current manifest), taking roughly 40s off the PR critical path with no reduction in coverage ## Linked Issues or Issue Description - Refs #8360 (introduced the 3-way general-server sharding this PR rebalances) - No public issue exists. Problem: the general-server test lane's round-robin shard assignment ignores per-suite duration, so one shard can carry multiple 30s+ suites while another finishes in half the time; the slowest shard alone determines the check's wall time. ## What Changed - `scripts/general-server-shard.mjs` (new): manifest loader and deterministic LPT (longest-processing-time) partitioner; suites missing from the manifest get the median recorded weight, and a missing or malformed manifest degrades to uniform weights so the lane never fails on stale data - `scripts/general-server-shard-durations.json` (new): per-suite duration manifest sampled from a real PR run (240 suites); the `$comment` field documents how to regenerate it - `scripts/run-vitest-stable.mjs`: both shard-selection sites (run and `--dry-run`) now use the balanced partition instead of index round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: 6 new tests covering skew-balance vs round-robin, determinism, median fallback for unlisted suites, malformed-manifest degradation, manifest coverage of the current suite set, and real-partition balance - `server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts`: hardened the `afterEach` sweep — post-run bookkeeping (run-event records, follow-up wake scheduling) can still insert rows briefly after a run reaches a terminal status, and a late insert landing between the `agent_wakeup_requests` and `agents` deletes failed teardown with a foreign-key violation on the first CI attempt of this PR; the sweep now retries so a late background write cannot take down the shard - `release-verify.yml` shares the same runner script and inherits the balancing with no workflow change ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs` — 9/9 pass (run against current master) - `npx vitest run src/__tests__/heartbeat-issue-rewake-throttle.test.ts` — 6/6 pass against embedded Postgres with the hardened teardown - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 2/2 pass - `node scripts/run-vitest-stable.mjs --dry-run` with each shard flag shows every suite assigned exactly once across the 3 shards, with projected weights ~113s each ## Risks - Low risk: partition changes which runner executes which suite, not what runs; a completeness test asserts every suite is assigned to exactly one shard - The duration manifest will drift as suites are added/changed; unlisted suites get the median weight and a coverage test flags when the manifest covers less than half the suite set, so drift degrades balance gracefully rather than breaking the lane > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking enabled, agentic tool use (file edits, shell, test execution) via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude (Paperclip SWE) <noreply@paperclip.ing> |
||
|
|
8775bde4ce |
Add runtime asset build-gap guard
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server package ships runtime asset trees used by built-in agents
and onboarding templates.
> - A prior server build omitted those source asset trees from `dist`,
allowing a built artifact to differ from runtime expectations.
> - The copy step is now present, but the existing build-gap gate only
checked TypeScript coverage for packages whose build skips `tsc`.
> - This pull request extends that standing gate so source asset files
under the server runtime asset trees must exist at the matching `dist`
paths after build.
> - The benefit is that future server asset additions fail loudly in CI
instead of silently shipping an incomplete `dist`.
## Linked Issues or Issue Description
### What happened?
After a server build, runtime asset files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**` could be
missing from `dist/` with no build failure. The existing build-gap gate
only checked TypeScript coverage for packages that skip `tsc`; it did
not verify that non-TypeScript source assets were copied to `dist`. A
server build that forgot the `cp -R` step, or that added a new asset
tree without updating the copy command, would produce an incomplete
`dist` without any CI signal.
### Expected behavior
After `pnpm --filter @paperclipai/server build`, every
non-TypeScript/non-JavaScript source file under `server/src/built-ins/`
and `server/src/onboarding-assets/` must exist at the matching path
under `server/dist/`. If any file is missing, the build-gap gate must
exit non-zero with a diagnostic listing the missing files and the
command to fix them.
### Steps to reproduce
1. Remove a copied runtime asset: `rm
server/dist/built-ins/agents/reflection-coach/AGENTS.md`
2. Run the guard: `node scripts/run-typecheck-build-gaps.mjs
--runtime-assets-only`
3. Before this fix: the command exits 0 and the missing file goes
undetected.
### Paperclip version or commit
Reproduced on `master` at `c36f1a4af` (`@paperclipai/server` 0.3.1).
### Deployment mode
Not deployment-specific — the build-gap check runs in CI on any
checkout.
## What Changed
- Extended `scripts/run-typecheck-build-gaps.mjs` with a source-derived
server runtime asset parity check for non-`.ts`/non-`.js` files under
`server/src/built-ins/**` and `server/src/onboarding-assets/**`.
- Added a guard-only mode, `--runtime-assets-only`, for focused
pass/fail verification after a server build.
- Wired `pnpm run typecheck:build-gaps` to prepare plugin SDK build
deps, build the server package, then run the existing build-gap gate
plus the new asset check.
## Verification
Pass path:
```text
$ pnpm --filter @paperclipai/plugin-sdk ensure-build-deps
> @paperclipai/plugin-sdk@1.0.0 ensure-build-deps .../packages/plugins/sdk
> node ../../../scripts/ensure-plugin-build-deps.mjs
$ pnpm --filter @paperclipai/server build
> @paperclipai/server@0.3.1 build .../server
> tsc && mkdir -p dist/onboarding-assets dist/built-ins && cp -R src/onboarding-assets/. dist/onboarding-assets/ && cp -R src/built-ins/. dist/built-ins/
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
Regression simulation (guard catches the missing file):
```text
$ rm server/dist/built-ins/agents/reflection-coach/AGENTS.md
$ node scripts/run-typecheck-build-gaps.mjs --runtime-assets-only
[typecheck:build-gaps] Missing server runtime asset(s) in dist:
- source: server/src/built-ins/agents/reflection-coach/AGENTS.md
expected dist: server/dist/built-ins/agents/reflection-coach/AGENTS.md
Run pnpm --filter @paperclipai/server build and ensure source runtime asset trees are copied into dist.
```
Standing gate (full end-to-end):
```text
$ pnpm run typecheck:build-gaps
[typecheck:build-gaps] typechecking 4 workspace(s): paperclipai, @paperclipai/plugin-authoring-smoke-example, @paperclipai/plugin-llm-wiki, @paperclipai/ui
[typecheck:build-gaps] server runtime assets present in dist: 7 file(s)
```
## Risks
Low risk. The check only reads source and dist files during the
build-gap gate. The main tradeoff is that the gate now builds
`@paperclipai/server` so a clean checkout has generated `dist` content
to validate.
## Model Used
OpenAI Codex, GPT-5 based coding agent with repository tool use and
shell execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
cec0fc249a |
[codex] Parallelize release verify workflow (#9168)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Releases publish the same app and package set that operators install, so release verification should keep full release-strength coverage. > - The release workflow currently verifies stable and canary releases with one serial job that typechecks, runs all tests, and builds. > - The PR workflow already proves the test surface can be split into grouped general suites and serialized shards without changing coverage. > - This pull request extracts the release verify work into a reusable workflow and fans out the independent lanes. > - The benefit is faster stable and canary release verification while preserving the existing publish and preview gates. ## Linked Issues or Issue Description No public GitHub issue exists for this CI improvement. **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Release verification spends most of its wall time in a single serial test step even though the same stable test surface is already partitioned for PR CI. Stable dispatches and master-push canaries therefore wait on one long runner after setup, typecheck, tests, and build run sequentially. **Proposed solution** Add a reusable release verification workflow with parallel typecheck, grouped general tests, serialized test shards, and build lanes. Have both stable and canary release verification call it with the ref they need to verify. **Alternatives considered** Keeping the serial `pnpm test:run` job preserves the old shape but keeps stable and canary releases waiting on one long runner. Skipping verification when a source SHA already has green CI would be faster, but adds stale-check and lookup risk beyond this change. **Roadmap alignment** No overlapping item found in `ROADMAP.md`; this is release CI maintenance. **Additional context** The new workflow keeps the release-strength full `pnpm -r typecheck`, uses the existing stable test grouping/sharding entry points, and leaves publish/preview jobs unchanged. ## What Changed - Added `.github/workflows/release-verify.yml` as a `workflow_call` workflow accepting a `ref` input. - Split release verification into parallel `typecheck`, `general_tests`, `serialized_tests`, and `build` jobs with 20-minute lane timeouts. - Mirrored the PR workflow's stable test partition: `general-server` shards 1-3, `general-workspaces-a`, `general-workspaces-b`, and four serialized shards. - Replaced `release.yml` `verify_canary` and `verify_stable` job bodies with calls to the reusable workflow while leaving publish and preview jobs unchanged. - Added a Node test that guards the release workflow delegation and split verify surface. ## Verification - `actionlint 1.7.12 .github/workflows/release.yml .github/workflows/release-verify.yml` - `node ./scripts/release-package-map.mjs check` - `node --test ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check` ## Risks - Release verification now starts more jobs per release event, increasing total runner setup/install minutes. This matches the existing PR CI tradeoff and should reduce release wall time substantially. - The called workflow checks out the requested ref shallowly. That is intentional for verify lanes; publish and preview jobs still retain their existing full-history checkouts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in local tool-use mode with shell execution, repository editing, GitHub connector access, and medium reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d3919713bc |
[codex] Document Storybook visual baseline platform lock (#9216)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Storybook visual baselines protect UI surfaces from unintended visual drift. > - Pixel-perfect screenshot baselines are sensitive to OS, font rasterization, and browser environment. > - The suite already stores external baseline artifacts and has an opt-in CI path. > - Local runs on non-matching platforms can report false-positive diffs unless the platform lock is explicit. > - This pull request documents the Linux/Ubuntu baseline constraint and makes the local static server port explicit. > - The benefit is clearer visual-review guidance and more predictable Playwright web server startup. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? The Storybook visual baseline suite requires a matching Linux capture environment for pixel-exact comparisons, but the docs did not clearly warn local users that non-Linux environments can produce false-positive diffs. The Playwright web server command also relied on the static server's default port instead of passing the configured port explicitly. ### Expected behavior Developers should see clear Linux/Ubuntu baseline guidance before running the visual suite locally, and Playwright should start the Storybook static server on the same explicit port that the test config expects. ### Steps to reproduce 1. Review the Storybook visual docs before this PR. 2. Run or inspect the Storybook visual Playwright config. 3. Notice the missing platform guidance and implicit static server port coupling. ### Paperclip version or commit Reproducible on `master` before this branch. ### Deployment mode Local dev (pnpm dev) / built from source. ## What Changed - Documents the Linux/Ubuntu-only baseline limitation in the developer docs and visual-suite README. - Adds `--port` parsing and validation to the Storybook static server helper. - Adds regression coverage for `--port` followed by another flag. - Passes the Playwright web server port explicitly from the Storybook visual config. ## Verification - Passed: `node --check scripts/serve-storybook-static.mjs` - Passed: `node --test scripts/__tests__/serve-storybook-static.test.mjs` - Passed: `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` - Greptile: 5/5 with no unresolved review threads after commit `94a649755a2ae7c4a34a3e8a1f16ec4d26d738fd`. - Not run: full `pnpm test:storybook-visual`, because it builds Storybook and runs the browser visual suite; this PR only changes docs plus server port plumbing. ## Risks Low risk. The server still defaults to port 6106 when no explicit port is provided, and invalid port values now fail fast with a clear error before the Playwright server waits for an unreachable URL. ## Model Used OpenAI GPT-5 Codex coding agent with local command execution and repository editing tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c07e650cd7 |
feat(ui): single-source design tokens, visual regression suite, and theme retune (#9134)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ef37203a48 |
perf(ci): build standalone public packages concurrently (#8567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - CI runs a Canary Dry Run job that exercises `release.sh`, which builds the standalone sandbox-provider packages for publish > - That step (`scripts/build-standalone-public-packages.mjs`) built the 7 provider plugins serially — each doing `rm -rf dist && tsc` — making it the dominant cost (~49s) inside the slowest PR check (~4.9m wall) after the general-server lane was already sharded > - The packages are independent (their own `node_modules` via `--ignore-workspace`, their own `dist`), so the serial build is pure latency with no correctness benefit > - This pull request builds them with a bounded-concurrency pool sized to the runner CPU count (overridable via `STANDALONE_BUILD_CONCURRENCY`), buffering each package's output and flushing it as one block so parallel logs stay readable, and aggregating failures by original index > - The benefit is a faster Canary Dry Run / PR feedback loop without changing what gets built or published ## Linked Issues or Issue Description No public GitHub issue exists. Inline feature/perf description: ### Problem or motivation `build-standalone-public-packages.mjs` builds standalone provider packages serially, making it the largest single cost inside the slowest PR check. ### Proposed solution Run independent per-package builds through a bounded-concurrency worker pool sized to runner CPU count, with an env override and readable buffered logs. ### Alternatives considered Keep the serial build for simpler logs, but that preserves the avoidable CI latency. ### Roadmap alignment This is CI maintenance and does not overlap planned core roadmap work. ## What Changed - `scripts/build-standalone-public-packages.mjs`: replaced the serial per-package build loop with a bounded-concurrency pool (default = runner CPU count, override via `STANDALONE_BUILD_CONCURRENCY`); per-package stdout/stderr is buffered and flushed as a single block; failures are aggregated by original package index so one failure neither aborts the others mid-flight nor obscures which package broke. - `scripts/__tests__/build-standalone-concurrency.test.mjs`: new `node:test` unit suite covering the pool (limit respected, all items run, ordered failure aggregation, env-override resolution). - `.github/workflows/pr.yml`: wired the new unit test into the policy job. ## Verification - `node --test ./scripts/__tests__/build-standalone-concurrency.test.mjs` → 6/6 pass - `node ./scripts/release-package-map.mjs check` → OK (29 enabled for CI publish) - `git diff --check origin/master..HEAD` → clean ## Risks - Low risk. Build inputs/outputs are unchanged; only scheduling differs. The concurrency is bounded by CPU count and overridable; output is buffered per package so logs remain attributable. If a package fails, all failures are still reported with their package index. ## Model Used - Claude (Anthropic), `claude-opus-4-8`, extended thinking with tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2853a9ae69 |
perf(ci): shard the general-server test lane across 3 runners (#8360)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every PR runs the `PR` GitHub Actions workflow, whose `verify` gate fans out into parallel test lanes (general tests, serialized server route suites, build, typecheck) > - The `General tests (server)` lane had grown into the run's critical path: it executed all ~213 non-route server suites serially in a single job (~7.2m of test time), more than 2x any other job > - It runs serially because `server/vitest.config.ts` pins `maxWorkers: 1`, so server suites cannot parallelize within a single runner — the only lever is spreading them across runners > - This pull request shards that lane into 3 even partitions that run on separate runners, mirroring the 4-way sharding already used for the serialized route suites > - The benefit is the lane drops from ~7.7m to ~2.4m/shard, cutting overall PR wall time roughly in half (~8.5m → ~4.2m) ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the underlying issue is described inline following the feature-request template. ### Problem or motivation PR CI wall time had crept back up to ~8.5m. On a recent fully-green run, the `General tests (server)` job took 7.72m — more than double any other job and the clear critical path. Of that, 7.23m was pure test execution (dependency install was a cached 0.27m). The job ran all server suites that are not route/authz tests (213 files) one after another, because the server vitest project pins `maxWorkers: 1`, making these suites inherently serial within a single runner. ### Proposed solution Shard the general-server lane across 3 parallel runners — the same technique the route/authz suites already use — so the suite set is split into even, deterministic partitions that run concurrently. Add a regression test that proves the shards always cover the full suite set with no gaps or overlap. ### Alternatives considered - **Raise `maxWorkers` for the server project** to parallelize within one runner — rejected: the server suites share process-level state (DB/port), which is exactly why `maxWorkers: 1` is pinned. - **Two shards instead of three** — would leave the lane at ~3.6m, still above the next bottleneck (Canary Dry Run, ~4.1m wouldn't be the gate). Three lands the lane comfortably below it. - **Do nothing / accept the slow lane** — rejected: it gates every PR. ### Roadmap alignment Developer-experience / CI tooling. Not core product roadmap work; does not overlap with planned features in `ROADMAP.md`. ## What Changed - `scripts/run-vitest-stable.mjs`: the `general-server` general-test group now accepts `--shard-index` / `--shard-count`. It enumerates the full server test set (the whole `server/src` tree, minus the route/authz suites that already run in their own serialized shards) and splits it deterministically by modulo. The non-sharded local invocation (`pnpm test:run:general --group general-server`) is unchanged. - `.github/workflows/pr.yml`: the `general_tests` matrix runs `general-server` as 3 parallel shards (1/3, 2/3, 3/3). Workspace groups are unchanged. The `verify` gate already aggregates the whole matrix result, so the required check name is unaffected. - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: a `node:test` suite asserting the 3 shards form a complete, non-overlapping partition of the general-server set, that no route/authz suite leaks into it, and that shard flags are rejected for the parallel workspace groups. Wired into the `policy` job. ## Verification - New partition test passes locally: `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` (3/3). - Confirmed the 3 shards form a complete, non-overlapping partition of all 213 files (71/71/71). - Ran a live thin shard (3 real server suites, including one outside `__tests__`) — 23 tests passed, confirming positional-include execution works end to end. - This PR's own CI is the authoritative check: all three `General tests (server (n/3))` jobs went green on the prior run, collectively covering every suite the old single job ran. ## Risks - Low risk. No product code changes — only test orchestration and CI matrix. Shard partitioning is deterministic and is now covered by an automated test that fails if the partition ever develops a gap or overlap. Modulo-on-sorted-filenames balances duration reasonably, matching the approach already proven by the serialized route shards. ## Model Used - Claude (Anthropic), `claude-opus-4-8`, extended thinking + tool use (agentic coding via Paperclip). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots (N/A — no UI change) - [x] I have updated relevant documentation to reflect my changes (inline comments explain the sharding rationale) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending this PR's run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |