mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 00:54:38 +02:00
codex/opencode-pre-routing-skill
17
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d23cb6962 |
ci: pin a checkout-independent Rust cache path so PR lanes hit master's cache (#14394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip ships a native Runner binary, written in Rust, and seven CI lanes build it on every pull request > - `Canary Dry Run` is the slowest check on every green PR run, and most of its time is `cargo build --release` on third-party crates > - Master saves a Rust dependency cache for these lanes, but every PR lane logs `No cache found` and compiles every crate from zero > - The cache key matches, but GitHub also compares a hash of the absolute cache paths, and the master writer (RunsOn fleet, `/home/runner/_work/...`) and the PR readers (GitHub-hosted, `/home/runner/work/...`) hash different paths > - This pull request gives both sides a checkout-independent workspace path, so the hashes match and the PR lanes restore master's cache > - The benefit is about 2.5 minutes less wall clock per PR run and about 18 fewer runner-minutes per run ## Linked Issues or Issue Description No public issue exists for this problem. The description below follows the enhancement template. Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs #13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the workspace path, so none of them fixes this miss. **What existing behavior does this improve?** The `Swatinem/rust-cache` restore step in the PR workflow lanes that build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release Registry`, and the four `Verify Paperclip Runner` lanes. **Subsystem affected** CI workflows under `.github/workflows/`, their guard tests under `.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`. **Current behavior** Every PR lane logs `No cache found` although master holds an entry with the exact key. Run 36424309181 computed `v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master holds a 678 MB entry with that key. GitHub matches a cache entry on the key and on a version hash of the absolute paths in the cache. The master writer runs on the RunsOn fleet, where the checkout is `/home/runner/_work/paperclip/paperclip`. The PR readers run on GitHub-hosted `ubuntu-latest`, where the checkout is `/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…` is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The `work` paths hash to `1656e9ee…`. The key can never match, so each lane compiles every third-party crate again. **Proposed behavior** The writer and the readers pass the same checkout-independent path to `rust-cache`. Both runner layouts then produce the same version hash, and the PR lanes restore master's cache. **Reason and benefit** `Canary Dry Run` takes 533s on a green run. 251s of that is dependency compilation that a warm cache removes. Seven lanes pay this cost in every PR run. **Breaking changes** None. This change affects CI only. ## What Changed - Add a `Pin the Runner Rust workspace path` step before `rust-cache` in the master writer (`release-verify.yml`, typecheck and runner lanes) and in all four PR readers (`pr-trusted.yml`). The step creates the symlink `$HOME/paperclip-runner-rust` → `$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache` resolves the input with `path.resolve`, which does not follow symlinks, so both runner layouts now produce the same cache paths and the same version hash. `$HOME` is `/home/runner` on both images, which is why the `~/.cargo` paths already agreed. - Bump the shared keys `release-runner-v1` → `release-runner-v2` and `release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable entries are then visibly orphaned instead of sharing a key with the new ones. - Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`, and `typecheck-rust-cache`. They now require the pin step in both workflows with identical text, placed before the cache step, and they reject a `workspaces:` value that resolves under the checkout. The `pr-runner-rust-cache` test checks all four PR reader jobs and fails if a `rust-cache` step appears in a PR job that is not in its reader list. - Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the `release-runner-v2` key and to explain the pinned workspace path. ### Expected savings once merged Measured from run 36424309181. "Removed" is the dependency-compile time that a warm restore removes, minus about 18s to restore the 680 MB entry. The fleet writer's own restore shows this cost. | Lane | Today | Removed | Expected | |---|---|---|---| | Canary Dry Run | 533s | ~150s | ~380s | | Typecheck + Release Registry | 462s | ~155s | ~305s | | Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s | | Verify Paperclip Runner (rust) | 400s | ~175s | ~225s | | Build | 348s | ~130s | ~220s | | Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s | | Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s | - Wall clock per PR run: about 533s → about 385s. That is about 2.5 minutes faster to a green check set. `Canary Dry Run` stays the longest check. The rest is the non-cargo work in `release.sh` (standalone package builds ~30s, publish-payload preview ~73s). - Runner time: about 18 runner-minutes saved per PR run across the seven lanes. - The first master push after merge compiles from zero once in the fleet writer (about 4 extra minutes on that one run) and saves the v2 entry. Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key still gives a partial hit, as before. ## Verification - Run the guard tests for the three cache lanes: `node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/release-runner-cache.test.mjs .github/scripts/tests/typecheck-rust-cache.test.mjs` Result: 21 pass, 0 fail. - Run the full guard suite: `node --test '.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3 failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox temp-file ENOENT and fail the same way on the unmodified branch. - Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`. Result: 14 pass. - Local archive test: create a tar from the `_work` layout through the symlink (relative `../../../paperclip-runner-rust/target` entries, `tar -P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract it on the `work` layout. The files land in the real target directory and the symlink stays intact. - After merge, open any GitHub-hosted PR run and confirm that the seven Rust lanes log `Restored from cache key ...release-runner-v2...` in place of `No cache found`. ## Risks - Low risk. The change touches CI workflows, their tests, and one doc page. No product code changes. - If the pin step fails, `rust-cache` reports a miss and the lane compiles from zero, as it does today. The build does not break. - Both runner layouts sit four levels under `/home/runner`, so the relative `../../../` archive entries line up. The existing `~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this property. A future runner image with a different `$HOME` depth would miss the cache but would not fail the job. - `rm -rf "$pinned"` acts on the symlink itself (no trailing slash), never on the checkout behind it. It only matters on a reused runner. - Squash-merge note: the branch carries commits by `Bender (Fable)`. Add `Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the squash body to keep that authorship. ## Model Used - Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache listings, and PR operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal ticket id. The agent execution workspace fixed this branch name, so I cannot rename it. Squash-merge drops the branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Bender (Fable) <bender-fable@paperclip.local> |
||
|
|
bd51f157e9 |
fix(ci): make the Runner Rust cache key independent of the image toolchain (#13500)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every change goes through the pull request CI workflow, and its `Verify Paperclip Runner` lane builds and tests the native Rust runner > - https://github.com/paperclipai/paperclip/pull/13457 made that lane restore master's prebuilt Rust dependency cache instead of recompiling 313 crates > - Measured over 26 runs since it went live, the cache hits on the RunsOn fleet and misses on every GitHub-hosted runner > - The cause is the cache key: `rust-cache` hashes every installed toolchain, and each runner image ships a different stable Rust next to the pinned one > - This pull request removes the extra toolchains before the key is computed, in the reader and the writer > - The benefit is that the saving the cache already delivers, 4.8 minutes per run, reaches the 73% of runs that currently miss it ## Linked Issues or Issue Description No public GitHub issue exists for this. The problem follows the enhancement issue template below. - Refs https://github.com/paperclipai/paperclip/pull/13457 — added the cache this pull request repairs - Refs https://github.com/paperclipai/paperclip/pull/13459 — activated it - Refs https://github.com/paperclipai/paperclip/pull/13194 — created the `release-runner-v1` entry on master I searched this repository for other pull requests touching this cache and found no duplicate and nothing in flight. **What existing behavior does this improve?** The Rust dependency cache added in #13457 misses on GitHub-hosted runners, so most pull requests still recompile the whole dependency tree. **Subsystem affected** Cross-cutting (multiple of the above). The change touches CI workflow configuration only. It does not change product code. **Current behavior** The cache works, but only on one runner class. Across 26 successful `Verify Paperclip Runner` jobs since #13459 merged: | Runner | Runs | Before | After | Change | Cache | |---|---|---|---|---|---| | RunsOn fleet | 7 | 16.4m | 11.6m | −4.8m | 4 of 4 full hit | | ubuntu-latest | 19 | 14.9m | 14.7m | −0.2m | 0 of 6 hit | | All | 26 | 15.0m | 13.9m | −1.1m | | Only 27% of runs reach the fleet, so the fleet-wide saving is 1.1 minutes rather than the 4.8 minutes the cache delivers where it lands. The two runners compute different keys: ``` fleet: v0-rust-release-runner-v1-Linux-x64-4bb3b8ea-a95b0328 ubuntu-latest: v0-rust-release-runner-v1-Linux-x64-9fdc73e3-a95b0328 ``` The lockfile half agrees. The environment half does not. `rust-cache` logs why, under `Environment considered`: | Runner | Toolchains it found | |---|---| | fleet | 1.97.1 and **1.98.0** | | ubuntu-latest | 1.97.1 and **1.98.1** | `rust-cache` hashes every installed toolchain, not only the active one. Both images carry the pinned 1.97.1. Each also ships its own stable Rust, and those differ by a patch version. The post-merge writer runs on a RunsOn image, so the fleet agrees with it and GitHub-hosted runners cannot. Pinning `RUSTUP_TOOLCHAIN` in #13457 was necessary but not sufficient. It fixes which toolchain builds the code. It does not change which toolchains exist on the image. **Proposed behavior** Remove every toolchain except the pin, before the cache step, in both the reader and the `release-runner-v1` writer. The key then depends on the pinned compiler and the lockfile alone, not on what the image happens to carry. **Reason and benefit** The cache already proves its value where it lands: release compile drops from 5m07s to 1m30s, debug from 1m48s to 13s, and the job from 16.4m to 11.6m. This change extends that to the other 73% of runs. Expected fleet-wide mean: about 11.5m, against 13.9m today and 15.0m before #13457. It also removes a standing fragility. The fleet hits today only because two RunsOn images happen to agree. If either image updates its stable Rust on its own, the hit rate drops to zero with no code change. **Breaking changes** None. The change only affects cache key computation. A miss reproduces the current behavior. **Additional context** The typecheck writer in `release-verify.yml` keeps its current step on purpose. It restores and saves on the same post-merge image, so its key never disagrees with itself. ## What Changed - Added a toolchain normalization block to `Select the pinned Runner Rust toolchain` in `.github/workflows/pr-trusted.yml`, before the cache restore. It keeps the pinned toolchain and uninstalls the rest. - Added the identical block to the same step in the `verify_paperclip_runner` job of `.github/workflows/release-verify.yml`, which writes `release-runner-v1`. The reader and the writer must agree, or the key matches nothing. - Made the block tolerant. If a toolchain cannot be removed it prints a notice and continues, so a pull request loses the cache rather than the run. - Extended `.github/scripts/tests/pr-runner-rust-cache.test.mjs` with two tests: the block is byte-identical in both workflows, and it runs before the cache step in each. ## Verification Run the workflow shape tests: ```bash node --test '.github/scripts/tests/*.test.mjs' ./scripts/__tests__/e2e-shard.test.mjs ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs ./scripts/cloud-source-verification.test.mjs ``` Result: 471 pass, 0 fail. I ran the step body against a stub `rustup` to confirm the logic, rather than only checking syntax. Three paths, all exit 0: | Case | Result | |---|---| | Pin plus an extra stable toolchain | Uninstalls only the extra, exports `RUSTUP_TOOLCHAIN=1.97.1-x86_64-unknown-linux-gnu` | | Pin only, already normalized | No uninstall calls, no error | | No `rustup` on `PATH` | Prints the notice and continues | I also mutation-tested the new parity assertion. Each mutation edits only the writer, then the reader and writer disagree: | Mutation to `release-verify.yml` | Result | |---|---| | One word changed in the shared comment | Caught | | `uninstall` changed to `remove` | Caught | | Trailing `rustup toolchain list` deleted | Caught | After this merges, confirm the repair in CI. Take a `Verify Paperclip Runner` job that ran on `ubuntu-latest` and check the restore step for `full match: true`. Under `Environment considered`, `Rust Versions` must list only 1.97.1. The job should finish near 11.5m rather than 14.7m. ## Risks - **Master must republish the cache once.** This changes the key, so the existing `release-runner-v1` entry no longer matches. `cloud-readiness.yml` runs `release-verify.yml` on every push to master, so the first push after this lands writes the new entry. Pull requests merged before that miss the cache, which is exactly what most of them do today. No run breaks. - **The reader and the writer must stay in step.** If they diverge, no run hits the cache. The new parity test fails on any difference in the block, including a comment. - **Uninstalling the image toolchain is deliberate.** Everything in this job builds through the pinned 1.97.1, resolved from `packages/paperclip-runner/rust-toolchain.toml` and `RUSTUP_TOOLCHAIN`. Nothing in the job uses the image default. - **Low blast radius.** The change only affects cache key inputs. A miss compiles from scratch, as today. - **Image drift is now handled.** The key no longer depends on the image's own Rust version, so a future image update cannot silently disable the cache. ## Model Used Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M context window. Adaptive thinking was on. I used tool use throughout: the `gh` CLI and the GitHub API to pull 26 post-merge job records and 10 full job logs, log parsing in Python to isolate the per-phase timings and the two cache keys, a stub `rustup` on `PATH` to exercise the new step, and local `node --test` runs to verify and mutation-test the guards. Run through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Note on the unchecked boxes. The CI and Greptile boxes stay unchecked until those checks finish. On documentation: no document describes the Runner cache keys, so there is nothing to update. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
37d7dfb0e3 |
ci: allow dependency changes in cloud eval verification (#13286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployment requires source verification for the exact merged commit. > - Contributor PRs leave lockfile updates to a separate bot PR. > - Most release checks can refresh an outdated lockfile while installing dependencies. > - Two Runner checks still require a frozen lockfile and fail after dependency changes. > - This PR gives those checks the same install policy as the other release checks. > - A valid dependency change can become deployable without waiting for another merge. ## Linked Issues or Issue Description Refs #13257. The dependency change in #13256 exposed this gap. The separate lockfile update is #13279. Related #12115 addresses the bot PR check trigger; this PR fixes exact-source cloud verification itself. **What happened?** [Cloud readiness for |
||
|
|
bc68312327 |
ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.
## Linked Issues or Issue Description
Refs #13243.
**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.
**Current behavior**
For merge
|
||
|
|
a23ae894a5 |
ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud images become deployable only after source verification passes. > - The typecheck job builds the native Runner binary through the server package. > - Fresh runners repeatedly compile Rust dependencies for that binary. > - This PR caches those dependencies for exact-source master verification. > - Workspace code and every typecheck still rebuild or run as before. ## Linked Issues or Issue Description Refs #13243 and #13257. **What existing behavior does this improve?** The typecheck portion of post-merge cloud source verification. **Current behavior** The typecheck job has no Rust dependency cache. An observed release build in this job took 4m 13s, including dependency compilation. **Proposed behavior** Restore dependency build outputs for canonical master pushes with the exact source SHA. Use a separate cache key from the Runner verification job, which builds other profiles. **Reason and benefit** A warm cache should remove roughly 2–3 minutes of dependency compilation from this job. Overall deployment gains depend on the remaining critical path. The first cache population still compiles from scratch. **Breaking changes** None. All checks remain enabled. Non-master callers compile without restoring or saving this cache. ## What Changed - Select the pinned Rust toolchain before the typecheck cache lookup. - Reuse the existing pinned Rust cache action with a typecheck-specific key. - Exclude workspace crates and installed cargo executables. - Test restore/save trust boundaries and document cache behavior. ## Verification - All workflow script tests pass locally, including nine new cache trust/contract cases. - actionlint passes for release-verify.yml. - Full local typecheck and build pass on the same source base (167s and 206s); `pnpm test:run` is still running and is recorded with #13257. This PR changes only the workflow, cache guard tests, and documentation. - All 32 current-head checks are successful or intentionally skipped, including the complete Linux test matrix, build, and Greptile 5/5 with no unresolved findings. Verify cache population and subsequent restore on actual master runs. ## Risks - The first run and any toolchain/dependency invalidation compile from scratch. - Cache restore/save overhead reduces the benefit for small dependency graphs. - Disable the cache step to roll back; the existing uncached build remains valid. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. Exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — focused change tests pass; full-suite local permission failures are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
974949a39b |
ci: spread cloud server verification across ten runners (#13227)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud waits for source verification before deploying a new image. > - The slowest server verification job spends about ten minutes running tests. > - Each job uses one test worker to preserve test isolation. > - This pull request distributes those suites across ten standard hosted runners. > - The benefit is a shorter verification path with the same test coverage. ## Linked Issues or Issue Description **Current behavior** In [readiness run 34572340764](https://github.com/paperclipai/paperclip/actions/runs/34572340764), the slowest server job ran for 638 seconds. Test execution used 594 seconds. This held readiness behind the image job. **Proposed behavior** Use ten general server jobs in the reusable release verification workflow. Keep the three chat jobs and every existing prerequisite. The complete partition test verifies that no server suite is omitted or duplicated. **Reason and benefit** Reduce merge-to-deployable time on the existing runner type. The next longest prerequisite was Runner verification at 526 seconds, so the initial expected total gain is about two minutes rather than a halving of readiness time. Measure actual queue and execution time before claiming a result. Related: #13198 introduced the separate chat lane. #12577 refreshes duration estimates; this change leaves that manifest alone. ## What Changed - Increase the general server matrix from five jobs to ten. - Verify the ten-way partition covers the complete server suite when combined with the chat lane. - Document runner demand and the unchanged local and PR grouping. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/__tests__/run-vitest-stable-shard.test.mjs`: 29 passed. - `actionlint .github/workflows/release-verify.yml`: passed. - Full local `pnpm -r typecheck` and `pnpm build`: passed. - All latest-head GitHub CI checks passed, including the complete Linux test partition, build, typecheck, and browser gates. Greptile: 5/5 with zero open findings. - [Ten-shard timing probe](https://github.com/paperclipai/paperclip/actions/runs/34606772388): all 16 jobs passed; slowest server job 6m 23s versus 10m 38s in the earlier five-shard sample. This compares the server lane, not total readiness, and is not a controlled same-source A/B. - The full local `pnpm test:run` is also running. It has reproduced previously observed macOS-only failures in unchanged skill-cache and native-session suites; the corresponding Linux CI suites passed. Final local results will be attached separately. No affected-workflow test failed. ## Risks Five additional concurrent jobs per release verification run increase runner demand and repeated setup work. Queueing can offset the gain. Test workers, timeouts, permissions, and readiness requirements stay unchanged. Revert the matrix and its partition test to restore the previous split. ## Model Used OpenAI GPT-6 / Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected workflow tests locally and they pass; full-suite macOS limitations are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
42961b6ef1 |
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e970df4c5 |
ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c5c80e1feb |
ci(release-verify): split server tests five ways like pr-trusted (#13185)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every master push publishes a canary through release.yml, gated by release-verify.yml — the fleet's staging deploys and the nightly/beta/stable chain all start from those canaries > - release-verify splits the server test suite across three shards with a 20-minute job cap, while pr-trusted splits the same suite across five > - The server suite grew on 2026-09-10 and the three shards moved to 17-19 minutes; that evening every push-triggered canary run was cancelled by the 20-minute cap mid-verify, and no canary published after 18:50 UTC > - This pull request mirrors pr-trusted's five-way server split in release-verify, putting shards back at the 10-15 minute range with real headroom > - The benefit is a canary lane that reports test verdicts instead of dying on an infrastructure cap ## Linked Issues or Issue Description **What happened?** Push-triggered Release runs stopped publishing canaries on 2026-09-10. Runs at 19:34, 22:30, and 22:37 UTC were all cancelled by "The job has exceeded the maximum execution time of 20m0s" on a `verify_canary / General tests (server (N/3))` shard. No canary published after 18:50 UTC, which also starves the staging fleet's continuous deploys. **Expected behavior** release-verify's server shards finish well inside the 20-minute cap and runs conclude with a test verdict, as pr-trusted's five-way split of the same suite does (10-15 minutes per shard). **Steps to reproduce** 1. Compare server shard durations in the `verify_canary` job across 2026-09-10: 11-14 minutes in the morning, 17-19 minutes from 15:06 UTC, over 20 minutes by evening. 2. Observe runs 34521169020, 34537798488, and 34538332689 cancelled at the cap. **Paperclip version or commit** `master` at `d1ba17eec` (current tip; its canary run was one of the cancelled ones). ## What Changed - `release-verify.yml`: the `general-server` matrix goes from three shards to five, byte-for-byte the shape `pr-trusted.yml` already runs, with a comment recording why. ## Verification - The identical five-way split runs green on every pr-trusted run (10-15 minutes per shard today, including on PRs merged this evening). - The suite's own growth (slower chat-connector tests) is being addressed separately; this PR only removes the artificial cliff. ## Risks - Low risk: two more runners per verify run; no test content changes. If shard durations regress further, the cap fires again — which is the correct signal once shards have honest headroom. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), extended thinking, tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2a05b5ed34 |
ci: split runner verification from build (#13142)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip uses GitHub Actions to verify changes before release. > - The Paperclip Runner has a separate verification boundary. > - The build job currently runs this verification before the workspace build. > - This pull request moves runner verification into its own parallel job. > - The benefit is clearer CI results and less wait time for independent work. ## Linked Issues or Issue Description **What existing behavior does this improve?** The trusted PR and release verification workflows run Paperclip Runner verification inside the Build job. **Subsystem affected** Cross-cutting (GitHub Actions CI workflows). **Current behavior** The Build job runs `pnpm --filter @paperclipai/paperclip-runner check:all` before it builds the workspace. A runner verification failure appears as a Build failure. The workspace build cannot run in parallel with runner verification. **Proposed behavior** Each workflow has a `Verify Paperclip Runner` job with the same checkout, dependency install, and command. The Build job only builds its required outputs. Both jobs run after the same gate and policy jobs. **Reason and benefit** The runner command is an independent verification boundary. A dedicated job gives it a clear status and allows it to run in parallel with Build. **Breaking changes** None. The same runner verification command still runs in both workflows. ## What Changed - Added a dedicated `Verify Paperclip Runner` job to the trusted PR workflow. - Added a dedicated `Verify Paperclip Runner` job to the release verification workflow. - Kept the Build jobs independent and retained their existing build commands. - Updated the trusted-workflow policy test for the additional dependency-install job. ## Verification - Ran `git diff --check`. - Ran `node --test ./scripts/__tests__/e2e-shard.test.mjs`. - Ran `pnpm exec prettier --check .github/workflows/pr-trusted.yml .github/workflows/release-verify.yml`. - Confirmed both jobs retain their prior runner, dependency, and policy prerequisites. ## Risks Low risk. The runner verification job repeats the existing setup. It adds one parallel GitHub Actions runner to each affected workflow. ## Model Used OpenAI Codex, GPT-5.6, 128k context window, reasoning and tool-use capabilities. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ffff1fe6e3 |
feat(runner): define package API and verification boundary (#12129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package now has protocol, transport, provider, catalog, and authorization foundations. > - Its first upstream package boundary should expose only the implemented runtime and test-helper surfaces. > - Rust correctness belongs in the repository existing build verification, without introducing a parallel release process. > - Direct package creation must build the files declared by the package manifest. > - This pull request defines the minimal package API and verifies the optimized runner binaries in the existing PR and release Build jobs. > - The benefit is a production-ready runner package boundary with minimal build-process change. ## Linked Issues or Issue Description Refs #11962 This pull request replaces one bounded part of the archived large runner change. It follows the package-local authorization change in #12126. ## What Changed - Export only `@paperclipai/paperclip-runner` and `@paperclipai/paperclip-runner/testing`. - Keep Node-only fixture loading and semantic conformance helpers out of the runtime root. - Add a provider-neutral semantic conformance kit with stable JSON comparison and fail-closed input checks. - Keep deferred SDK, eval, browser, React, lab, and command surfaces private. - Pin the runner Rust toolchain to 1.97.1 with the minimal profile and `rustfmt`. - Run the Rust workspace tests in release mode. - Launch the optimized `paperclip-runnerd` and fake-harness binaries in process-level integration coverage. - Add one `pnpm --filter @paperclipai/paperclip-runner check:all` step to each existing PR and release Build job. - Make the existing server `prepack` lifecycle run its existing build after it prepares UI assets. - Document that no production adapter starts runnerd yet. This revision adds no standalone GitHub Actions job. It adds no server runner dependency or runner vendoring. It adds no Docker bootstrap or clean-consumer harness. It does not change `pnpm-lock.yaml`. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` - 66 TypeScript tests - 8 protocol contract tests - 56 Rust unit and integration tests - Release-mode integration coverage launches the optimized runnerd and fake-harness binaries. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-package-build-script.test.ts` (2 tests) - Clean `pnpm pack` from `server/` rebuilt the server and produced both `package/dist/index.js` and `package/dist/index.d.ts`. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` (8 tests) - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `git diff --check` - No `pnpm-lock.yaml` diff. - The diff changes 12 files. ## Risks The runner adds Rust work to the existing Build jobs. These jobs can take longer on a cold cache. The pinned toolchain makes contributor and CI behavior reproducible. Cargo tests use `--release` to verify optimized executables. The server prepack lifecycle now performs the build that its published entry points require. This can make direct server packing slower. This pull request does not wire runnerd into the server. It does not select runnerd for any adapter. Existing application execution and finalization paths remain unchanged. ## Model Used OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b6e58019f2 |
ci: split serialized tests into five shards (#10923)
## Thinking Path > - Paperclip uses CI to keep control-plane changes safe and mergeable. > - The PR workflow splits serialized server tests across isolated runners. > - A recent successful run spent 305 seconds in serialized shard 2/4. > - That job was the slowest check in the run. > - The four shards reported about 739 seconds of Vitest suite time. > - This pull request adds a fifth serialized shard and keeps release verification aligned. > - The benefit is a shorter PR critical path with no loss of test coverage. ## Linked Issues or Issue Description **What existing behavior does this improve?** The PR and release verification workflows run serialized server tests in four shards. **Current behavior** Successful PR run 30876682788 spent 305 seconds in `Verify serialized server suites (2/4)`. The test step used 256 seconds and made this job the slowest check. **Proposed behavior** Run the same serialized suite set in five complete and non-overlapping shards. **Reason and benefit** The measured suites reported about 739 seconds of total Vitest time. Five runners reduce the expected average suite time from about 185 seconds to about 148 seconds before setup overhead. **Breaking changes** None. The change only alters CI partition size. ## What Changed - Split serialized server tests into five shards in the PR workflow. - Apply the same five-shard layout to release verification. - Add a partition test that proves complete and non-overlapping serialized coverage. - Update release workflow coverage tests for five shards. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` - `git diff --check` ## Risks - Low risk. CI uses one additional runner for the serialized lane. - Round-robin partition weights can still vary as suite timings change. > This change does not overlap with planned core work in `ROADMAP.md`. Related PR #10663 optimized the separate general-server lane. ## Model Used - OpenAI Codex, GPT-5, agentic coding with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Devin Foley <139239+devinfoley@users.noreply.github.com> |
||
|
|
18fe9486a0 |
build(deps): bump actions/setup-node from 6 to 7 (#9884)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 6 to 7. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/actions/setup-node/releases">actions/setup-node's releases</a>.</em></p> <blockquote> <h2>v7.0.0</h2> <h2>What's Changed</h2> <h3>Enhancements:</h3> <ul> <li>Add cache-primary-key and cache-matched-key as outputs by <a href="https://github.com/gowridurgad"><code>@gowridurgad</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1577">actions/setup-node#1577</a></li> <li>Migrate to ESM and upgrade dependencies by <a href="https://github.com/gowridurgad"><code>@gowridurgad</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1574">actions/setup-node#1574</a></li> </ul> <h3>Bug fixes:</h3> <ul> <li>Remove dummy NODE_AUTH_TOKEN export by <a href="https://github.com/gowridurgad"><code>@gowridurgad</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1558">actions/setup-node#1558</a></li> <li>Only use <code>mirrorToken</code> in <code>getManifest</code> if it's provided by <a href="https://github.com/deiga"><code>@deiga</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1548">actions/setup-node#1548</a></li> </ul> <h3>Documentation updates:</h3> <ul> <li>Add documentation for publishing to npm with Trusted Publisher (OIDC) by <a href="https://github.com/chiranjib-swain"><code>@chiranjib-swain</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1536">actions/setup-node#1536</a></li> <li>docs: Update restore-only cache documentation by <a href="https://github.com/priya-kinthali"><code>@priya-kinthali</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1550">actions/setup-node#1550</a></li> <li>docs: Update caching recommendations to mitigate cache poisoning risks by <a href="https://github.com/chiranjib-swain"><code>@chiranjib-swain</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1567">actions/setup-node#1567</a></li> </ul> <h3>Dependency update:</h3> <ul> <li>Upgrade <code>@actions/cache</code> to 5.1.0, log cache write denied by <a href="https://github.com/jasongin"><code>@jasongin</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1569">actions/setup-node#1569</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/chiranjib-swain"><code>@chiranjib-swain</code></a> made their first contribution in <a href="https://redirect.github.com/actions/setup-node/pull/1536">actions/setup-node#1536</a></li> <li><a href="https://github.com/deiga"><code>@deiga</code></a> made their first contribution in <a href="https://redirect.github.com/actions/setup-node/pull/1548">actions/setup-node#1548</a></li> <li><a href="https://github.com/jasongin"><code>@jasongin</code></a> made their first contribution in <a href="https://redirect.github.com/actions/setup-node/pull/1569">actions/setup-node#1569</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/setup-node/compare/v6...v7.0.0">https://github.com/actions/setup-node/compare/v6...v7.0.0</a></p> <h2>v6.5.0</h2> <h2>What's Changed</h2> <ul> <li>Update <code>@actions/cache</code> to 5.1.0 and add security overrides for undici and fast-xml-parser by <a href="https://github.com/HarithaVattikuti"><code>@HarithaVattikuti</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1579">actions/setup-node#1579</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/setup-node/compare/v6.4.0...v6.5.0">https://github.com/actions/setup-node/compare/v6.4.0...v6.5.0</a></p> <h2>v6.4.0</h2> <h2>What's Changed</h2> <h3>Dependency updates:</h3> <ul> <li>Upgrade <a href="https://github.com/actions"><code>@actions</code></a> dependencies by <a href="https://github.com/Copilot"><code>@Copilot</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1525">actions/setup-node#1525</a></li> <li>Update Node.js versions in versions.yml and bump package to v6.4.0 by <a href="https://github.com/priya-kinthali"><code>@priya-kinthali</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1533">actions/setup-node#1533</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/Copilot"><code>@Copilot</code></a> made their first contribution in <a href="https://redirect.github.com/actions/setup-node/pull/1525">actions/setup-node#1525</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/setup-node/compare/v6...v6.4.0">https://github.com/actions/setup-node/compare/v6...v6.4.0</a></p> <h2>v6.3.0</h2> <h2>What's Changed</h2> <h3>Enhancements:</h3> <ul> <li>Support parsing <code>devEngines</code> field by <a href="https://github.com/susnux"><code>@susnux</code></a> in <a href="https://redirect.github.com/actions/setup-node/pull/1283">actions/setup-node#1283</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/actions/setup-node/commit/820762786026740c76f36085b0efc47a31fe5020"><code>8207627</code></a> Migrate to ESM and upgrade dependencies (<a href="https://redirect.github.com/actions/setup-node/issues/1574">#1574</a>)</li> <li><a href="https://github.com/actions/setup-node/commit/04be95cf3511ea51ebf9f224ddfb99cc7ab87cd4"><code>04be95c</code></a> Add cache-primary-key and cache-matched-key as outputs (<a href="https://redirect.github.com/actions/setup-node/issues/1577">#1577</a>)</li> <li><a href="https://github.com/actions/setup-node/commit/7c2c68d20d402ed6a201ada70a81341941093140"><code>7c2c68d</code></a> docs: Update caching recommendations to mitigate cache poisoning risks (<a href="https://redirect.github.com/actions/setup-node/issues/1567">#1567</a>)</li> <li><a href="https://github.com/actions/setup-node/commit/6a61c0375d66246de94630495909f12cf8dac84d"><code>6a61c03</code></a> Merge pull request <a href="https://redirect.github.com/actions/setup-node/issues/1569">#1569</a> from jasongin/update-actions-cache-5.1.0</li> <li><a href="https://github.com/actions/setup-node/commit/30eb73b41ded577900c1ebf968ef95cdf8f7434f"><code>30eb73b</code></a> Resolve high-severity audit issues</li> <li><a href="https://github.com/actions/setup-node/commit/4e1a87a501d0302f99e30e2748568adcb388d09f"><code>4e1a87a</code></a> Update dist</li> <li><a href="https://github.com/actions/setup-node/commit/360237f0c01778d0c17291f75c56d6feae4f7574"><code>360237f</code></a> Strict equality</li> <li><a href="https://github.com/actions/setup-node/commit/4f8aac5beb2f0854bc79651567a18c67eb0b9de3"><code>4f8aac5</code></a> Bump <code>@actions/cache</code> to 5.1.0, log cache write denied</li> <li><a href="https://github.com/actions/setup-node/commit/f4a67bbeca970f103397d3d2b9462cf787cd2980"><code>f4a67bb</code></a> Only use <code>mirrorToken</code> in <code>getManifest</code> if it's provided (<a href="https://redirect.github.com/actions/setup-node/issues/1548">#1548</a>)</li> <li><a href="https://github.com/actions/setup-node/commit/0355742c943ddb13ca8a6b700f824231caa91e75"><code>0355742</code></a> Remove dummy NODE_AUTH_TOKEN export (<a href="https://redirect.github.com/actions/setup-node/issues/1558">#1558</a>)</li> <li>Additional commits viewable in <a href="https://github.com/actions/setup-node/compare/v6...v7">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
cec0fc249a |
[codex] Parallelize release verify workflow (#9168)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Releases publish the same app and package set that operators install, so release verification should keep full release-strength coverage. > - The release workflow currently verifies stable and canary releases with one serial job that typechecks, runs all tests, and builds. > - The PR workflow already proves the test surface can be split into grouped general suites and serialized shards without changing coverage. > - This pull request extracts the release verify work into a reusable workflow and fans out the independent lanes. > - The benefit is faster stable and canary release verification while preserving the existing publish and preview gates. ## Linked Issues or Issue Description No public GitHub issue exists for this CI improvement. **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Release verification spends most of its wall time in a single serial test step even though the same stable test surface is already partitioned for PR CI. Stable dispatches and master-push canaries therefore wait on one long runner after setup, typecheck, tests, and build run sequentially. **Proposed solution** Add a reusable release verification workflow with parallel typecheck, grouped general tests, serialized test shards, and build lanes. Have both stable and canary release verification call it with the ref they need to verify. **Alternatives considered** Keeping the serial `pnpm test:run` job preserves the old shape but keeps stable and canary releases waiting on one long runner. Skipping verification when a source SHA already has green CI would be faster, but adds stale-check and lookup risk beyond this change. **Roadmap alignment** No overlapping item found in `ROADMAP.md`; this is release CI maintenance. **Additional context** The new workflow keeps the release-strength full `pnpm -r typecheck`, uses the existing stable test grouping/sharding entry points, and leaves publish/preview jobs unchanged. ## What Changed - Added `.github/workflows/release-verify.yml` as a `workflow_call` workflow accepting a `ref` input. - Split release verification into parallel `typecheck`, `general_tests`, `serialized_tests`, and `build` jobs with 20-minute lane timeouts. - Mirrored the PR workflow's stable test partition: `general-server` shards 1-3, `general-workspaces-a`, `general-workspaces-b`, and four serialized shards. - Replaced `release.yml` `verify_canary` and `verify_stable` job bodies with calls to the reusable workflow while leaving publish and preview jobs unchanged. - Added a Node test that guards the release workflow delegation and split verify surface. ## Verification - `actionlint 1.7.12 .github/workflows/release.yml .github/workflows/release-verify.yml` - `node ./scripts/release-package-map.mjs check` - `node --test ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs` - `git diff --check` ## Risks - Release verification now starts more jobs per release event, increasing total runner setup/install minutes. This matches the existing PR CI tradeoff and should reduce release wall time substantially. - The called workflow checks out the requested ref shallowly. That is intentional for verify lanes; publish and preview jobs still retain their existing full-history checkouts. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-class coding agent in local tool-use mode with shell execution, repository editing, GitHub connector access, and medium reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |