Commit Graph
4 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 4265cb3a2b fix(ci): cache native server integration builds in release verification (#15619)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Release verification must run the same server integration coverage
as PR verification.
> - Three server suites use real Rust Runner binaries.
> - PR checks already run them with the shared Rust dependency cache,
but release checks cold-build them inside ordinary server test shards.
> - A cold Dot Runner build can consume its entire setup deadline before
any test runs.
> - This pull request gives release verification a required cached
native integration lane and preserves every test.

## Linked Issues or Issue Description

**What happened?**

On master commit `1881894973a2b25838d8abed9bd8aeebc3af4441`, [Cloud
readiness run
37837810707](https://github.com/paperclipai/paperclip/actions/runs/37837810707)
failed in server shard 7. Dot Runner's `cargo build --release` exceeded
its 300-second child-process deadline. The other 1,640 tests in that
shard passed. The failed setup then tried to remove an undefined
temporary path and emitted a second error. Cargo output was captured as
an opaque buffer, which obscured build progress.

**What did you expect to happen?**

Build fixture binaries before tests in a lane with the existing Rust
dependency cache. Run all native integration tests and make their result
required for release readiness. A setup failure should keep its original
diagnostic.

**Steps to reproduce**

Run release verification on a clean runner. The existing
`general-server-without-chat` group retains the three Cargo-backed
suites outside the PR workflow. The Dot suite can time out during a cold
release build. Run the new workflow and partition tests against the
previous source to reproduce five routing and coverage failures
deterministically.

**Version / commit**

Observed at `1881894973a2b25838d8abed9bd8aeebc3af4441`; this change is
based on `ed6abbf158b`.

**Deployment mode**

GitHub Actions Release and Cloud readiness verification. No application
runtime or deployment action changes.

Related work: #15581 also edits Dot integration tests for onboarding
behavior. It does not repair release test routing. Searches found no
open PR for this failure.

## What Changed

- Add an explicit server test group that excludes the dedicated chat and
native suites. Preserve the existing PR and local test groups.
- Run all three native server suites in one required matrix lane with
the existing trusted Rust dependency cache.
- Build debug and release fixture binaries in a visible step with a
10-minute limit before tests start. Keep Cargo freshness checks and the
existing test deadlines.
- Stream Cargo diagnostics and clean up safely when Dot setup stops
before creating its temporary directory.
- Test complete, non-overlapping partitions for Release, Cloud
readiness, other callers, and existing PR/local groups. Check cache
restrictions, build order, and the source verification dependency.

## Verification

- Passed 44 workflow and partition tests: `node --test
scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- The same tests fail in five relevant cases against the previous
workflow and selector. They pass after this correction.
- An independent review repeated all 44 tests successfully.
- Passed `actionlint .github/workflows/release-verify.yml`, `git diff
--check`, and a secret scan.
- Passed all 381 workflow policy tests: `node --test
'.github/scripts/tests/*.test.mjs'`.
- Passed full `pnpm build` in a clean worktree.
- Built debug and release fixture binaries, then passed all 32 tests in
the three real native integration suites: `pnpm test:run:general --
--group general-server-native-runner` (49.5 seconds, no skips).
- Passed full `pnpm -r typecheck`.
- Injected a synthetic Cargo setup failure. The suite reports that
failure without the secondary undefined-path cleanup error.
- Exact-head CI completed: 53 successful checks, including all server
shards, both native Runner lanes, build, typecheck, browser tests, and
the canary dry run. Two optional Storybook checks were skipped.
- Started the duplicate local `pnpm test:run` aggregate and stopped it
after exact-head CI passed. No completed local aggregate result is
claimed. All test processes owned by that run exited.
- Greptile scored the final head
`ffe5ab5a28af2dfea9db1c1952c652f070bf52f8` at 5/5 with no review
threads.
- `Superagent Supply Chain Scan` is neutral, not passed: it only
supports exact dependency pin replacements and cannot verify structural
edits to `.github/workflows/release-verify.yml`. It reported no
annotations. Independent source review, Greptile, actionlint, and the
workflow policy tests cover this structural change.
- GitHub still requires a code-owner review for the workflow files.
There are no merge conflicts.

## Risks

- Adds one CI matrix job, which increases concurrent runner demand. The
existing cache writer and trust restrictions stay in place.
- Missing dependencies still compile from the lockfile. A build that
exceeds its step limit fails visibly; failed tests still block source
verification.
- Workflow execution uses the new lane after merge. PR tests validate
the workflow contract and run the existing native test lane.
- The supply chain scanner cannot analyze this workflow structure. Its
neutral result is an explicit coverage limit; it is not counted as a
successful scan.
- No runtime, schema, deployment, publication, credential, or
test-timeout changes.

## Model Used

OpenAI Codex, based on GPT-6, with repository inspection, code
execution, and independent agent review. The exact deployed model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 14:33:10 -07:00
1d23cb6962 ci: pin a checkout-independent Rust cache path so PR lanes hit master's cache (#14394)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip ships a native Runner binary, written in Rust, and seven
CI lanes build it on every pull request
> - `Canary Dry Run` is the slowest check on every green PR run, and
most of its time is `cargo build --release` on third-party crates
> - Master saves a Rust dependency cache for these lanes, but every PR
lane logs `No cache found` and compiles every crate from zero
> - The cache key matches, but GitHub also compares a hash of the
absolute cache paths, and the master writer (RunsOn fleet,
`/home/runner/_work/...`) and the PR readers (GitHub-hosted,
`/home/runner/work/...`) hash different paths
> - This pull request gives both sides a checkout-independent workspace
path, so the hashes match and the PR lanes restore master's cache
> - The benefit is about 2.5 minutes less wall clock per PR run and
about 18 fewer runner-minutes per run

## Linked Issues or Issue Description

No public issue exists for this problem. The description below follows
the enhancement template.

Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs
#13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the
workspace path, so none of them fixes this miss.

**What existing behavior does this improve?**

The `Swatinem/rust-cache` restore step in the PR workflow lanes that
build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release
Registry`, and the four `Verify Paperclip Runner` lanes.

**Subsystem affected**

CI workflows under `.github/workflows/`, their guard tests under
`.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`.

**Current behavior**

Every PR lane logs `No cache found` although master holds an entry with
the exact key. Run 36424309181 computed
`v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master
holds a 678 MB entry with that key. GitHub matches a cache entry on the
key and on a version hash of the absolute paths in the cache. The master
writer runs on the RunsOn fleet, where the checkout is
`/home/runner/_work/paperclip/paperclip`. The PR readers run on
GitHub-hosted `ubuntu-latest`, where the checkout is
`/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…`
is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The
`work` paths hash to `1656e9ee…`. The key can never match, so each lane
compiles every third-party crate again.

**Proposed behavior**

The writer and the readers pass the same checkout-independent path to
`rust-cache`. Both runner layouts then produce the same version hash,
and the PR lanes restore master's cache.

**Reason and benefit**

`Canary Dry Run` takes 533s on a green run. 251s of that is dependency
compilation that a warm cache removes. Seven lanes pay this cost in
every PR run.

**Breaking changes**

None. This change affects CI only.

## What Changed

- Add a `Pin the Runner Rust workspace path` step before `rust-cache` in
the master writer (`release-verify.yml`, typecheck and runner lanes) and
in all four PR readers (`pr-trusted.yml`). The step creates the symlink
`$HOME/paperclip-runner-rust` →
`$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that
path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache`
resolves the input with `path.resolve`, which does not follow symlinks,
so both runner layouts now produce the same cache paths and the same
version hash. `$HOME` is `/home/runner` on both images, which is why the
`~/.cargo` paths already agreed.
- Bump the shared keys `release-runner-v1` → `release-runner-v2` and
`release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable
entries are then visibly orphaned instead of sharing a key with the new
ones.
- Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`,
and `typecheck-rust-cache`. They now require the pin step in both
workflows with identical text, placed before the cache step, and they
reject a `workspaces:` value that resolves under the checkout. The
`pr-runner-rust-cache` test checks all four PR reader jobs and fails if
a `rust-cache` step appears in a PR job that is not in its reader list.
- Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the
`release-runner-v2` key and to explain the pinned workspace path.

### Expected savings once merged

Measured from run 36424309181. "Removed" is the dependency-compile time
that a warm restore removes, minus about 18s to restore the 680 MB
entry. The fleet writer's own restore shows this cost.

| Lane | Today | Removed | Expected |
|---|---|---|---|
| Canary Dry Run | 533s | ~150s | ~380s |
| Typecheck + Release Registry | 462s | ~155s | ~305s |
| Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s |
| Verify Paperclip Runner (rust) | 400s | ~175s | ~225s |
| Build | 348s | ~130s | ~220s |
| Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s |
| Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s |

- Wall clock per PR run: about 533s → about 385s. That is about 2.5
minutes faster to a green check set. `Canary Dry Run` stays the longest
check. The rest is the non-cargo work in `release.sh` (standalone
package builds ~30s, publish-payload preview ~73s).
- Runner time: about 18 runner-minutes saved per PR run across the seven
lanes.
- The first master push after merge compiles from zero once in the fleet
writer (about 4 extra minutes on that one run) and saves the v2 entry.
Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key
still gives a partial hit, as before.

## Verification

- Run the guard tests for the three cache lanes:
`node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs
.github/scripts/tests/release-runner-cache.test.mjs
.github/scripts/tests/typecheck-rust-cache.test.mjs`
  Result: 21 pass, 0 fail.
- Run the full guard suite: `node --test
'.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3
failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox
temp-file ENOENT and fail the same way on the unmodified branch.
- Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`.
Result: 14 pass.
- Local archive test: create a tar from the `_work` layout through the
symlink (relative `../../../paperclip-runner-rust/target` entries, `tar
-P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract
it on the `work` layout. The files land in the real target directory and
the symlink stays intact.
- After merge, open any GitHub-hosted PR run and confirm that the seven
Rust lanes log `Restored from cache key ...release-runner-v2...` in
place of `No cache found`.

## Risks

- Low risk. The change touches CI workflows, their tests, and one doc
page. No product code changes.
- If the pin step fails, `rust-cache` reports a miss and the lane
compiles from zero, as it does today. The build does not break.
- Both runner layouts sit four levels under `/home/runner`, so the
relative `../../../` archive entries line up. The existing
`~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this
property. A future runner image with a different `$HOME` depth would
miss the cache but would not fail the job.
- `rm -rf "$pinned"` acts on the symlink itself (no trailing slash),
never on the checkout behind it. It only matters on a reused runner.
- Squash-merge note: the branch carries commits by `Bender (Fable)`. Add
`Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the
squash body to keep that authorship.

## Model Used

- Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude
Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool
use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache
listings, and PR operations.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
ticket id. The agent execution workspace fixed this branch name, so I
cannot rename it. Squash-merge drops the branch name.
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
2026-09-30 16:04:23 -07:00
Devin FoleyandPaperclip d2e940f4c1 ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys verified images from merged source commits.
> - Cloud readiness waits for every release verification check.
> - Runner verification currently runs long TypeScript tests before Rust
checks.
> - These checks can run on independent runners with their own build
directories.
> - This PR runs them in parallel while preserving all checks and the
shared dependency cache.

## Linked Issues or Issue Description

Refs #13194. Related prior work: #13142 and #13259. A search found no
duplicate parallel release-check change.

**What existing behavior does this improve?**

Time from merge to Cloud source verification and deployment readiness.

**Current behavior**

Recent successful runs take roughly 13 minutes from merge to deployable.
In run 34705914878, Runner verification took 11m23s. Protocol tests
finished before Rust tests and API authority checks started.

**Proposed behavior**

Run protocol and Rust verification in two matrix jobs. Cloud readiness
still requires both jobs to pass.

**Reason and benefit**

Remove the serial dependency between independent checks. Expected
improvement is about 2–3 minutes on a typical cached run, until the
image build or server tests become the longest job. This is an estimate;
post-merge timing will confirm it.

**Breaking changes**

Individual release Runner job names gain a lane suffix. Cloud source and
readiness marker names stay the same. PR runner routing is unchanged.

## What Changed

- Split release Runner checks into protocol and Rust lanes. Keep every
constituent of `check:all` exactly once.
- Restore the existing Rust dependency cache in both lanes. Allow only
the Rust lane to save it after warming both build profiles.
- Add coverage and cache authorization regressions. Document the
parallel verification and single cache writer.

## Verification

- Passed 477 workflow and source-verification tests with `node --test
.github/scripts/tests/*.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- Passed `actionlint`, `git diff --check`, and the private AWS routing
regression suite.
- Passed local `pnpm -r typecheck` and the standalone `check:runner &&
check:api-authority` lane, including all 1,671 API tests before the
protocol lane had built TypeScript output.
- The broad local protocol run under Node 25 had four failures. The two
affected files passed under CI's Node 24.19.0: 67 passed, 6 platform
skips.
- Local `pnpm test:run` aborted when disk space ran out; local `pnpm
build` could not run afterward. These are local verification limits.
[Linux CI run
34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424)
passed full typecheck, all grouped tests, native verification, build,
release dry run, and browser checks. Native protocol CI passed 1,986
tests, plus 1,671 API tests and the Rust suites.
- Latest-head Greptile is 5/5 with no open findings. All 33 current-head
checks are successful or intentionally skipped.

## Risks

- Uses one additional short-lived verification runner per release
verification. The existing AWS exact-master restriction remains in
place.
- The Rust lane warms debug dependencies so its cache save also serves
protocol tests. Both lanes always rebuild workspace code.
- A workflow regression could omit a check. The new coverage test
compares the matrix checks directly with `check:all`; Cloud readiness
depends on the complete reusable workflow.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:33:20 -07:00
Devin FoleyandPaperclip 9e970df4c5 ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:10:36 -07:00