## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip ships a native Runner binary, written in Rust, and seven CI lanes build it on every pull request > - `Canary Dry Run` is the slowest check on every green PR run, and most of its time is `cargo build --release` on third-party crates > - Master saves a Rust dependency cache for these lanes, but every PR lane logs `No cache found` and compiles every crate from zero > - The cache key matches, but GitHub also compares a hash of the absolute cache paths, and the master writer (RunsOn fleet, `/home/runner/_work/...`) and the PR readers (GitHub-hosted, `/home/runner/work/...`) hash different paths > - This pull request gives both sides a checkout-independent workspace path, so the hashes match and the PR lanes restore master's cache > - The benefit is about 2.5 minutes less wall clock per PR run and about 18 fewer runner-minutes per run ## Linked Issues or Issue Description No public issue exists for this problem. The description below follows the enhancement template. Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs #13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the workspace path, so none of them fixes this miss. **What existing behavior does this improve?** The `Swatinem/rust-cache` restore step in the PR workflow lanes that build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release Registry`, and the four `Verify Paperclip Runner` lanes. **Subsystem affected** CI workflows under `.github/workflows/`, their guard tests under `.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`. **Current behavior** Every PR lane logs `No cache found` although master holds an entry with the exact key. Run 36424309181 computed `v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master holds a 678 MB entry with that key. GitHub matches a cache entry on the key and on a version hash of the absolute paths in the cache. The master writer runs on the RunsOn fleet, where the checkout is `/home/runner/_work/paperclip/paperclip`. The PR readers run on GitHub-hosted `ubuntu-latest`, where the checkout is `/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…` is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The `work` paths hash to `1656e9ee…`. The key can never match, so each lane compiles every third-party crate again. **Proposed behavior** The writer and the readers pass the same checkout-independent path to `rust-cache`. Both runner layouts then produce the same version hash, and the PR lanes restore master's cache. **Reason and benefit** `Canary Dry Run` takes 533s on a green run. 251s of that is dependency compilation that a warm cache removes. Seven lanes pay this cost in every PR run. **Breaking changes** None. This change affects CI only. ## What Changed - Add a `Pin the Runner Rust workspace path` step before `rust-cache` in the master writer (`release-verify.yml`, typecheck and runner lanes) and in all four PR readers (`pr-trusted.yml`). The step creates the symlink `$HOME/paperclip-runner-rust` → `$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache` resolves the input with `path.resolve`, which does not follow symlinks, so both runner layouts now produce the same cache paths and the same version hash. `$HOME` is `/home/runner` on both images, which is why the `~/.cargo` paths already agreed. - Bump the shared keys `release-runner-v1` → `release-runner-v2` and `release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable entries are then visibly orphaned instead of sharing a key with the new ones. - Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`, and `typecheck-rust-cache`. They now require the pin step in both workflows with identical text, placed before the cache step, and they reject a `workspaces:` value that resolves under the checkout. The `pr-runner-rust-cache` test checks all four PR reader jobs and fails if a `rust-cache` step appears in a PR job that is not in its reader list. - Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the `release-runner-v2` key and to explain the pinned workspace path. ### Expected savings once merged Measured from run 36424309181. "Removed" is the dependency-compile time that a warm restore removes, minus about 18s to restore the 680 MB entry. The fleet writer's own restore shows this cost. | Lane | Today | Removed | Expected | |---|---|---|---| | Canary Dry Run | 533s | ~150s | ~380s | | Typecheck + Release Registry | 462s | ~155s | ~305s | | Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s | | Verify Paperclip Runner (rust) | 400s | ~175s | ~225s | | Build | 348s | ~130s | ~220s | | Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s | | Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s | - Wall clock per PR run: about 533s → about 385s. That is about 2.5 minutes faster to a green check set. `Canary Dry Run` stays the longest check. The rest is the non-cargo work in `release.sh` (standalone package builds ~30s, publish-payload preview ~73s). - Runner time: about 18 runner-minutes saved per PR run across the seven lanes. - The first master push after merge compiles from zero once in the fleet writer (about 4 extra minutes on that one run) and saves the v2 entry. Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key still gives a partial hit, as before. ## Verification - Run the guard tests for the three cache lanes: `node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/release-runner-cache.test.mjs .github/scripts/tests/typecheck-rust-cache.test.mjs` Result: 21 pass, 0 fail. - Run the full guard suite: `node --test '.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3 failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox temp-file ENOENT and fail the same way on the unmodified branch. - Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`. Result: 14 pass. - Local archive test: create a tar from the `_work` layout through the symlink (relative `../../../paperclip-runner-rust/target` entries, `tar -P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract it on the `work` layout. The files land in the real target directory and the symlink stays intact. - After merge, open any GitHub-hosted PR run and confirm that the seven Rust lanes log `Restored from cache key ...release-runner-v2...` in place of `No cache found`. ## Risks - Low risk. The change touches CI workflows, their tests, and one doc page. No product code changes. - If the pin step fails, `rust-cache` reports a miss and the lane compiles from zero, as it does today. The build does not break. - Both runner layouts sit four levels under `/home/runner`, so the relative `../../../` archive entries line up. The existing `~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this property. A future runner image with a different `$HOME` depth would miss the cache but would not fail the job. - `rm -rf "$pinned"` acts on the symlink itself (no trailing slash), never on the checkout behind it. It only matters on a reused runner. - Squash-merge note: the branch carries commits by `Bender (Fable)`. Add `Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the squash body to keep that authorship. ## Model Used - Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache listings, and PR operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal ticket id. The agent execution workspace fixed this branch name, so I cannot rename it. Squash-merge drops the branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
13 KiB
Cloud build readiness
The Cloud readiness workflow starts for every master push and retains the
versioned Cloud source verified v1 job. It calls the full Release Verify
workflow for that exact commit, including typecheck, builds, general and
serialized tests, and Runner verification. The source proof depends on every
source check and fails closed if verification fails, is cancelled, or is skipped.
The recurring public -cloud publisher and its Cloud deployable v1 gate are
retired. The workflow no longer builds a legacy image or waits for one.
Keep its filename and source-proof job name stable: npm canary publication and
downstream image composers consume that exact contract.
Standard images still publish independently through docker.yml. The
cloud-migrator-artifacts.yml workflow still publishes signed exact-source
migrators independently. A downstream composer must verify those artifacts,
build and test its own image, and record separate deployment readiness.
Verification runs outside the full npm release's concurrency group. Different
commits have independent groups. The npm canary release reuses the source proof
for its exact master push instead of starting another Release Verify run.
Stable releases and candidate-branch betas still run full verification.
The canary consumer requires the expected workflow ID and path, upstream source repository, master push event, full SHA, and a successful job in the latest run attempt. It checks the run again after reading the jobs to reject a concurrent rerun. Missing proof waits for up to 45 minutes; failed, skipped, cancelled, ambiguous, or mismatched proof cannot authorize publication. API failures fail closed. If a source check fails, fix it and rerun Cloud readiness before retrying the release. Use Re-run all jobs when a later attempt did not rerun the source proof; an earlier attempt's successful job is not accepted. This avoids duplicate test jobs on standard runners. Measure queue time to assess the timing gain.
Release verification spreads the general server suites across ten standard hosted runners, with the long chat suite split separately across three jobs. Each server job still runs one test worker. The partition covers every suite exactly once; normal PR and local test groups keep their existing shape. More jobs increase concurrent runner demand, so compare queue time as well as test duration.
All release verification installs, including the Runner scorer and chaos evals, allow pnpm to refresh an outdated lockfile. Contributor PRs leave lockfile updates to the separate refresh bot, so a dependency-changing master commit can arrive before that bot's PR merges. Verification must install and test that commit without waiting for another merge. The generated lockfile stays in the job's workspace; these checks do not commit it back to the repository.
Consumer contract and retirement boundary
Accept Cloud source verified v1 only from the latest attempt of the canonical
cloud-readiness.yml master push for the expected repository identity and full
source SHA. Check the job itself and reject failed, skipped, cancelled, or
ambiguous proof. This signal verifies source only. It creates no release record,
certifies no composed image, and deploys no instance.
Downstream deployment consumers must separately verify the standard image's immutable digest and attestation, the exact-source migrator's signature and integrity, migration compatibility, their own image composition, and target health. Order automatic candidates by master ancestry, not completion time.
Merge this retirement only after all active automatic deployment consumers use
the standard-image composition contract. A consumer still selecting
Cloud deployable v1 will stop advancing at the last legacy-ready commit.
Do not rename the source proof to the old readiness name or weaken a consumer
check to hide that dependency.
Existing release records, immutable image digests, migrator artifacts, and
registry tags are retained for rollback. No registry deletion or live deployment
is part of this change. Explicit release.yml preview requests still use the
legacy cloud Dockerfile target for a specified source commit. They do not
restart recurring legacy publication. Keep that compatibility path until its
operator consumers migrate separately.
The old nightly-cloud, beta-cloud, latest-cloud, and canary-cloud aliases
stop advancing. Self-hosted standard release aliases continue unchanged. A
rollback to an already published image needs no rebuild; restoring recurring
legacy publication would require reverting the publisher retirement.
Timing and rollout
The reusable Runner chaos workflow scopes concurrency to the caller workflow and source ref. Cloud readiness, stable verification, and standalone evals can verify the same commit at the same time. They must not cancel each other's required test job.
Measure the complete path from a master merge to a healthy target running that exact commit. Keep readiness and deployment as separate milestones:
| Milestone | Evidence | Elapsed time starts at |
|---|---|---|
| Merge | Merged PR timestamp and full merge commit SHA | Merge |
| Image available | Successful full-SHA image publication and verification | Merge |
| Source verified | Successful Cloud source verified v1 job in the accepted push run and attempt |
Merge |
| Composed image ready | Downstream composition verification and publication succeed | Merge |
| Canary healthy | Deployment consumer's canary health gate confirms the target commit | Merge |
| Fleet complete | Campaign succeeds for all eligible targets at that commit | Merge |
Record the source SHA, workflow run ID and attempt, readiness job completion time, and deployment campaign identity together. Verify the run against the consumer contract above. A manual dispatch can test wiring, but its timestamp does not measure automatic merge-to-deploy latency. A preparation-only run resolves artifacts without deploying a target and must not be counted as a successful deployment.
Record queue time and the image, source-verification, migrator, and composition durations separately. The slowest prerequisite determines readiness; shortening an already faster prerequisite may have no effect on the total. After readiness, measure consumer discovery delay, artifact resolution, canary health, and fleet rollout. An automatic consumer that still waits for the full npm canary publication has that queue on its critical path even if cloud artifacts are ready earlier.
For a target health measurement, confirm the deployed source SHA as well as service health. A proxy health response alone may describe the control plane while the tenant still runs the previous image. Report the eligible target count, excluded or sleeping targets, retries, and failures with the fleet result. Record runner queue conditions and cache state; one warm or cold run is a sample, not a latency guarantee.
Reserved AWS verification capacity
AWS_POST_MERGE_CI_ENABLED=true routes cloud source verification and
exact-master migrator preparation to the
paperclip-post-merge runner group. The separate Fleet label is
runs-on/fleet=paperclip-post-merge-x64/env=public-ci. Its 36 reserved slots use
the same four-vCPU, 16-GiB machines as approved PR jobs. PR capacity is reduced
to 64; the separately provisioned image capacity is unchanged by this retirement.
This keeps PR bursts from consuming every post-merge verification slot.
Every selector checks the canonical repository name and ID, master ref, and a
push or manual event. Reusable verification also requires inputs.ref to equal
that event's github.sha. The migrator route requires cloud-migrator and
inputs.source_ref == github.sha. Branch/tag refs, PR events, arbitrary preview
sources, and missing or disabled switches use GitHub-hosted runners. If another
merge lands before a migrator dispatch resolves master, the older source uses
GitHub-hosted runners too. npm publication always remains GitHub-hosted to keep
its trusted-publisher identity.
Before enabling the switch, deploy the separate Fleet and restrict its GitHub
runner group to repository ID 1170821064 and these workflows at
refs/heads/master: cloud-readiness.yml,
release-verify.yml, runner-chaos-evals.yml, and release.yml. Do not authorize
PR-controlled workflow versions. The direct migrator producer always uses
GitHub-hosted runners and needs no AWS runner-group authorization. PR placement retains its independent pinned
workflow and six-account author/actor allowlist.
Disable the switch and rerun the whole workflow to restore GitHub-hosted placement. Assigned jobs keep their original runners. Readiness requirements, source checks, and npm integrity checks are unchanged.
Retired AWS cloud build routing
AWS_CLOUD_BUILDS_ENABLED and the paperclip-cloud-build runner group no longer
route a public image job after this retirement. This source change does not
delete runner groups, Fleets, credentials, registry images, or cache tags. Review
shared infrastructure ownership separately before removing those resources.
Typecheck Rust dependency cache
Source verification's typecheck job builds the native Runner binary through the
server's prepare:runner-vendor command. It restores and saves compiled Rust
dependencies only for canonical master pushes that verify the event's exact SHA.
The release-typecheck-v2 cache is separate from Runner verification because
those jobs compile different profiles. The pinned toolchain is selected before
cache lookup, and the cache action receives the Runner crate through the same
checkout-independent $HOME/paperclip-runner-rust path as the Runner lanes
(see RELEASE-AUTOMATION-SETUP.md). Workspace crates and installed cargo
binaries are excluded, and all typechecks still execute. A missing or
invalidated cache triggers compilation.
pnpm dependency store cache
The Refresh Lockfile workflow does not cache the pnpm store. Its resolution-only command does not download packages and can save an empty default-branch cache before full install jobs finish. The PR policy job also leaves store caching off.
PR install jobs restore the pnpm store without saving it. They hash the checked-in
lockfile before downloading the policy job's regenerated lockfile, matching the
key format used by master install jobs. A same-OS, same-architecture pnpm fallback
can reuse older package downloads when the exact key is absent. Each job still
installs with --frozen-lockfile against the policy artifact when one exists;
cache contents do not select dependency versions. A cache miss downloads packages
normally. New PR-only dependencies may be downloaded again on each PR run until
master populates a cache that contains them.
This avoids storing a full dependency archive under every PR merge ref. Those
copies competed with the Rust caches for the repository's storage limit. Keep
master cache writes enabled so trusted post-merge installs refresh shared stores.
After activating the new trusted workflow pin, verify cache restores and package
reuse in an allowlisted PR, and verify that no new node-cache- entries appear
under its refs/pull/<number>/merge ref. Existing copies can expire normally.
The repository cache storage ceiling is managed in GitHub Settings, separately from this workflow. Check it with:
gh api repos/paperclipai/paperclip/actions/cache/storage-limit
Increasing the repository limit above 10 GB can require an organization owner to raise the maximum in organization Settings → Actions → General first. Repository administration access alone cannot override that maximum. Paid cache storage also requires a payment method and sufficient Actions Cache Storage budget; see the GitHub cache storage documentation. Preserve populated master pnpm and Rust caches when inspecting pressure.
After deploying this correction, remove any existing empty default-branch entry for the current lockfile key. List cache IDs, branches, and archive sizes first:
gh api --paginate 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&key=node-cache-Linux-x64-pnpm-&per_page=100' \
--jq '.actions_caches[] | {id, ref, key, size_in_bytes}'
Match the key and upload size against the cache-creation job's logs. The
September 11 incident was cache ID 7559920987, a 216-byte archive. This guarded
command deletes only that observed entry. It leaves a populated replacement or
an entry on another branch untouched, and does nothing if the old ID is absent:
bad_cache_id=7559920987
bad_cache_key=node-cache-Linux-x64-pnpm-c3096ecb02a34aaa9782baaadafcb731510e1dba10dd661618c3a2ee91e58fa5
entries="$(gh api --paginate --slurp 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&per_page=100')"
if printf '%s\n' "$entries" | jq -e --argjson id "$bad_cache_id" --arg key "$bad_cache_key" '
[.[].actions_caches[] | select(.id == $id)] |
length == 1 and .[0].ref == "refs/heads/master" and
.[0].key == $key and .[0].size_in_bytes == 216
' >/dev/null; then
gh api --method DELETE "repos/paperclipai/paperclip/actions/caches/$bad_cache_id"
fi
A subsequent master install can populate the missing entry. Check the saved archive size and package reuse in install logs; a cache hit alone does not prove that the entry contains dependencies.