mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip ships a native Runner binary, written in Rust, and seven CI lanes build it on every pull request > - `Canary Dry Run` is the slowest check on every green PR run, and most of its time is `cargo build --release` on third-party crates > - Master saves a Rust dependency cache for these lanes, but every PR lane logs `No cache found` and compiles every crate from zero > - The cache key matches, but GitHub also compares a hash of the absolute cache paths, and the master writer (RunsOn fleet, `/home/runner/_work/...`) and the PR readers (GitHub-hosted, `/home/runner/work/...`) hash different paths > - This pull request gives both sides a checkout-independent workspace path, so the hashes match and the PR lanes restore master's cache > - The benefit is about 2.5 minutes less wall clock per PR run and about 18 fewer runner-minutes per run ## Linked Issues or Issue Description No public issue exists for this problem. The description below follows the enhancement template. Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs #13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the workspace path, so none of them fixes this miss. **What existing behavior does this improve?** The `Swatinem/rust-cache` restore step in the PR workflow lanes that build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release Registry`, and the four `Verify Paperclip Runner` lanes. **Subsystem affected** CI workflows under `.github/workflows/`, their guard tests under `.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`. **Current behavior** Every PR lane logs `No cache found` although master holds an entry with the exact key. Run 36424309181 computed `v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master holds a 678 MB entry with that key. GitHub matches a cache entry on the key and on a version hash of the absolute paths in the cache. The master writer runs on the RunsOn fleet, where the checkout is `/home/runner/_work/paperclip/paperclip`. The PR readers run on GitHub-hosted `ubuntu-latest`, where the checkout is `/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…` is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The `work` paths hash to `1656e9ee…`. The key can never match, so each lane compiles every third-party crate again. **Proposed behavior** The writer and the readers pass the same checkout-independent path to `rust-cache`. Both runner layouts then produce the same version hash, and the PR lanes restore master's cache. **Reason and benefit** `Canary Dry Run` takes 533s on a green run. 251s of that is dependency compilation that a warm cache removes. Seven lanes pay this cost in every PR run. **Breaking changes** None. This change affects CI only. ## What Changed - Add a `Pin the Runner Rust workspace path` step before `rust-cache` in the master writer (`release-verify.yml`, typecheck and runner lanes) and in all four PR readers (`pr-trusted.yml`). The step creates the symlink `$HOME/paperclip-runner-rust` → `$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache` resolves the input with `path.resolve`, which does not follow symlinks, so both runner layouts now produce the same cache paths and the same version hash. `$HOME` is `/home/runner` on both images, which is why the `~/.cargo` paths already agreed. - Bump the shared keys `release-runner-v1` → `release-runner-v2` and `release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable entries are then visibly orphaned instead of sharing a key with the new ones. - Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`, and `typecheck-rust-cache`. They now require the pin step in both workflows with identical text, placed before the cache step, and they reject a `workspaces:` value that resolves under the checkout. The `pr-runner-rust-cache` test checks all four PR reader jobs and fails if a `rust-cache` step appears in a PR job that is not in its reader list. - Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the `release-runner-v2` key and to explain the pinned workspace path. ### Expected savings once merged Measured from run 36424309181. "Removed" is the dependency-compile time that a warm restore removes, minus about 18s to restore the 680 MB entry. The fleet writer's own restore shows this cost. | Lane | Today | Removed | Expected | |---|---|---|---| | Canary Dry Run | 533s | ~150s | ~380s | | Typecheck + Release Registry | 462s | ~155s | ~305s | | Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s | | Verify Paperclip Runner (rust) | 400s | ~175s | ~225s | | Build | 348s | ~130s | ~220s | | Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s | | Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s | - Wall clock per PR run: about 533s → about 385s. That is about 2.5 minutes faster to a green check set. `Canary Dry Run` stays the longest check. The rest is the non-cargo work in `release.sh` (standalone package builds ~30s, publish-payload preview ~73s). - Runner time: about 18 runner-minutes saved per PR run across the seven lanes. - The first master push after merge compiles from zero once in the fleet writer (about 4 extra minutes on that one run) and saves the v2 entry. Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key still gives a partial hit, as before. ## Verification - Run the guard tests for the three cache lanes: `node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/release-runner-cache.test.mjs .github/scripts/tests/typecheck-rust-cache.test.mjs` Result: 21 pass, 0 fail. - Run the full guard suite: `node --test '.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3 failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox temp-file ENOENT and fail the same way on the unmodified branch. - Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`. Result: 14 pass. - Local archive test: create a tar from the `_work` layout through the symlink (relative `../../../paperclip-runner-rust/target` entries, `tar -P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract it on the `work` layout. The files land in the real target directory and the symlink stays intact. - After merge, open any GitHub-hosted PR run and confirm that the seven Rust lanes log `Restored from cache key ...release-runner-v2...` in place of `No cache found`. ## Risks - Low risk. The change touches CI workflows, their tests, and one doc page. No product code changes. - If the pin step fails, `rust-cache` reports a miss and the lane compiles from zero, as it does today. The build does not break. - Both runner layouts sit four levels under `/home/runner`, so the relative `../../../` archive entries line up. The existing `~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this property. A future runner image with a different `$HOME` depth would miss the cache but would not fail the job. - `rm -rf "$pinned"` acts on the symlink itself (no trailing slash), never on the checkout behind it. It only matters on a reused runner. - Squash-merge note: the branch carries commits by `Bender (Fable)`. Add `Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the squash body to keep that authorship. ## Model Used - Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache listings, and PR operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal ticket id. The agent execution workspace fixed this branch name, so I cannot rename it. Squash-merge drops the branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
234 lines
13 KiB
Markdown
234 lines
13 KiB
Markdown
# Cloud build readiness
|
|
|
|
The `Cloud readiness` workflow starts for every master push and retains the
|
|
versioned `Cloud source verified v1` job. It calls the full `Release Verify`
|
|
workflow for that exact commit, including typecheck, builds, general and
|
|
serialized tests, and Runner verification. The source proof depends on every
|
|
source check and fails closed if verification fails, is cancelled, or is skipped.
|
|
|
|
The recurring public `-cloud` publisher and its `Cloud deployable v1` gate are
|
|
retired. The workflow no longer builds a legacy image or waits for one.
|
|
Keep its filename and source-proof job name stable: npm canary publication and
|
|
downstream image composers consume that exact contract.
|
|
|
|
Standard images still publish independently through `docker.yml`. The
|
|
`cloud-migrator-artifacts.yml` workflow still publishes signed exact-source
|
|
migrators independently. A downstream composer must verify those artifacts,
|
|
build and test its own image, and record separate deployment readiness.
|
|
|
|
Verification runs outside the full npm release's concurrency group. Different
|
|
commits have independent groups. The npm canary release reuses the source proof
|
|
for its exact master push instead of starting another `Release Verify` run.
|
|
Stable releases and candidate-branch betas still run full verification.
|
|
|
|
The canary consumer requires the expected workflow ID and path, upstream source
|
|
repository, master push event, full SHA, and a successful job in the latest run
|
|
attempt. It checks the run again after reading the jobs to reject a concurrent
|
|
rerun. Missing proof waits for up to 45 minutes; failed, skipped, cancelled,
|
|
ambiguous, or mismatched proof cannot authorize publication. API failures fail
|
|
closed. If a source check fails, fix it and rerun Cloud readiness before retrying
|
|
the release. Use **Re-run all jobs** when a later attempt did not rerun the source
|
|
proof; an earlier attempt's successful job is not accepted. This avoids duplicate
|
|
test jobs on standard runners. Measure queue time to assess the timing gain.
|
|
|
|
Release verification spreads the general server suites across ten standard hosted
|
|
runners, with the long chat suite split separately across three jobs. Each server
|
|
job still runs one test worker. The partition covers every suite exactly once;
|
|
normal PR and local test groups keep their existing shape. More jobs increase
|
|
concurrent runner demand, so compare queue time as well as test duration.
|
|
|
|
All release verification installs, including the Runner scorer and chaos evals,
|
|
allow pnpm to refresh an outdated lockfile. Contributor PRs leave lockfile updates
|
|
to the separate refresh bot, so a dependency-changing master commit can arrive
|
|
before that bot's PR merges. Verification must install and test that commit
|
|
without waiting for another merge. The generated lockfile stays in the job's
|
|
workspace; these checks do not commit it back to the repository.
|
|
|
|
## Consumer contract and retirement boundary
|
|
|
|
Accept `Cloud source verified v1` only from the latest attempt of the canonical
|
|
`cloud-readiness.yml` master push for the expected repository identity and full
|
|
source SHA. Check the job itself and reject failed, skipped, cancelled, or
|
|
ambiguous proof. This signal verifies source only. It creates no release record,
|
|
certifies no composed image, and deploys no instance.
|
|
|
|
Downstream deployment consumers must separately verify the standard image's
|
|
immutable digest and attestation, the exact-source migrator's signature and
|
|
integrity, migration compatibility, their own image composition, and target
|
|
health. Order automatic candidates by master ancestry, not completion time.
|
|
|
|
Merge this retirement only after all active automatic deployment consumers use
|
|
the standard-image composition contract. A consumer still selecting
|
|
`Cloud deployable v1` will stop advancing at the last legacy-ready commit.
|
|
Do not rename the source proof to the old readiness name or weaken a consumer
|
|
check to hide that dependency.
|
|
|
|
Existing release records, immutable image digests, migrator artifacts, and
|
|
registry tags are retained for rollback. No registry deletion or live deployment
|
|
is part of this change. Explicit `release.yml` preview requests still use the
|
|
legacy `cloud` Dockerfile target for a specified source commit. They do not
|
|
restart recurring legacy publication. Keep that compatibility path until its
|
|
operator consumers migrate separately.
|
|
|
|
The old `nightly-cloud`, `beta-cloud`, `latest-cloud`, and `canary-cloud` aliases
|
|
stop advancing. Self-hosted standard release aliases continue unchanged. A
|
|
rollback to an already published image needs no rebuild; restoring recurring
|
|
legacy publication would require reverting the publisher retirement.
|
|
|
|
## Timing and rollout
|
|
|
|
The reusable Runner chaos workflow scopes concurrency to the caller workflow
|
|
and source ref. Cloud readiness, stable verification, and standalone evals can
|
|
verify the same commit at the same time. They must not cancel each other's
|
|
required test job.
|
|
|
|
Measure the complete path from a master merge to a healthy target running that
|
|
exact commit. Keep readiness and deployment as separate milestones:
|
|
|
|
| Milestone | Evidence | Elapsed time starts at |
|
|
| --- | --- | --- |
|
|
| Merge | Merged PR timestamp and full merge commit SHA | Merge |
|
|
| Image available | Successful full-SHA image publication and verification | Merge |
|
|
| Source verified | Successful `Cloud source verified v1` job in the accepted push run and attempt | Merge |
|
|
| Composed image ready | Downstream composition verification and publication succeed | Merge |
|
|
| Canary healthy | Deployment consumer's canary health gate confirms the target commit | Merge |
|
|
| Fleet complete | Campaign succeeds for all eligible targets at that commit | Merge |
|
|
|
|
Record the source SHA, workflow run ID and attempt, readiness job completion
|
|
time, and deployment campaign identity together. Verify the run against the
|
|
consumer contract above. A manual dispatch can test wiring, but its timestamp
|
|
does not measure automatic merge-to-deploy latency. A preparation-only run
|
|
resolves artifacts without deploying a target and must not be counted as a
|
|
successful deployment.
|
|
|
|
Record queue time and the image, source-verification, migrator, and composition durations
|
|
separately. The slowest prerequisite determines readiness; shortening an already
|
|
faster prerequisite may have no effect on the total. After readiness, measure
|
|
consumer discovery delay, artifact resolution, canary health, and fleet rollout.
|
|
An automatic consumer that still waits for the full npm canary publication has
|
|
that queue on its critical path even if cloud artifacts are ready earlier.
|
|
|
|
For a target health measurement, confirm the deployed source SHA as well as
|
|
service health. A proxy health response alone may describe the control plane
|
|
while the tenant still runs the previous image. Report the eligible target count,
|
|
excluded or sleeping targets, retries, and failures with the fleet result. Record
|
|
runner queue conditions and cache state; one warm or cold run is a sample, not a
|
|
latency guarantee.
|
|
|
|
## Reserved AWS verification capacity
|
|
|
|
`AWS_POST_MERGE_CI_ENABLED=true` routes cloud source verification and
|
|
exact-master migrator preparation to the
|
|
`paperclip-post-merge` runner group. The separate Fleet label is
|
|
`runs-on/fleet=paperclip-post-merge-x64/env=public-ci`. Its 36 reserved slots use
|
|
the same four-vCPU, 16-GiB machines as approved PR jobs. PR capacity is reduced
|
|
to 64; the separately provisioned image capacity is unchanged by this retirement.
|
|
This keeps PR bursts from consuming every post-merge verification slot.
|
|
|
|
Every selector checks the canonical repository name and ID, master ref, and a
|
|
push or manual event. Reusable verification also requires `inputs.ref` to equal
|
|
that event's `github.sha`. The migrator route requires `cloud-migrator` and
|
|
`inputs.source_ref == github.sha`. Branch/tag refs, PR events, arbitrary preview
|
|
sources, and missing or disabled switches use GitHub-hosted runners. If another
|
|
merge lands before a migrator dispatch resolves master, the older source uses
|
|
GitHub-hosted runners too. npm publication always remains GitHub-hosted to keep
|
|
its trusted-publisher identity.
|
|
|
|
Before enabling the switch, deploy the separate Fleet and restrict its GitHub
|
|
runner group to repository ID `1170821064` and these workflows at
|
|
`refs/heads/master`: `cloud-readiness.yml`,
|
|
`release-verify.yml`, `runner-chaos-evals.yml`, and `release.yml`. Do not authorize
|
|
PR-controlled workflow versions. The direct migrator producer always uses
|
|
GitHub-hosted runners and needs no AWS runner-group authorization. PR placement retains its independent pinned
|
|
workflow and six-account author/actor allowlist.
|
|
|
|
Disable the switch and rerun the whole workflow to restore GitHub-hosted
|
|
placement. Assigned jobs keep their original runners. Readiness requirements,
|
|
source checks, and npm integrity checks are unchanged.
|
|
|
|
## Retired AWS cloud build routing
|
|
|
|
`AWS_CLOUD_BUILDS_ENABLED` and the `paperclip-cloud-build` runner group no longer
|
|
route a public image job after this retirement. This source change does not
|
|
delete runner groups, Fleets, credentials, registry images, or cache tags. Review
|
|
shared infrastructure ownership separately before removing those resources.
|
|
|
|
### Typecheck Rust dependency cache
|
|
|
|
Source verification's typecheck job builds the native Runner binary through the
|
|
server's `prepare:runner-vendor` command. It restores and saves compiled Rust
|
|
dependencies only for canonical master pushes that verify the event's exact SHA.
|
|
The `release-typecheck-v2` cache is separate from Runner verification because
|
|
those jobs compile different profiles. The pinned toolchain is selected before
|
|
cache lookup, and the cache action receives the Runner crate through the same
|
|
checkout-independent `$HOME/paperclip-runner-rust` path as the Runner lanes
|
|
(see `RELEASE-AUTOMATION-SETUP.md`). Workspace crates and installed cargo
|
|
binaries are excluded, and all typechecks still execute. A missing or
|
|
invalidated cache triggers compilation.
|
|
|
|
### pnpm dependency store cache
|
|
|
|
The Refresh Lockfile workflow does not cache the pnpm store. Its resolution-only
|
|
command does not download packages and can save an empty default-branch cache
|
|
before full install jobs finish. The PR policy job also leaves store caching off.
|
|
|
|
PR install jobs restore the pnpm store without saving it. They hash the checked-in
|
|
lockfile before downloading the policy job's regenerated lockfile, matching the
|
|
key format used by master install jobs. A same-OS, same-architecture pnpm fallback
|
|
can reuse older package downloads when the exact key is absent. Each job still
|
|
installs with `--frozen-lockfile` against the policy artifact when one exists;
|
|
cache contents do not select dependency versions. A cache miss downloads packages
|
|
normally. New PR-only dependencies may be downloaded again on each PR run until
|
|
master populates a cache that contains them.
|
|
|
|
This avoids storing a full dependency archive under every PR merge ref. Those
|
|
copies competed with the Rust caches for the repository's storage limit. Keep
|
|
master cache writes enabled so trusted post-merge installs refresh shared stores.
|
|
After activating the new trusted workflow pin, verify cache restores and package
|
|
reuse in an allowlisted PR, and verify that no new `node-cache-` entries appear
|
|
under its `refs/pull/<number>/merge` ref. Existing copies can expire normally.
|
|
|
|
The repository cache storage ceiling is managed in GitHub Settings, separately
|
|
from this workflow. Check it with:
|
|
|
|
```sh
|
|
gh api repos/paperclipai/paperclip/actions/cache/storage-limit
|
|
```
|
|
|
|
Increasing the repository limit above 10 GB can require an organization owner to
|
|
raise the maximum in organization Settings → Actions → General first. Repository
|
|
administration access alone cannot override that maximum. Paid cache storage also
|
|
requires a payment method and sufficient Actions Cache Storage budget; see the
|
|
[GitHub cache storage documentation](https://docs.github.com/en/actions/reference/workflows-and-actions/dependency-caching#increasing-cache-size).
|
|
Preserve populated master pnpm and Rust caches when inspecting pressure.
|
|
|
|
After deploying this correction, remove any existing empty default-branch entry
|
|
for the current lockfile key. List cache IDs, branches, and archive sizes first:
|
|
|
|
```sh
|
|
gh api --paginate 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&key=node-cache-Linux-x64-pnpm-&per_page=100' \
|
|
--jq '.actions_caches[] | {id, ref, key, size_in_bytes}'
|
|
```
|
|
|
|
Match the key and upload size against the cache-creation job's logs. The
|
|
September 11 incident was cache ID `7559920987`, a 216-byte archive. This guarded
|
|
command deletes only that observed entry. It leaves a populated replacement or
|
|
an entry on another branch untouched, and does nothing if the old ID is absent:
|
|
|
|
```sh
|
|
bad_cache_id=7559920987
|
|
bad_cache_key=node-cache-Linux-x64-pnpm-c3096ecb02a34aaa9782baaadafcb731510e1dba10dd661618c3a2ee91e58fa5
|
|
entries="$(gh api --paginate --slurp 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&per_page=100')"
|
|
if printf '%s\n' "$entries" | jq -e --argjson id "$bad_cache_id" --arg key "$bad_cache_key" '
|
|
[.[].actions_caches[] | select(.id == $id)] |
|
|
length == 1 and .[0].ref == "refs/heads/master" and
|
|
.[0].key == $key and .[0].size_in_bytes == 216
|
|
' >/dev/null; then
|
|
gh api --method DELETE "repos/paperclipai/paperclip/actions/caches/$bad_cache_id"
|
|
fi
|
|
```
|
|
|
|
A subsequent master install can populate the missing entry. Check the saved
|
|
archive size and package reuse in install logs; a cache hit alone does not prove
|
|
that the entry contains dependencies.
|