Files
PaperClipAI/doc/cloud-build-readiness.md
T
Devin FoleyandPaperclip 19c76bfc3f fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - CI installs dependencies before it verifies and builds cloud
artifacts.
> - Install jobs share a pnpm package-store cache with lockfile refresh.
> - Lockfile refresh resolves versions without downloading packages.
> - That job saved an empty cache before full install jobs could save
theirs.
> - This PR prevents lockfile refresh from publishing that empty entry.
> - Full install jobs can then populate the cache and reuse
dependencies.

## Linked Issues or Issue Description

Refs #13259 for the related cloud verification cache work. No duplicate
empty-cache fix was found.

**What happened?**

Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm
cache at 18:58:08 UTC on September 10. Full install jobs still restore
that empty entry. The cache API reports 216 bytes for master and about
703 MB for populated entries with the same key and cache version in PR
scopes.

**Expected behavior**

A job that installs dependencies should populate the shared
package-store cache.

**Steps to reproduce**

1. Run lockfile refresh with a new lockfile cache key.
2. Its resolution-only command leaves the package store empty.
3. The Node action saves the empty archive before a full install
finishes.
4. Later jobs report a cache hit but download packages again.

**Paperclip version or commit**

Observed on master 6728e133f8 and still
present at a23ae894a5.

**Deployment mode**

GitHub Actions cloud verification and release workflows.

## What Changed

- Disable package-manager caching in Refresh Lockfile.
- Document how to remove the existing empty default-branch entry and
verify a populated replacement.
- Add regression coverage for explicit and automatic package-manager
cache selection in a resolution-only job.

## Verification

- actionlint and git diff checks pass.
- All 175 existing workflow-script tests pass. Both new regression cases
pass and fail against the original workflow, covering the explicit pnpm
cache and automatic npm cache paths. This change adds no application
behavior.
- [The cache creator
job](https://github.com/paperclipai/paperclip/actions/runs/34517514932/job/103006542158)
logs a 216-byte upload under the same key still used by cloud
verification.
- The batch-wide local full typecheck and build passed. The local full
test run reported 10,600 passed, 65 skipped, and 13 permission failures
in unchanged runtime-skill suites. These checks were not repeated in
this dependency-free worktree. Current-head Linux CI passes. The
unchanged chat and browser suites passed on their single retry; all
final checks are green. Greptile is 5/5 with all threads resolved.
- After merge, delete only the existing empty master cache entry. Verify
that a master install saves a populated archive and subsequent jobs
reuse packages. Measure the net install-time change before claiming a
latency gain.

## Risks

- Lockfile resolution can require fresh registry metadata. It does not
need a cached package store.
- The existing empty cache must be removed once; this change prevents
its recreation by this workflow.
- Cache benefits vary with download speed and archive extraction time.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — focused workflow checks
pass; the batch-wide local test limitation is disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 14:35:28 -07:00

13 KiB

Cloud build readiness

The Cloud readiness workflow starts for every master push. Its versioned Cloud deployable v1 job succeeds only after all three prerequisites succeed:

  • The existing Release Verify workflow checks that exact commit, including typecheck, builds, general and serialized tests, and Runner verification.
  • The reusable Docker cloud workflow builds and verifies its Linux AMD64 image, including Sentry resolution and orphan reaping, then publishes the full-SHA cloud tag. Cloud readiness owns the master trigger so there is one cloud build per push. Release tags and manual Docker runs retain their callers.
  • The full-SHA image and both exact-source npm packages are visible. The packages are @paperclipai/shared and @paperclipai/db at 0.0.0-preview.g<FULL_SHA>, published through the migrator-only release lane. Registry metadata must match the full commit, and the database package must pin the matching shared package.

The Cloud workflow builds the image with USER_UID=1001 and USER_GID=1001, matching the managed runtime. This avoids a startup user remap, which can walk the mounted home and delay health checks. Before publishing the full-SHA tag, the workflow checks the baked identity without running the entrypoint, then checks the normal entrypoint's effective user and writable home. Volume ownership repair still runs when needed. The Dockerfile defaults remain 1000:1000 for self-hosted builds, and runtime identity overrides remain supported. The first build with the new identity must rebuild layers that depend on the base image; later builds can reuse those layers.

Verification and image building run concurrently, outside the full npm release's concurrency group. Different commits have independent groups. The npm canary release reuses Cloud source verified v1 for the exact master push instead of starting a second copy of Release Verify. This source-only job depends on every source check but does not wait for Docker or migrator publication. npm canary publication remains possible when source verification passes and an image build fails. Stable releases and candidate-branch betas still run full verification.

The canary consumer requires the expected workflow ID and path, upstream source repository, master push event, full SHA, and a successful job in the latest run attempt. It checks the run again after reading the jobs to reject a concurrent rerun. Missing proof waits for up to 45 minutes; failed, skipped, cancelled, ambiguous, or mismatched proof cannot authorize publication. API failures fail closed. If a source check fails, fix it and rerun Cloud readiness before retrying the release. Use Re-run all jobs when a later attempt did not rerun the source proof; an earlier attempt's successful job is not accepted. This avoids duplicate test jobs on standard runners. Measure queue time to assess the timing gain.

Release verification spreads the general server suites across ten standard hosted runners, with the long chat suite split separately across three jobs. Each server job still runs one test worker. The partition covers every suite exactly once; normal PR and local test groups keep their existing shape. More jobs increase concurrent runner demand, so compare queue time as well as test duration.

The artifact wait runs for up to 30 minutes and reports what is missing. Only an HTTP 404 means publication is pending; authorization errors, upstream outages, and identity mismatches fail the job. A failed, cancelled, or skipped prerequisite cannot produce a successful readiness job. Retry the failed publication or build, then rerun the failed readiness workflow jobs to check the same commit again.

Consumer contract

Cloud deployable v1 is a source-and-artifact readiness signal. A deployment consumer must still resolve and pin the image digest and npm integrity/lockfile, validate migration contents and compatibility, and apply its target health gates. The check creates no release record and deploys no instance. A full-SHA tag by itself, or a successful migrator dispatch, is not this readiness signal.

For automatic selection, accept only a successful job named exactly Cloud deployable v1 in the latest attempt of a successful .github/workflows/cloud-readiness.yml run in paperclipai/paperclip, with event push, head branch master, and the expected full head SHA and repository. Do not trust a similarly named check from another workflow or a manual branch run. Order candidates by master ancestry, not job completion time: an older commit finishing late must not roll a fleet backward. Fail closed on API errors.

Existing npm canary discovery is unchanged by this producer workflow. Consumers can adopt the versioned signal separately after the workflow has landed and successfully verified a real master commit.

Timing and rollout

The reusable Runner chaos workflow scopes concurrency to the caller workflow and source ref. Cloud readiness, stable verification, and standalone evals can verify the same commit at the same time. They must not cancel each other's required test job.

Measure the complete path from a master merge to a healthy target running that exact commit. Keep readiness and deployment as separate milestones:

Milestone Evidence Elapsed time starts at
Merge Merged PR timestamp and full merge commit SHA Merge
Image available Successful full-SHA image publication and verification Merge
Cloud deployable Successful Cloud deployable v1 job in the accepted push run and attempt Merge
Canary healthy Deployment consumer's canary health gate confirms the target commit Merge
Fleet complete Campaign succeeds for all eligible targets at that commit Merge

Record the source SHA, workflow run ID and attempt, readiness job completion time, and deployment campaign identity together. Verify the run against the consumer contract above. A manual dispatch can test wiring, but its timestamp does not measure automatic merge-to-deploy latency. A preparation-only run resolves artifacts without deploying a target and must not be counted as a successful deployment.

Record queue time and the image, source-verification, and artifact-wait durations separately. The slowest prerequisite determines readiness; shortening an already faster prerequisite may have no effect on the total. After readiness, measure consumer discovery delay, artifact resolution, canary health, and fleet rollout. An automatic consumer that still waits for the full npm canary publication has that queue on its critical path even if cloud artifacts are ready earlier.

For a target health measurement, confirm the deployed source SHA as well as service health. A proxy health response alone may describe the control plane while the tenant still runs the previous image. Report the eligible target count, excluded or sleeping targets, retries, and failures with the fleet result. Record runner queue conditions and cache state; one warm or cold run is a sample, not a latency guarantee.

Land full-SHA image publication, independent cloud builds, and migrator-only publication before enabling this workflow. Until those producers are present, the artifact wait cannot succeed. A manual dispatch on master can verify the wiring, but automatic consumers should use push runs. Source verification and registry checks can be rerun without deploying or changing mutable npm channels.

When reverting this workflow, restore the master push trigger in docker-cloud.yml in the same change so master images continue to build.

Reserved AWS verification capacity

AWS_POST_MERGE_CI_ENABLED=true routes cloud source verification, artifact waiting, readiness signals, and exact-master migrator preparation to the paperclip-post-merge runner group. The separate Fleet label is runs-on/fleet=paperclip-post-merge-x64/env=public-ci. Its 36 reserved slots use the same four-vCPU, 16-GiB machines as approved PR jobs. PR capacity is reduced to 64; image capacity stays at eight. The total ceiling remains 108 runners. This keeps PR bursts from consuming every post-merge verification slot.

Every selector checks the canonical repository name and ID, master ref, and a push or manual event. Reusable verification also requires inputs.ref to equal that event's github.sha. The migrator route requires cloud-migrator and inputs.source_ref == github.sha. Branch/tag refs, PR events, arbitrary preview sources, and missing or disabled switches use GitHub-hosted runners. If another merge lands before a migrator dispatch resolves master, the older source uses GitHub-hosted runners too. npm publication always remains GitHub-hosted to keep its trusted-publisher identity.

Before enabling the switch, deploy the separate Fleet and restrict its GitHub runner group to repository ID 1170821064 and these workflows at refs/heads/master: cloud-readiness.yml, cloud-artifacts.yml, release-verify.yml, runner-chaos-evals.yml, and release.yml. Do not authorize PR-controlled workflow versions. PR placement retains its independent pinned workflow and six-account author/actor allowlist.

Disable the switch and rerun the whole workflow to restore GitHub-hosted placement. Assigned jobs keep their original runners. Readiness requirements, source checks, and npm integrity checks are unchanged.

AWS cloud build routing

AWS_CLOUD_BUILDS_ENABLED=true routes the Docker cloud job to the paperclip-cloud-build-x64 RunsOn Fleet for canonical paperclipai/paperclip master pushes and manual master runs. Forks, pull requests, and release tags retain GitHub-hosted runners. The separate AWS_CI_ENABLED and AWS_CI_TRUSTED_USER_IDS variables control PR routing.

The cloud Fleet uses a separate runner group, paperclip-cloud-build, restricted to this repository and .github/workflows/docker-cloud.yml@refs/heads/master. Provision that group and Fleet before enabling the variable. The cloud runners need at least 64 GiB free for Docker and the workspace; the initial configuration uses 120 GiB disks with the existing 4-vCPU, 16-GiB machine size. AWS jobs have a 40-minute workflow timeout so they finish before the 45-minute instance lifetime; GitHub-hosted jobs retain their 60-minute timeout. Keep the registry cache and all pushed-image verification steps enabled.

To roll back routing, set AWS_CLOUD_BUILDS_ENABLED=false, then rerun the cloud workflow. Changing the variable does not migrate an already assigned job. Check the Actions job's runner name and runner group to verify placement. Record queue time, image verification completion, and Cloud deployable v1 separately; source verification and the migrator still run on GitHub-hosted runners.

Typecheck Rust dependency cache

Source verification's typecheck job builds the native Runner binary through the server's prepare:runner-vendor command. It restores and saves compiled Rust dependencies only for canonical master pushes that verify the event's exact SHA. The release-typecheck-v1 cache is separate from Runner verification because those jobs compile different profiles. The pinned toolchain is selected before cache lookup. Workspace crates and installed cargo binaries are excluded, and all typechecks still execute. A missing or invalidated cache triggers compilation.

pnpm dependency store cache

The Refresh Lockfile workflow does not cache the pnpm store. Its resolution-only command does not download packages and can save an empty default-branch cache before full install jobs finish. Jobs that install dependencies retain caching.

After deploying this correction, remove any existing empty default-branch entry for the current lockfile key. List cache IDs, branches, and archive sizes first:

gh api --paginate 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&key=node-cache-Linux-x64-pnpm-&per_page=100' \
  --jq '.actions_caches[] | {id, ref, key, size_in_bytes}'

Match the key and upload size against the cache-creation job's logs. The September 11 incident was cache ID 7559920987, a 216-byte archive. This guarded command deletes only that observed entry. It leaves a populated replacement or an entry on another branch untouched, and does nothing if the old ID is absent:

bad_cache_id=7559920987
bad_cache_key=node-cache-Linux-x64-pnpm-c3096ecb02a34aaa9782baaadafcb731510e1dba10dd661618c3a2ee91e58fa5
entries="$(gh api --paginate --slurp 'repos/paperclipai/paperclip/actions/caches?ref=refs/heads/master&per_page=100')"
if printf '%s\n' "$entries" | jq -e --argjson id "$bad_cache_id" --arg key "$bad_cache_key" '
  [.[].actions_caches[] | select(.id == $id)] |
  length == 1 and .[0].ref == "refs/heads/master" and
  .[0].key == $key and .[0].size_in_bytes == 216
' >/dev/null; then
  gh api --method DELETE "repos/paperclipai/paperclip/actions/caches/$bad_cache_id"
fi

A subsequent master install can populate the missing entry. Check the saved archive size and package reuse in install logs; a cache hit alone does not prove that the entry contains dependencies.