Files
PaperClipAI/.github/workflows/docker.yml
T
5488a79eb5 ci(docker): build each architecture on a native runner instead of QEMU (#12821)
## Thinking Path

A fleet rollout failed with `image_manifest_not_found` for a commit that
had merged and gone green. Tracing that back: the fleet resolves
releases against the `-cloud` image, that image comes from `docker.yml`,
and `docker.yml`'s runs on master have been reading `cancelled` for a
long stretch. The cloud half was fine; the production half was hanging
and taking the run down with it — and, because a run holds the
concurrency slot for its whole duration, starving later commits of a
build at all.

## Linked Issues or Issue Description

No tracking issue — described inline, per CONTRIBUTING.md.

**What's wrong.** `build-and-push` builds `linux/amd64,linux/arm64` on
an x86 runner, so arm64 runs under QEMU. It hangs there —
deterministically, in the same step:

```
#111 [linux/arm64 build  8/10] RUN pnpm --filter @paperclipai/server build
```

…then emits nothing until `timeout-minutes: 60` kills it. Three
consecutive runs on 2026-09-04, silent for **38, 43 and 45 minutes**
respectively. The amd64 leg reached `production 5/5` minutes earlier in
every one.

**Why it stayed hidden.** A timed-out job is reported by GitHub as
**cancelled, not failed**. The run conclusion reads "cancelled", which
looks like supersession rather than breakage, so the production image
quietly stopped publishing.

**The knock-on.** A run that burns the full hour holds the top-level
concurrency slot for that hour. `cancel-in-progress: false` keeps
exactly one pending slot, so merges arriving faster than one an hour
supersede each other while queued. Sampling the last ten master commits,
**five produced no image at all** — their Docker runs have zero job
records because they never started.

**Expected.** Both architectures publish, and a commit merged during a
busy period still gets an image.

## What Changed

`build-and-push` becomes a two-leg matrix, each on a runner of its own
architecture:

| platform | runner |
|---|---|
| `linux/amd64` | `ubuntu-latest` |
| `linux/arm64` | `ubuntu-24.04-arm` |

Each leg pushes **by digest** (`push-by-digest=true`, untagged), and a
new `merge-and-push` job names the digests into one manifest list with
the real lane tags. Nothing is publicly tagged until the merge, so a
half-published multi-arch image is never a pullable state.

Two supporting changes:

- **Per-arch BuildKit cache refs** (`:buildcache-amd64` /
`:buildcache-arm64`). Separate runners sharing one ref would overwrite
each other on every build.
- **The PID-1 orphan-reaping check moves to the merge job**, since that
is where a tagged, pullable image first exists. It still runs against
the pushed image rather than a local build, for the same reason as
before.

**arm64 is kept, not dropped.** The cloud variant is amd64-only and can
be — managed hosts are amd64. This is the self-hosted image and ARM
hosts consume it, so dropping arm64 would break them. GitHub-hosted
arm64 runners are free for public repositories, which this is.

`build-and-push-cloud` is untouched. It was already `platforms:
linux/amd64` and has been succeeding in ~14 minutes throughout — that is
why `-cloud` images exist at all.

## Verification

Parsed the workflow and asserted its shape (jobs, matrix, `needs`, step
order, that the cloud job is unchanged). The artifact actions are pinned
by SHA with version comments, matching the repo's dominant convention —
`upload-artifact` v7 and `download-artifact` v8, the same pins used
across the other workflows; v8 is required for the `pattern` /
`merge-multiple` inputs the merge job uses.

**This PR's CI does not exercise the change.** `docker.yml` triggers on
master and tag pushes, never on pull requests — deliberately, since it
publishes release images. The first real run is after merge, so the
check is: the next master push produces a `Docker` run whose
`build-and-push (amd64)`, `build-and-push (arm64)` and `merge-and-push`
jobs all succeed, and whose conclusion is `success` rather than
`cancelled`.

## Risks

- **Not testable before merge**, per above. If the matrix is wrong the
next master push fails loudly rather than silently — which is already
better than the current state, where the failure mode is an invisible
"cancelled".
- **First use of `ubuntu-24.04-arm` in this repo.** No other workflow
uses an ARM runner. They are free for public repos, but if the label is
unavailable the arm64 leg will fail to schedule and the merge will not
run — no image, same as today, and visible.
- **Digest-push changes the publish shape.** Between the legs finishing
and the merge running, digests exist untagged in ghcr. Anything watching
for tags sees no intermediate state; anything enumerating untagged
manifests will see more of them.
- **Cache refs change name**, so the first build after this lands is
cold on both legs and will be slower than steady state.
- **Does not fix the underlying QEMU hang** — it avoids it. If arm64
ever has to build under emulation again, the same stall is presumably
still there.
- No application code, schema, server or persistence change.

## Model Used

Anthropic Claude — Opus 5, model ID `claude-opus-5`, run through Claude
Code.

Extended thinking enabled. Tool use throughout: GitHub Actions API to
correlate run/job outcomes and read build logs, `git` for ancestry
checks, and a YAML parser to validate the rewritten workflow's
structure.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

A workflow change has no unit test to add, and `docker.yml` cannot run
on a PR; the verification section states what to check on the first
master run instead.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Devin Foley <devin@paperclip.ing>
2026-09-09 15:56:28 -07:00

638 lines
28 KiB
YAML

name: Docker
on:
push:
branches:
- "master"
tags:
- "v*"
- "nightly/v*"
- "beta/v*"
# Release workflows push lane tags with GITHUB_TOKEN, and GitHub suppresses
# push-triggered runs for those, so release.yml dispatches this workflow at
# the new tag ref instead. The tag mapping below keys off github.ref either
# way.
workflow_dispatch:
# Least privilege: nothing at the workflow level; each job declares exactly
# the token scopes it uses (checkout needs contents:read, GHCR pushes need
# packages:write).
permissions: {}
# Serialise builds per ref without killing an in-flight one: a newer push
# supersedes only the pending slot, so the image build that is already
# running always finishes and publishes. Canary TAG refs each get their
# own group on purpose: their builds run in parallel so every published
# canary gets its sha images regardless of merge cadence. The mutable
# `:canary` channel tags are NOT written by the build matrix (which
# would race across parallel runs) — each canary-tag run retags the
# channel afterwards, only if it still matches the npm `canary`
# dist-tag, so the channel moves monotonically and always mirrors npm.
concurrency:
group: docker-${{ github.ref }}
cancel-in-progress: false
jobs:
# Multi-arch by native runner, not QEMU.
#
# This was one job building linux/amd64,linux/arm64 together on an x86
# runner. The arm64 half is emulated there, and it did not merely run slow:
# it wedged, every time, in `RUN pnpm --filter @paperclipai/server build`,
# emitting nothing for 38-45 minutes until `timeout-minutes: 60` killed the
# job. Verified across three consecutive runs on 2026-09-04; the amd64 half
# reached `production 5/5` minutes earlier in every one.
#
# A timed-out job is reported as *cancelled*, not failed, so the run read
# "cancelled" and the production image simply stopped publishing without
# anything going red in an obvious way.
#
# It also starved the queue. A run that burns the full hour holds the
# top-level concurrency slot for that hour, and `cancel-in-progress: false`
# keeps exactly one pending slot — so with merges arriving faster than one
# an hour, most runs were superseded before they ever started a job. Five of
# ten master commits sampled that day never produced an image at all.
#
# Each platform now builds on a runner of its own architecture and pushes by
# digest; `merge` assembles the manifest list. arm64 is kept rather than
# dropped (the cloud variant below dropped it and is amd64-only) because
# this is the self-hosted image, and ARM hosts consume it.
build-and-push:
strategy:
# Independent legs: one architecture failing should still publish
# nothing, but it must not also hide the other's logs behind a
# cancellation.
fail-fast: false
matrix:
include:
- platform: linux/amd64
runner: ubuntu-latest
arch: amd64
- platform: linux/arm64
runner: ubuntu-24.04-arm
arch: arm64
runs-on: ${{ matrix.runner }}
# Native builds land well inside this; it is a backstop, not a budget.
# (The interim fix while this PR landed raised the single QEMU job's cap
# to 120 minutes; native per-arch legs make that headroom unnecessary.)
timeout-minutes: 60
permissions:
contents: read
packages: write
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history and tags so `git describe` below can compute the
# release version to stamp into the image.
fetch-depth: 0
# `.git` is dockerignored, so a running image cannot derive its own
# version and otherwise reports the source package.json placeholder in
# analytics and the debug panel. Compute it here from the pristine
# checkout (real CalVer drift from the nearest release tag) and pass it
# into both builds. Empty when no release tag is reachable — the server
# then keeps its existing fallbacks.
- name: Compute build version
id: build-version
run: |
set -euo pipefail
case "${GITHUB_REF}" in
refs/tags/nightly/v*)
# Lane tags carry the exact published version; stamp it verbatim
# instead of describing drift from the nearest stable tag.
version="${GITHUB_REF#refs/tags/nightly/v}"
;;
refs/tags/beta/v*)
version="${GITHUB_REF#refs/tags/beta/v}"
;;
*)
version="$(git describe --tags --match 'v*' --long --dirty 2>/dev/null || true)"
;;
esac
echo "version=${version}" >> "$GITHUB_OUTPUT"
echo "Stamping build version: ${version:-<none>}"
# ISO week stamp for the Dockerfile's tool layer: the layer caches
# across commits and re-pulls the @latest CLI tools when the week rolls
# over, instead of on every build.
- name: Compute tool cache epoch
id: tools-epoch
run: echo "epoch=$(date -u +%G-W%V)" >> "$GITHUB_OUTPUT"
- name: Setup pnpm
uses: pnpm/action-setup@v6
with:
version: 9.15.4
run_install: false
# No dependency cache here: this workflow publishes release images, and
# restoring a shared Actions cache into the build inputs would let a
# poisoned cache entry reach the published artifact.
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: 24
- name: Refresh lockfile for Docker build context
run: |
set -euo pipefail
pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile
changed="$(git status --porcelain)"
if [ -z "$changed" ]; then
echo "Lockfile already matches package metadata."
exit 0
fi
if printf '%s\n' "$changed" | grep -Fvq ' pnpm-lock.yaml'; then
echo "Unexpected files changed during lockfile refresh:"
echo "$changed"
exit 1
fi
echo "Using refreshed pnpm-lock.yaml in the Docker build context."
- name: Free runner disk
run: |
set -euo pipefail
echo "Disk before cleanup:"
df -h
pnpm store prune || true
sudo apt-get clean || true
sudo rm -rf \
/usr/share/dotnet \
/usr/share/swift \
/usr/local/lib/android \
/usr/local/share/boost \
/usr/local/share/powershell \
/opt/ghc \
/opt/hostedtoolcache/CodeQL \
/opt/hostedtoolcache/PyPy \
/opt/hostedtoolcache/Ruby || true
docker system prune -af || true
echo "Disk after cleanup:"
df -h
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
# Deployment tooling reads these labels from the registry to verify an
# image's schema expectations against a migrator before deploying it,
# without pulling the image. The server refuses to start when the
# database is missing bundled migrations, so orchestrators need a cheap
# way to check image/migrator compatibility up front.
- name: Compute schema migration labels
id: schema
run: |
set -euo pipefail
last=$(ls packages/db/src/migrations/*.sql | sed 's|.*/||' | LC_ALL=C sort | tail -1)
count=$(ls packages/db/src/migrations/*.sql | wc -l | tr -d ' ')
echo "last=${last}" >> "$GITHUB_OUTPUT"
echo "count=${count}" >> "$GITHUB_OUTPUT"
# Lane tag mapping: nightly/v* tags publish `:nightly`, and only
# stable v* tags move `:latest` and the versioned tags.
# `:sha-<short>` is published on every build. `:canary` is
# deliberately absent here — the channel tag is moved by the
# dist-tag-checked retag step below, never by the build matrix,
# so parallel canary builds cannot race it backwards.
- name: Docker meta
id: meta
uses: docker/metadata-action@v6
with:
images: ghcr.io/${{ github.repository }}
tags: |
type=raw,value=nightly,enable=${{ startsWith(github.ref, 'refs/tags/nightly/v') }}
type=raw,value=beta,enable=${{ startsWith(github.ref, 'refs/tags/beta/v') }}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{version}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{major}}.{{minor}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=sha
labels: |
io.github.paperclipai.schema.last-migration=${{ steps.schema.outputs.last }}
io.github.paperclipai.schema.migration-count=${{ steps.schema.outputs.count }}
- name: Build and push by digest
id: build
uses: docker/build-push-action@v7
with:
context: .
# Pin the self-hosted image to the production stage explicitly:
# the Dockerfile now declares a later `cloud` stage, and without a
# target the default would silently become that stage.
target: production
build-args: |
PAPERCLIP_BUILD_VERSION=${{ steps.build-version.outputs.version }}
PAPERCLIP_BUILD_COMMIT=${{ github.sha }}
CLI_TOOLS_CACHE_EPOCH=${{ steps.tools-epoch.outputs.epoch }}
platforms: ${{ matrix.platform }}
# By digest, not by tag: two runners cannot each push the same tag
# and end up with a manifest list. Each leg publishes an untagged
# image and `merge` names them together.
outputs: type=image,name=ghcr.io/${{ github.repository }},push-by-digest=true,name-canonical=true,push=true
# Registry-backed BuildKit cache instead of type=gha: the Actions
# cache is capped at 10GB per repo, and two multi-arch mode=max jobs
# evict each other, so most builds ran effectively cold. The cache
# ref lives in ghcr next to the image and is written only by this
# workflow (docker.yml runs on master/tag pushes, never on PRs).
#
# Per-arch refs now the legs are separate runners: a shared ref would
# have each leg overwrite the other's cache on every build.
cache-from: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-${{ matrix.arch }}
cache-to: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-${{ matrix.arch }},mode=max
labels: ${{ steps.meta.outputs.labels }}
# The digest is the only thing `merge` needs from this job. Carried as an
# empty file named for it, which is the upstream pattern — the name is
# the payload, so several legs can upload without colliding on content.
- name: Export digest
run: |
set -euo pipefail
mkdir -p /tmp/digests
digest="${{ steps.build.outputs.digest }}"
test -n "$digest"
touch "/tmp/digests/${digest#sha256:}"
- name: Upload digest
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: digests-production-${{ matrix.arch }}
path: /tmp/digests/*
if-no-files-found: error
retention-days: 1
# Names the per-architecture digests as one manifest list under the real
# tags. Nothing is publicly tagged until this runs, so a half-published
# multi-arch image is not a state anything can pull.
merge-and-push:
needs: build-and-push
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
contents: read
packages: write
steps:
# Checked out for `packages/db` (schema labels) and the orphan-reaping
# script the verification step pipes in.
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
fetch-depth: 0
- name: Download digests
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
path: /tmp/digests
pattern: digests-production-*
merge-multiple: true
- name: Login to GitHub Container Registry
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@37fe631027851001ddb9b187196cc803df7f5f0e # v4
# Repeated from the build job rather than passed between them: job
# outputs would have to survive a matrix, and this is two `ls` calls.
- name: Compute schema migration labels
id: schema
run: |
set -euo pipefail
last=$(ls packages/db/src/migrations/*.sql | sed 's|.*/||' | LC_ALL=C sort | tail -1)
count=$(ls packages/db/src/migrations/*.sql | wc -l | tr -d ' ')
echo "last=${last}" >> "$GITHUB_OUTPUT"
echo "count=${count}" >> "$GITHUB_OUTPUT"
# Same lane mapping as the build job; this is where it is actually
# applied, since the legs push untagged. `:canary` is deliberately
# absent, exactly as in the build job's mapping: the channel tag is
# moved only by the dist-tag-checked promote_canary_channel job below,
# so parallel canary-tag builds can never race the channel backwards.
- name: Docker meta
id: meta
uses: docker/metadata-action@dc802804100637a589fabce1cb79ff13a1411302 # v6
with:
images: ghcr.io/${{ github.repository }}
tags: |
type=raw,value=nightly,enable=${{ startsWith(github.ref, 'refs/tags/nightly/v') }}
type=raw,value=beta,enable=${{ startsWith(github.ref, 'refs/tags/beta/v') }}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{version}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{major}}.{{minor}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=sha
labels: |
io.github.paperclipai.schema.last-migration=${{ steps.schema.outputs.last }}
io.github.paperclipai.schema.migration-count=${{ steps.schema.outputs.count }}
- name: Create manifest list and push
working-directory: /tmp/digests
run: |
set -euo pipefail
docker buildx imagetools create \
$(jq -cr '.tags | map("-t " + .) | join(" ")' <<< "$DOCKER_METADATA_OUTPUT_JSON") \
$(printf 'ghcr.io/${{ github.repository }}@sha256:%s ' *)
# PID 1 must be an init that reaps adopted orphans. With node there, the
# orphans agent runs leave behind are never wait()ed and pin as zombies
# until the cgroup pid limit is exhausted and every fork() in the
# container fails. Run against the pushed manifest rather than a local
# build: the legs push by digest, so nothing is loaded into this
# runner's daemon. The cloud variant is FROM production and inherits the
# same ENTRYPOINT, so checking this image covers both.
- name: Verify PID 1 reaps orphaned processes
env:
# Through the environment, not interpolated into the script body, so
# the tag text is data rather than shell.
IMAGE_TAGS: ${{ steps.meta.outputs.tags }}
run: |
set -euo pipefail
image="$(printf '%s\n' "$IMAGE_TAGS" | head -n 1)"
test -n "$image"
echo "Verifying orphan reaping in $image"
docker run --rm -i --pull always "$image" sh -s < scripts/assert-orphan-reaping.sh
# The cloud variant carries built bundled plugins for managed deployments
# (see the `cloud` stage in the Dockerfile). It runs as its own job with no
# `needs:` on the stock publish above, so the two builds run in parallel and
# a failure or slow build in one never gates, delays, or skips the other.
# Both jobs share only the single top-level concurrency slot. Each job is a
# separate runner, so this one carries its own copy of the prep steps
# (checkout through schema labels) — the accepted cost of that isolation.
build-and-push-cloud:
runs-on: ubuntu-latest
timeout-minutes: 60
permissions:
contents: read
packages: write
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history and tags so `git describe` below can compute the
# release version to stamp into the image.
fetch-depth: 0
# `.git` is dockerignored, so a running image cannot derive its own
# version and otherwise reports the source package.json placeholder in
# analytics and the debug panel. Compute it here from the pristine
# checkout (real CalVer drift from the nearest release tag) and pass it
# into the build. Empty when no release tag is reachable — the server
# then keeps its existing fallbacks.
- name: Compute build version
id: build-version
run: |
set -euo pipefail
case "${GITHUB_REF}" in
refs/tags/nightly/v*)
# Lane tags carry the exact published version; stamp it verbatim
# instead of describing drift from the nearest stable tag.
version="${GITHUB_REF#refs/tags/nightly/v}"
;;
refs/tags/beta/v*)
version="${GITHUB_REF#refs/tags/beta/v}"
;;
*)
version="$(git describe --tags --match 'v*' --long --dirty 2>/dev/null || true)"
;;
esac
echo "version=${version}" >> "$GITHUB_OUTPUT"
echo "Stamping build version: ${version:-<none>}"
# ISO week stamp for the Dockerfile's tool layer: the layer caches
# across commits and re-pulls the @latest CLI tools when the week rolls
# over, instead of on every build.
- name: Compute tool cache epoch
id: tools-epoch
run: echo "epoch=$(date -u +%G-W%V)" >> "$GITHUB_OUTPUT"
- name: Setup pnpm
uses: pnpm/action-setup@v6
with:
version: 9.15.4
run_install: false
# No dependency cache here: this workflow publishes release images, and
# restoring a shared Actions cache into the build inputs would let a
# poisoned cache entry reach the published artifact.
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: 24
- name: Refresh lockfile for Docker build context
run: |
set -euo pipefail
pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile
changed="$(git status --porcelain)"
if [ -z "$changed" ]; then
echo "Lockfile already matches package metadata."
exit 0
fi
if printf '%s\n' "$changed" | grep -Fvq ' pnpm-lock.yaml'; then
echo "Unexpected files changed during lockfile refresh:"
echo "$changed"
exit 1
fi
echo "Using refreshed pnpm-lock.yaml in the Docker build context."
- name: Free runner disk
run: |
set -euo pipefail
echo "Disk before cleanup:"
df -h
pnpm store prune || true
sudo apt-get clean || true
sudo rm -rf \
/usr/share/dotnet \
/usr/share/swift \
/usr/local/lib/android \
/usr/local/share/boost \
/usr/local/share/powershell \
/opt/ghc \
/opt/hostedtoolcache/CodeQL \
/opt/hostedtoolcache/PyPy \
/opt/hostedtoolcache/Ruby || true
docker system prune -af || true
echo "Disk after cleanup:"
df -h
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
# Deployment tooling reads these labels from the registry to verify an
# image's schema expectations against a migrator before deploying it,
# without pulling the image. The server refuses to start when the
# database is missing bundled migrations, so orchestrators need a cheap
# way to check image/migrator compatibility up front.
- name: Compute schema migration labels
id: schema
run: |
set -euo pipefail
last=$(ls packages/db/src/migrations/*.sql | sed 's|.*/||' | LC_ALL=C sort | tail -1)
count=$(ls packages/db/src/migrations/*.sql | wc -l | tr -d ' ')
echo "last=${last}" >> "$GITHUB_OUTPUT"
echo "count=${count}" >> "$GITHUB_OUTPUT"
# Published under the same lane tag set as the self-hosted image, with a
# `-cloud` suffix (nightly-cloud, latest-cloud, <version>-cloud,
# sha-<short>-cloud). `:canary-cloud` follows the same retag-step
# ownership rule as `:canary` above.
- name: Docker meta (cloud)
id: meta-cloud
uses: docker/metadata-action@v6
with:
images: ghcr.io/${{ github.repository }}
flavor: |
suffix=-cloud,onlatest=true
tags: |
type=raw,value=nightly,enable=${{ startsWith(github.ref, 'refs/tags/nightly/v') }}
type=raw,value=beta,enable=${{ startsWith(github.ref, 'refs/tags/beta/v') }}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{version}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{major}}.{{minor}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=sha
labels: |
io.github.paperclipai.schema.last-migration=${{ steps.schema.outputs.last }}
io.github.paperclipai.schema.migration-count=${{ steps.schema.outputs.count }}
- name: Build and push (cloud)
uses: docker/build-push-action@v7
with:
context: .
target: cloud
# Space-separated sandbox-provider directory names to build into
# the variant; add here when managed deployments need another.
# CLOUD_BUNDLED_SERVER_DEPS names the optional peer packages the
# variant installs from server/package.json's declared version;
# add another name there when a managed tenant needs it.
build-args: |
CLOUD_BUNDLED_PLUGINS=daytona
CLOUD_BUNDLED_SERVER_DEPS=@sentry/node
PAPERCLIP_BUILD_VERSION=${{ steps.build-version.outputs.version }}
PAPERCLIP_BUILD_COMMIT=${{ github.sha }}
CLI_TOOLS_CACHE_EPOCH=${{ steps.tools-epoch.outputs.epoch }}
# amd64 only, unlike the self-hosted image above: the cloud variant
# is consumed exclusively by managed-deployment hosts, which run
# amd64. The QEMU-emulated arm64 half dominated this job's wall
# clock, and dropping it roughly halves time-to-deployable-image.
platforms: linux/amd64
push: true
# Registry-backed BuildKit cache, separate ref from the self-hosted
# job so the two parallel builds never clobber each other's cache
# manifest (see the rationale on the job above).
cache-from: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-cloud
cache-to: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-cloud,mode=max
tags: ${{ steps.meta-cloud.outputs.tags }}
labels: ${{ steps.meta-cloud.outputs.labels }}
# The cloud target installs @sentry/node at the version
# server/package.json declares, into a directory the server's own
# module resolution walks. Verify the image this job just pushed, not
# a local build, so a build-cache or layer-ordering regression is
# caught before any tenant runs the image.
- name: Verify the pushed image resolves the declared Sentry version
env:
IMAGE_TAGS: ${{ steps.meta-cloud.outputs.tags }}
run: |
set -euo pipefail
image="$(printf '%s\n' "$IMAGE_TAGS" | head -n 1)"
test -n "$image"
expected="$(node -e "process.stdout.write(require('./server/package.json').peerDependencies['@sentry/node'])")"
test -n "$expected"
installed="$(docker run --rm --pull always \
-v "$PWD/scripts/assert-cloud-image-sentry.mjs:/app/server/.ci-sentry-probe.mjs:ro" \
--entrypoint node "$image" /app/server/.ci-sentry-probe.mjs)"
echo "Declared optional peer version: $expected"
echo "Installed in the pushed image: $installed"
if [ "$installed" != "$expected" ]; then
echo "ERROR: the pushed image resolves @sentry/node@$installed, expected @sentry/node@$expected" >&2
exit 1
fi
echo "The pushed image resolves the declared @sentry/node version."
# Moves the mutable `:canary` / `:canary-cloud` channel tags. Kept OUT
# of the build jobs and serialized in its own lane, and — the load-
# bearing property — CONVERGENT rather than self-interested: a
# promotion does not promote "its own" canary, it retags the channel
# to whatever the npm `canary` dist-tag names at execution time,
# provided that version's sha images are published. GitHub's shared
# concurrency lane keeps one running and one pending promotion and
# REPLACES the pending slot with the latest enqueued — an older build
# finishing late can therefore evict the newest canary's pending
# promotion. With convergent promotion that eviction is harmless:
# whichever promotion survives resolves the current dist-tag fresh
# and lands the channel there (the current canary's images always
# exist by the time any later promotion runs, because per-tag build
# groups mean canary builds are never superseded and each run's
# promotion is gated on its own completed pushes). Every interleaving
# converges the Docker channel onto the npm channel.
promote_canary_channel:
if: startsWith(github.ref, 'refs/tags/canary/v')
# merge-and-push, not build-and-push: the per-arch legs push untagged
# digests, and the production `sha-*` tags this promotion retags only
# exist once the manifest merge has named them.
needs: [merge-and-push, build-and-push-cloud]
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
packages: write
concurrency:
group: docker-canary-channel-promotion
cancel-in-progress: false
steps:
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Converge the channel tags onto the current npm canary
env:
IMAGE: ghcr.io/${{ github.repository }}
GH_TOKEN: ${{ github.token }}
run: |
current="$(curl -fsS "https://registry.npmjs.org/-/package/@paperclipai%2Fdb/dist-tags" | jq -er .canary)"
sha="$(gh api "repos/${GITHUB_REPOSITORY}/commits/$(printf 'canary/v%s' "$current" | jq -sRr @uri)" --jq .sha 2>/dev/null || true)"
if [ -z "$sha" ]; then
echo "canary/v${current} does not resolve yet; a later promotion converges the channel"
exit 0
fi
short="$(printf '%s' "$sha" | cut -c1-7)"
if ! docker buildx imagetools inspect "$IMAGE:sha-${short}-cloud" >/dev/null 2>&1; then
echo "images for ${current} (sha-${short}) not published yet; its own promotion converges the channel"
exit 0
fi
docker buildx imagetools create -t "$IMAGE:canary" "$IMAGE:sha-${short}"
docker buildx imagetools create -t "$IMAGE:canary-cloud" "$IMAGE:sha-${short}-cloud"
echo "channel tags moved to canary ${current} (sha-${short})"