Files
PaperClipAI/.github/workflows/docker.yml
T
Devin Foley 01ad858492 ci: raise the multi-arch Docker publish timeout to 120 minutes (#13114)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Docker workflow publishes the server images that all deployments
pull, including the `sha-*` images that downstream consumers deploy
> - The `build-and-push` job builds for linux/amd64 and QEMU-emulated
linux/arm64, and that build now takes more than its 60 minute job
timeout
> - Every run dies at the timeout, and each doomed hour-long run holds
the per-ref concurrency slot, so queued master pushes supersede each
other and no image publishes at all
> - This pull request raises the multi-arch job timeout to 120 minutes
> - The benefit is that image publishing works again, with headroom for
the build to grow

## Linked Issues or Issue Description

Related: #12821 replaces the QEMU-emulated arm64 build with native
runners — that is the durable fix for the build duration itself. This PR
is the immediate unblock so images publish again while #12821 lands.

No existing issue for the outage. Description follows the bug report
template:

**What happened?**

The `docker.yml` `build-and-push` job hits its 60 minute
`timeout-minutes` cap on every run. The last fully successful
`docker.yml` run was September 2. Since then almost every run ends
`cancelled`: the multi-arch build is killed at the timeout, and runs
queued behind it are superseded by newer master pushes before they can
start. The amd64-only `build-and-push-cloud` job often still succeeds
inside those cancelled runs, which masked the breakage.

**Expected behavior**

Every master push and canary tag dispatch publishes its `sha-*`
production and cloud images, and the `promote_canary_channel` job runs.

**Steps to reproduce**

Look at the runs of the Docker workflow on master: `gh run list
--workflow docker.yml --branch master`. Nearly every run since September
5 ends `cancelled` or `failure`. Open a cancelled run: the
`build-and-push` job runs for 61+ minutes and its "Build and push" step
ends `cancelled` at the job timeout. The last runs that succeeded
(September 2) took 39 to 54 minutes for the same job.

## What Changed

- Raise `timeout-minutes` on the `build-and-push` job from 60 to 120,
with a comment that explains why. The amd64-only `build-and-push-cloud`
job keeps its 60 minute cap.

## Verification

- `actionlint .github/workflows/docker.yml` reports no issues in this
change (only pre-existing info-level shellcheck notes in untouched
steps).
- Compared job durations across the last successful runs (39-54 minutes)
and the recent timeout kills (61+ minutes) to confirm the cap is the
failure cause.
- After merge, the next master push should produce a `docker.yml` run
that completes with both build jobs green.

## Risks

Low risk. The change only gives the existing build more time. A
genuinely hung build now occupies a runner for up to 120 minutes instead
of 60. The slow arm64 emulated build itself is worth a separate look
(native arm runners or splitting the platforms), but that is a larger
change than this outage fix.

## Model Used

Claude (Anthropic) — Fable 5 (`claude-fable-5`), extended thinking,
agentic tool use via Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (no code paths changed;
workflow linted with actionlint)
- [x] I have added or updated tests where applicable (not applicable for
a CI timeout value)
- [x] I have updated relevant documentation to reflect my changes (the
workflow comment documents the rationale)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-09 14:34:45 -07:00

498 lines
22 KiB
YAML

name: Docker
on:
push:
branches:
- "master"
tags:
- "v*"
- "nightly/v*"
- "beta/v*"
# Release workflows push lane tags with GITHUB_TOKEN, and GitHub suppresses
# push-triggered runs for those, so release.yml dispatches this workflow at
# the new tag ref instead. The tag mapping below keys off github.ref either
# way.
workflow_dispatch:
permissions:
contents: read
packages: write
# Serialise builds per ref without killing an in-flight one: a newer push
# supersedes only the pending slot, so the image build that is already
# running always finishes and publishes. Canary TAG refs each get their
# own group on purpose: their builds run in parallel so every published
# canary gets its sha images regardless of merge cadence. The mutable
# `:canary` channel tags are NOT written by the build matrix (which
# would race across parallel runs) — each canary-tag run retags the
# channel afterwards, only if it still matches the npm `canary`
# dist-tag, so the channel moves monotonically and always mirrors npm.
concurrency:
group: docker-${{ github.ref }}
cancel-in-progress: false
jobs:
build-and-push:
runs-on: ubuntu-latest
# The multi-arch (amd64 + QEMU-emulated arm64) production build has
# outgrown 60 minutes: the last runs to finish under the old cap took
# 39-54, and once the build crossed it every job died at the timeout.
# Each hour-long doomed run also held the per-ref concurrency slot, so
# queued master pushes superseded each other and the workflow published
# nothing at all. 120 restores headroom; the cloud job below is
# amd64-only (~10-15 minutes) and keeps its tighter cap.
timeout-minutes: 120
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history and tags so `git describe` below can compute the
# release version to stamp into the image.
fetch-depth: 0
# `.git` is dockerignored, so a running image cannot derive its own
# version and otherwise reports the source package.json placeholder in
# analytics and the debug panel. Compute it here from the pristine
# checkout (real CalVer drift from the nearest release tag) and pass it
# into both builds. Empty when no release tag is reachable — the server
# then keeps its existing fallbacks.
- name: Compute build version
id: build-version
run: |
set -euo pipefail
case "${GITHUB_REF}" in
refs/tags/nightly/v*)
# Lane tags carry the exact published version; stamp it verbatim
# instead of describing drift from the nearest stable tag.
version="${GITHUB_REF#refs/tags/nightly/v}"
;;
refs/tags/beta/v*)
version="${GITHUB_REF#refs/tags/beta/v}"
;;
*)
version="$(git describe --tags --match 'v*' --long --dirty 2>/dev/null || true)"
;;
esac
echo "version=${version}" >> "$GITHUB_OUTPUT"
echo "Stamping build version: ${version:-<none>}"
# ISO week stamp for the Dockerfile's tool layer: the layer caches
# across commits and re-pulls the @latest CLI tools when the week rolls
# over, instead of on every build.
- name: Compute tool cache epoch
id: tools-epoch
run: echo "epoch=$(date -u +%G-W%V)" >> "$GITHUB_OUTPUT"
- name: Setup pnpm
uses: pnpm/action-setup@v6
with:
version: 9.15.4
run_install: false
# No dependency cache here: this workflow publishes release images, and
# restoring a shared Actions cache into the build inputs would let a
# poisoned cache entry reach the published artifact.
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: 24
- name: Refresh lockfile for Docker build context
run: |
set -euo pipefail
pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile
changed="$(git status --porcelain)"
if [ -z "$changed" ]; then
echo "Lockfile already matches package metadata."
exit 0
fi
if printf '%s\n' "$changed" | grep -Fvq ' pnpm-lock.yaml'; then
echo "Unexpected files changed during lockfile refresh:"
echo "$changed"
exit 1
fi
echo "Using refreshed pnpm-lock.yaml in the Docker build context."
- name: Free runner disk
run: |
set -euo pipefail
echo "Disk before cleanup:"
df -h
pnpm store prune || true
sudo apt-get clean || true
sudo rm -rf \
/usr/share/dotnet \
/usr/share/swift \
/usr/local/lib/android \
/usr/local/share/boost \
/usr/local/share/powershell \
/opt/ghc \
/opt/hostedtoolcache/CodeQL \
/opt/hostedtoolcache/PyPy \
/opt/hostedtoolcache/Ruby || true
docker system prune -af || true
echo "Disk after cleanup:"
df -h
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
# Deployment tooling reads these labels from the registry to verify an
# image's schema expectations against a migrator before deploying it,
# without pulling the image. The server refuses to start when the
# database is missing bundled migrations, so orchestrators need a cheap
# way to check image/migrator compatibility up front.
- name: Compute schema migration labels
id: schema
run: |
set -euo pipefail
last=$(ls packages/db/src/migrations/*.sql | sed 's|.*/||' | LC_ALL=C sort | tail -1)
count=$(ls packages/db/src/migrations/*.sql | wc -l | tr -d ' ')
echo "last=${last}" >> "$GITHUB_OUTPUT"
echo "count=${count}" >> "$GITHUB_OUTPUT"
# Lane tag mapping: nightly/v* tags publish `:nightly`, and only
# stable v* tags move `:latest` and the versioned tags.
# `:sha-<short>` is published on every build. `:canary` is
# deliberately absent here — the channel tag is moved by the
# dist-tag-checked retag step below, never by the build matrix,
# so parallel canary builds cannot race it backwards.
- name: Docker meta
id: meta
uses: docker/metadata-action@v6
with:
images: ghcr.io/${{ github.repository }}
tags: |
type=raw,value=nightly,enable=${{ startsWith(github.ref, 'refs/tags/nightly/v') }}
type=raw,value=beta,enable=${{ startsWith(github.ref, 'refs/tags/beta/v') }}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{version}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{major}}.{{minor}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=sha
labels: |
io.github.paperclipai.schema.last-migration=${{ steps.schema.outputs.last }}
io.github.paperclipai.schema.migration-count=${{ steps.schema.outputs.count }}
- name: Build and push
uses: docker/build-push-action@v7
with:
context: .
# Pin the self-hosted image to the production stage explicitly:
# the Dockerfile now declares a later `cloud` stage, and without a
# target the default would silently become that stage.
target: production
build-args: |
PAPERCLIP_BUILD_VERSION=${{ steps.build-version.outputs.version }}
PAPERCLIP_BUILD_COMMIT=${{ github.sha }}
CLI_TOOLS_CACHE_EPOCH=${{ steps.tools-epoch.outputs.epoch }}
platforms: linux/amd64,linux/arm64
push: true
# Registry-backed BuildKit cache instead of type=gha: the Actions
# cache is capped at 10GB per repo, and two multi-arch mode=max jobs
# evict each other, so most builds ran effectively cold. The cache
# ref lives in ghcr next to the image and is written only by this
# workflow (docker.yml runs on master/tag pushes, never on PRs).
cache-from: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache
cache-to: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache,mode=max
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
# PID 1 must be an init that reaps adopted orphans. With node there, the
# orphans agent runs leave behind are never wait()ed and pin as zombies
# until the cgroup pid limit is exhausted and every fork() in the
# container fails. Run against the pushed image rather than a local
# build: the step above is multi-arch with `push: true`, so nothing is
# loaded into the runner's daemon. The cloud variant is FROM production
# and inherits the same ENTRYPOINT, so checking this image covers both.
- name: Verify PID 1 reaps orphaned processes
env:
# Through the environment, not interpolated into the script body, so
# the tag text is data rather than shell.
IMAGE_TAGS: ${{ steps.meta.outputs.tags }}
run: |
set -euo pipefail
image="$(printf '%s\n' "$IMAGE_TAGS" | head -n 1)"
test -n "$image"
echo "Verifying orphan reaping in $image"
docker run --rm -i --pull always "$image" sh -s < scripts/assert-orphan-reaping.sh
# The cloud variant carries built bundled plugins for managed deployments
# (see the `cloud` stage in the Dockerfile). It runs as its own job with no
# `needs:` on the stock publish above, so the two builds run in parallel and
# a failure or slow build in one never gates, delays, or skips the other.
# Both jobs share only the single top-level concurrency slot. Each job is a
# separate runner, so this one carries its own copy of the prep steps
# (checkout through schema labels) — the accepted cost of that isolation.
build-and-push-cloud:
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout
uses: actions/checkout@v7
with:
# Full history and tags so `git describe` below can compute the
# release version to stamp into the image.
fetch-depth: 0
# `.git` is dockerignored, so a running image cannot derive its own
# version and otherwise reports the source package.json placeholder in
# analytics and the debug panel. Compute it here from the pristine
# checkout (real CalVer drift from the nearest release tag) and pass it
# into the build. Empty when no release tag is reachable — the server
# then keeps its existing fallbacks.
- name: Compute build version
id: build-version
run: |
set -euo pipefail
case "${GITHUB_REF}" in
refs/tags/nightly/v*)
# Lane tags carry the exact published version; stamp it verbatim
# instead of describing drift from the nearest stable tag.
version="${GITHUB_REF#refs/tags/nightly/v}"
;;
refs/tags/beta/v*)
version="${GITHUB_REF#refs/tags/beta/v}"
;;
*)
version="$(git describe --tags --match 'v*' --long --dirty 2>/dev/null || true)"
;;
esac
echo "version=${version}" >> "$GITHUB_OUTPUT"
echo "Stamping build version: ${version:-<none>}"
# ISO week stamp for the Dockerfile's tool layer: the layer caches
# across commits and re-pulls the @latest CLI tools when the week rolls
# over, instead of on every build.
- name: Compute tool cache epoch
id: tools-epoch
run: echo "epoch=$(date -u +%G-W%V)" >> "$GITHUB_OUTPUT"
- name: Setup pnpm
uses: pnpm/action-setup@v6
with:
version: 9.15.4
run_install: false
# No dependency cache here: this workflow publishes release images, and
# restoring a shared Actions cache into the build inputs would let a
# poisoned cache entry reach the published artifact.
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: 24
- name: Refresh lockfile for Docker build context
run: |
set -euo pipefail
pnpm install --resolution-only --ignore-scripts --no-frozen-lockfile
changed="$(git status --porcelain)"
if [ -z "$changed" ]; then
echo "Lockfile already matches package metadata."
exit 0
fi
if printf '%s\n' "$changed" | grep -Fvq ' pnpm-lock.yaml'; then
echo "Unexpected files changed during lockfile refresh:"
echo "$changed"
exit 1
fi
echo "Using refreshed pnpm-lock.yaml in the Docker build context."
- name: Free runner disk
run: |
set -euo pipefail
echo "Disk before cleanup:"
df -h
pnpm store prune || true
sudo apt-get clean || true
sudo rm -rf \
/usr/share/dotnet \
/usr/share/swift \
/usr/local/lib/android \
/usr/local/share/boost \
/usr/local/share/powershell \
/opt/ghc \
/opt/hostedtoolcache/CodeQL \
/opt/hostedtoolcache/PyPy \
/opt/hostedtoolcache/Ruby || true
docker system prune -af || true
echo "Disk after cleanup:"
df -h
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
# Deployment tooling reads these labels from the registry to verify an
# image's schema expectations against a migrator before deploying it,
# without pulling the image. The server refuses to start when the
# database is missing bundled migrations, so orchestrators need a cheap
# way to check image/migrator compatibility up front.
- name: Compute schema migration labels
id: schema
run: |
set -euo pipefail
last=$(ls packages/db/src/migrations/*.sql | sed 's|.*/||' | LC_ALL=C sort | tail -1)
count=$(ls packages/db/src/migrations/*.sql | wc -l | tr -d ' ')
echo "last=${last}" >> "$GITHUB_OUTPUT"
echo "count=${count}" >> "$GITHUB_OUTPUT"
# Published under the same lane tag set as the self-hosted image, with a
# `-cloud` suffix (nightly-cloud, latest-cloud, <version>-cloud,
# sha-<short>-cloud). `:canary-cloud` follows the same retag-step
# ownership rule as `:canary` above.
- name: Docker meta (cloud)
id: meta-cloud
uses: docker/metadata-action@v6
with:
images: ghcr.io/${{ github.repository }}
flavor: |
suffix=-cloud,onlatest=true
tags: |
type=raw,value=nightly,enable=${{ startsWith(github.ref, 'refs/tags/nightly/v') }}
type=raw,value=beta,enable=${{ startsWith(github.ref, 'refs/tags/beta/v') }}
type=raw,value=latest,enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{version}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=semver,pattern={{major}}.{{minor}},enable=${{ startsWith(github.ref, 'refs/tags/v') }}
type=sha
labels: |
io.github.paperclipai.schema.last-migration=${{ steps.schema.outputs.last }}
io.github.paperclipai.schema.migration-count=${{ steps.schema.outputs.count }}
- name: Build and push (cloud)
uses: docker/build-push-action@v7
with:
context: .
target: cloud
# Space-separated sandbox-provider directory names to build into
# the variant; add here when managed deployments need another.
# CLOUD_BUNDLED_SERVER_DEPS names the optional peer packages the
# variant installs from server/package.json's declared version;
# add another name there when a managed tenant needs it.
build-args: |
CLOUD_BUNDLED_PLUGINS=daytona
CLOUD_BUNDLED_SERVER_DEPS=@sentry/node
PAPERCLIP_BUILD_VERSION=${{ steps.build-version.outputs.version }}
PAPERCLIP_BUILD_COMMIT=${{ github.sha }}
CLI_TOOLS_CACHE_EPOCH=${{ steps.tools-epoch.outputs.epoch }}
# amd64 only, unlike the self-hosted image above: the cloud variant
# is consumed exclusively by managed-deployment hosts, which run
# amd64. The QEMU-emulated arm64 half dominated this job's wall
# clock, and dropping it roughly halves time-to-deployable-image.
platforms: linux/amd64
push: true
# Registry-backed BuildKit cache, separate ref from the self-hosted
# job so the two parallel builds never clobber each other's cache
# manifest (see the rationale on the job above).
cache-from: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-cloud
cache-to: type=registry,ref=ghcr.io/${{ github.repository }}:buildcache-cloud,mode=max
tags: ${{ steps.meta-cloud.outputs.tags }}
labels: ${{ steps.meta-cloud.outputs.labels }}
# The cloud target installs @sentry/node at the version
# server/package.json declares, into a directory the server's own
# module resolution walks. Verify the image this job just pushed, not
# a local build, so a build-cache or layer-ordering regression is
# caught before any tenant runs the image.
- name: Verify the pushed image resolves the declared Sentry version
env:
IMAGE_TAGS: ${{ steps.meta-cloud.outputs.tags }}
run: |
set -euo pipefail
image="$(printf '%s\n' "$IMAGE_TAGS" | head -n 1)"
test -n "$image"
expected="$(node -e "process.stdout.write(require('./server/package.json').peerDependencies['@sentry/node'])")"
test -n "$expected"
installed="$(docker run --rm --pull always \
-v "$PWD/scripts/assert-cloud-image-sentry.mjs:/app/server/.ci-sentry-probe.mjs:ro" \
--entrypoint node "$image" /app/server/.ci-sentry-probe.mjs)"
echo "Declared optional peer version: $expected"
echo "Installed in the pushed image: $installed"
if [ "$installed" != "$expected" ]; then
echo "ERROR: the pushed image resolves @sentry/node@$installed, expected @sentry/node@$expected" >&2
exit 1
fi
echo "The pushed image resolves the declared @sentry/node version."
# Moves the mutable `:canary` / `:canary-cloud` channel tags. Kept OUT
# of the build jobs and serialized in its own lane, and — the load-
# bearing property — CONVERGENT rather than self-interested: a
# promotion does not promote "its own" canary, it retags the channel
# to whatever the npm `canary` dist-tag names at execution time,
# provided that version's sha images are published. GitHub's shared
# concurrency lane keeps one running and one pending promotion and
# REPLACES the pending slot with the latest enqueued — an older build
# finishing late can therefore evict the newest canary's pending
# promotion. With convergent promotion that eviction is harmless:
# whichever promotion survives resolves the current dist-tag fresh
# and lands the channel there (the current canary's images always
# exist by the time any later promotion runs, because per-tag build
# groups mean canary builds are never superseded and each run's
# promotion is gated on its own completed pushes). Every interleaving
# converges the Docker channel onto the npm channel.
promote_canary_channel:
if: startsWith(github.ref, 'refs/tags/canary/v')
needs: [build-and-push, build-and-push-cloud]
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
packages: write
concurrency:
group: docker-canary-channel-promotion
cancel-in-progress: false
steps:
- name: Login to GitHub Container Registry
uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Converge the channel tags onto the current npm canary
env:
IMAGE: ghcr.io/${{ github.repository }}
GH_TOKEN: ${{ github.token }}
run: |
current="$(curl -fsS "https://registry.npmjs.org/-/package/@paperclipai%2Fdb/dist-tags" | jq -er .canary)"
sha="$(gh api "repos/${GITHUB_REPOSITORY}/commits/$(printf 'canary/v%s' "$current" | jq -sRr @uri)" --jq .sha 2>/dev/null || true)"
if [ -z "$sha" ]; then
echo "canary/v${current} does not resolve yet; a later promotion converges the channel"
exit 0
fi
short="$(printf '%s' "$sha" | cut -c1-7)"
if ! docker buildx imagetools inspect "$IMAGE:sha-${short}-cloud" >/dev/null 2>&1; then
echo "images for ${current} (sha-${short}) not published yet; its own promotion converges the channel"
exit 0
fi
docker buildx imagetools create -t "$IMAGE:canary" "$IMAGE:sha-${short}"
docker buildx imagetools create -t "$IMAGE:canary-cloud" "$IMAGE:sha-${short}-cloud"
echo "channel tags moved to canary ${current} (sha-${short})"