mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path A fleet rollout failed with `image_manifest_not_found` for a commit that had merged and gone green. Tracing that back: the fleet resolves releases against the `-cloud` image, that image comes from `docker.yml`, and `docker.yml`'s runs on master have been reading `cancelled` for a long stretch. The cloud half was fine; the production half was hanging and taking the run down with it — and, because a run holds the concurrency slot for its whole duration, starving later commits of a build at all. ## Linked Issues or Issue Description No tracking issue — described inline, per CONTRIBUTING.md. **What's wrong.** `build-and-push` builds `linux/amd64,linux/arm64` on an x86 runner, so arm64 runs under QEMU. It hangs there — deterministically, in the same step: ``` #111 [linux/arm64 build 8/10] RUN pnpm --filter @paperclipai/server build ``` …then emits nothing until `timeout-minutes: 60` kills it. Three consecutive runs on 2026-09-04, silent for **38, 43 and 45 minutes** respectively. The amd64 leg reached `production 5/5` minutes earlier in every one. **Why it stayed hidden.** A timed-out job is reported by GitHub as **cancelled, not failed**. The run conclusion reads "cancelled", which looks like supersession rather than breakage, so the production image quietly stopped publishing. **The knock-on.** A run that burns the full hour holds the top-level concurrency slot for that hour. `cancel-in-progress: false` keeps exactly one pending slot, so merges arriving faster than one an hour supersede each other while queued. Sampling the last ten master commits, **five produced no image at all** — their Docker runs have zero job records because they never started. **Expected.** Both architectures publish, and a commit merged during a busy period still gets an image. ## What Changed `build-and-push` becomes a two-leg matrix, each on a runner of its own architecture: | platform | runner | |---|---| | `linux/amd64` | `ubuntu-latest` | | `linux/arm64` | `ubuntu-24.04-arm` | Each leg pushes **by digest** (`push-by-digest=true`, untagged), and a new `merge-and-push` job names the digests into one manifest list with the real lane tags. Nothing is publicly tagged until the merge, so a half-published multi-arch image is never a pullable state. Two supporting changes: - **Per-arch BuildKit cache refs** (`:buildcache-amd64` / `:buildcache-arm64`). Separate runners sharing one ref would overwrite each other on every build. - **The PID-1 orphan-reaping check moves to the merge job**, since that is where a tagged, pullable image first exists. It still runs against the pushed image rather than a local build, for the same reason as before. **arm64 is kept, not dropped.** The cloud variant is amd64-only and can be — managed hosts are amd64. This is the self-hosted image and ARM hosts consume it, so dropping arm64 would break them. GitHub-hosted arm64 runners are free for public repositories, which this is. `build-and-push-cloud` is untouched. It was already `platforms: linux/amd64` and has been succeeding in ~14 minutes throughout — that is why `-cloud` images exist at all. ## Verification Parsed the workflow and asserted its shape (jobs, matrix, `needs`, step order, that the cloud job is unchanged). The artifact actions are pinned by SHA with version comments, matching the repo's dominant convention — `upload-artifact` v7 and `download-artifact` v8, the same pins used across the other workflows; v8 is required for the `pattern` / `merge-multiple` inputs the merge job uses. **This PR's CI does not exercise the change.** `docker.yml` triggers on master and tag pushes, never on pull requests — deliberately, since it publishes release images. The first real run is after merge, so the check is: the next master push produces a `Docker` run whose `build-and-push (amd64)`, `build-and-push (arm64)` and `merge-and-push` jobs all succeed, and whose conclusion is `success` rather than `cancelled`. ## Risks - **Not testable before merge**, per above. If the matrix is wrong the next master push fails loudly rather than silently — which is already better than the current state, where the failure mode is an invisible "cancelled". - **First use of `ubuntu-24.04-arm` in this repo.** No other workflow uses an ARM runner. They are free for public repos, but if the label is unavailable the arm64 leg will fail to schedule and the merge will not run — no image, same as today, and visible. - **Digest-push changes the publish shape.** Between the legs finishing and the merge running, digests exist untagged in ghcr. Anything watching for tags sees no intermediate state; anything enumerating untagged manifests will see more of them. - **Cache refs change name**, so the first build after this lands is cold on both legs and will be slower than steady state. - **Does not fix the underlying QEMU hang** — it avoids it. If arm64 ever has to build under emulation again, the same stall is presumably still there. - No application code, schema, server or persistence change. ## Model Used Anthropic Claude — Opus 5, model ID `claude-opus-5`, run through Claude Code. Extended thinking enabled. Tool use throughout: GitHub Actions API to correlate run/job outcomes and read build logs, `git` for ancestry checks, and a YAML parser to validate the rewritten workflow's structure. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge A workflow change has no unit test to add, and `docker.yml` cannot run on a PR; the verification section states what to check on the first master run instead. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Devin Foley <devin@paperclip.ing>