mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Published container images are the deployable unit for self-hosted and managed instances, so merge-to-image latency bounds every deploy iteration > - The docker workflow configures BuildKit caching, but builds still ran ~12+ minutes essentially cold > - Two causes: the most expensive layer (four CLI toolchains + apt) is ordered after the always-changing app copy so it can never cache, and the type=gha cache's 10GB repo cap means the two multi-arch mode=max jobs evict each other > - This pull request reorders the tool layer above the app copy (with a weekly epoch so @latest tools keep advancing) and switches both jobs to registry-backed cache in ghcr > - The benefit is that warm builds shrink to roughly the app build + push, targeting the sub-5-minute range together with the amd64-only cloud variant ## Linked Issues or Issue Description No existing public issue — inline description following the feature request template: **Subsystem affected** CI / release publishing (docker workflow, Dockerfile) **Problem or motivation** Despite `cache-from/cache-to` being configured, image builds run effectively cold: (1) the production stage installs four CLI toolchains + apt packages *after* `COPY --from=build /app /app`, and since the app copy changes every commit, that most-expensive layer rebuilds every build, per arch; (2) the `type=gha` BuildKit cache is capped at 10GB per repository, and two multi-arch `mode=max` jobs overflow and evict each other's entries. **Proposed solution** Order the tool/OS layer before the app copy (it references nothing from `/app`), refresh it weekly via a `CLI_TOOLS_CACHE_EPOCH` build arg so the `@latest` tools don't freeze in the cache, and move both jobs to registry-backed BuildKit cache (`:buildcache` / `:buildcache-cloud` refs in ghcr, no size cap, separate refs so the parallel jobs don't clobber each other). **Alternatives considered** Pinning CLI tool versions instead of the weekly epoch — more deterministic, but adds a version-bump chore; the weekly epoch preserves current freshness semantics with bounded staleness. Keeping type=gha with `mode=min` — smaller cache but loses intermediate-stage reuse, which is where most of the win is. **Roadmap alignment** Not on ROADMAP.md; CI/publishing speed improvement only. ## What Changed - `Dockerfile`: the production stage's tool/OS `RUN` (npm --global CLIs, apt, `/paperclip` setup) moves above `COPY --from=build /app /app`; new `CLI_TOOLS_CACHE_EPOCH` arg consumed by that layer. The `cloud` stage is unaffected — it only layers plugin dists on top of the finished production stage. - `.github/workflows/docker.yml`: both jobs stamp the ISO week into `CLI_TOOLS_CACHE_EPOCH`, and both switch `cache-from/cache-to` from `type=gha` to `type=registry` with per-job refs. - Includes the one-line amd64-only cloud-variant commit from #10570 so the two PRs can't conflict; if #10570 merges first, this PR rebases down to a single commit automatically. ## Verification - Image content is unchanged by layer reordering: the moved `RUN` references nothing from `/app`, and Docker layer ordering only affects caching, not the final filesystem (tool installs and app copy touch disjoint paths). - The cache ref is written only by this workflow — `docker.yml` runs on master/tag pushes, never on PRs — so the workflow's existing "no shared caches into build inputs" supply-chain stance is unchanged (BuildKit layer cache was already accepted via type=gha; the registry backend has the same writer trust). - Runtime proof lands with the first two master builds after merge: the first warms the cache, the second should show the tool layer and deps/build stages as CACHED in the build log, with wall clock dropping accordingly. I'll be watching those as part of managed-deploy work. ## Risks - Low. Worst case the registry cache misses (cold-build behavior, same as today). The weekly epoch means CLI tools update at most a week late inside images; a release built mid-week ships the tools from that week's first build. Cache refs add two small artifacts to ghcr. ## Model Used Claude Fable 5 (`claude-fable-5`, extended thinking, via Claude Code with tool use and code execution). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (no test-affecting changes) - [x] I have added or updated tests where applicable (n/a — build config and layer ordering only) - [x] I have updated relevant documentation to reflect my changes (in-file comments document both mechanisms) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge