mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
codex/plugin-task-execution
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d23cb6962 |
ci: pin a checkout-independent Rust cache path so PR lanes hit master's cache (#14394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip ships a native Runner binary, written in Rust, and seven CI lanes build it on every pull request > - `Canary Dry Run` is the slowest check on every green PR run, and most of its time is `cargo build --release` on third-party crates > - Master saves a Rust dependency cache for these lanes, but every PR lane logs `No cache found` and compiles every crate from zero > - The cache key matches, but GitHub also compares a hash of the absolute cache paths, and the master writer (RunsOn fleet, `/home/runner/_work/...`) and the PR readers (GitHub-hosted, `/home/runner/work/...`) hash different paths > - This pull request gives both sides a checkout-independent workspace path, so the hashes match and the PR lanes restore master's cache > - The benefit is about 2.5 minutes less wall clock per PR run and about 18 fewer runner-minutes per run ## Linked Issues or Issue Description No public issue exists for this problem. The description below follows the enhancement template. Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs #13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the workspace path, so none of them fixes this miss. **What existing behavior does this improve?** The `Swatinem/rust-cache` restore step in the PR workflow lanes that build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release Registry`, and the four `Verify Paperclip Runner` lanes. **Subsystem affected** CI workflows under `.github/workflows/`, their guard tests under `.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`. **Current behavior** Every PR lane logs `No cache found` although master holds an entry with the exact key. Run 36424309181 computed `v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master holds a 678 MB entry with that key. GitHub matches a cache entry on the key and on a version hash of the absolute paths in the cache. The master writer runs on the RunsOn fleet, where the checkout is `/home/runner/_work/paperclip/paperclip`. The PR readers run on GitHub-hosted `ubuntu-latest`, where the checkout is `/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…` is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The `work` paths hash to `1656e9ee…`. The key can never match, so each lane compiles every third-party crate again. **Proposed behavior** The writer and the readers pass the same checkout-independent path to `rust-cache`. Both runner layouts then produce the same version hash, and the PR lanes restore master's cache. **Reason and benefit** `Canary Dry Run` takes 533s on a green run. 251s of that is dependency compilation that a warm cache removes. Seven lanes pay this cost in every PR run. **Breaking changes** None. This change affects CI only. ## What Changed - Add a `Pin the Runner Rust workspace path` step before `rust-cache` in the master writer (`release-verify.yml`, typecheck and runner lanes) and in all four PR readers (`pr-trusted.yml`). The step creates the symlink `$HOME/paperclip-runner-rust` → `$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache` resolves the input with `path.resolve`, which does not follow symlinks, so both runner layouts now produce the same cache paths and the same version hash. `$HOME` is `/home/runner` on both images, which is why the `~/.cargo` paths already agreed. - Bump the shared keys `release-runner-v1` → `release-runner-v2` and `release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable entries are then visibly orphaned instead of sharing a key with the new ones. - Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`, and `typecheck-rust-cache`. They now require the pin step in both workflows with identical text, placed before the cache step, and they reject a `workspaces:` value that resolves under the checkout. The `pr-runner-rust-cache` test checks all four PR reader jobs and fails if a `rust-cache` step appears in a PR job that is not in its reader list. - Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the `release-runner-v2` key and to explain the pinned workspace path. ### Expected savings once merged Measured from run 36424309181. "Removed" is the dependency-compile time that a warm restore removes, minus about 18s to restore the 680 MB entry. The fleet writer's own restore shows this cost. | Lane | Today | Removed | Expected | |---|---|---|---| | Canary Dry Run | 533s | ~150s | ~380s | | Typecheck + Release Registry | 462s | ~155s | ~305s | | Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s | | Verify Paperclip Runner (rust) | 400s | ~175s | ~225s | | Build | 348s | ~130s | ~220s | | Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s | | Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s | - Wall clock per PR run: about 533s → about 385s. That is about 2.5 minutes faster to a green check set. `Canary Dry Run` stays the longest check. The rest is the non-cargo work in `release.sh` (standalone package builds ~30s, publish-payload preview ~73s). - Runner time: about 18 runner-minutes saved per PR run across the seven lanes. - The first master push after merge compiles from zero once in the fleet writer (about 4 extra minutes on that one run) and saves the v2 entry. Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key still gives a partial hit, as before. ## Verification - Run the guard tests for the three cache lanes: `node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/release-runner-cache.test.mjs .github/scripts/tests/typecheck-rust-cache.test.mjs` Result: 21 pass, 0 fail. - Run the full guard suite: `node --test '.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3 failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox temp-file ENOENT and fail the same way on the unmodified branch. - Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`. Result: 14 pass. - Local archive test: create a tar from the `_work` layout through the symlink (relative `../../../paperclip-runner-rust/target` entries, `tar -P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract it on the `work` layout. The files land in the real target directory and the symlink stays intact. - After merge, open any GitHub-hosted PR run and confirm that the seven Rust lanes log `Restored from cache key ...release-runner-v2...` in place of `No cache found`. ## Risks - Low risk. The change touches CI workflows, their tests, and one doc page. No product code changes. - If the pin step fails, `rust-cache` reports a miss and the lane compiles from zero, as it does today. The build does not break. - Both runner layouts sit four levels under `/home/runner`, so the relative `../../../` archive entries line up. The existing `~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this property. A future runner image with a different `$HOME` depth would miss the cache but would not fail the job. - `rm -rf "$pinned"` acts on the symlink itself (no trailing slash), never on the checkout behind it. It only matters on a reused runner. - Squash-merge note: the branch carries commits by `Bender (Fable)`. Add `Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the squash body to keep that authorship. ## Model Used - Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache listings, and PR operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal ticket id. The agent execution workspace fixed this branch name, so I cannot rename it. Squash-merge drops the branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Bender (Fable) <bender-fable@paperclip.local> |
||
|
|
8ee8f1fd6e |
ci: retire recurring public cloud image builds (#13827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Core publishes standard images and source verification for downstream services. > - Managed services can now compose private images from the signed standard image. > - Core still builds a second public cloud image on every master push and release. > - That duplicate producer consumes build capacity and retains an obsolete readiness contract. > - This pull request retires recurring cloud publication while preserving the standard producer and rollback artifacts. ## Linked Issues or Issue Description Refs #13797 and #13789. Related: #12856 changes image dependency packaging; it does not retire this producer. **What existing behavior does this improve?** Core's recurring Docker publication and Cloud readiness workflow. **Current behavior** Master pushes call the legacy cloud publisher from Cloud readiness. Release tags and manual Docker runs call it too. Canary promotion also requires the legacy image. **Proposed behavior** Publish standard Core images and retain `Cloud source verified v1`. Let downstream services build their managed image. Keep explicit commit previews and existing images available. ## What Changed - Remove `docker-cloud.yml`, its master and release callers, and its unused cache selector. - Remove the legacy image/migrator wait and `Cloud deployable v1` job. Keep the full source verification workflow and exact source-proof name. - Make canary promotion inspect and promote the standard image only. - Preserve signed standard-image publication, direct migrator publication, and explicit `release.yml` previews. The preview path still uses the Dockerfile `cloud` target. - Update workflow, preview, build-stamp, and packaging tests. Exercise the promotion shell with mocked registry commands, including missing-image and missing-tag cases. - Document frozen legacy aliases, consumer requirements, preview compatibility, and rollback retention. ## Verification - All 377 workflow tests pass: `node --test .github/scripts/tests/*.test.mjs`. - All 129 release-registry tests pass: `pnpm test:release-registry`. - Focused source-proof, standard-image, preview, and workflow tests pass: 256 tests. - Focused image packaging/build-stamp tests pass: 16 tests. - Actionlint passes on all three changed workflow files. `git diff --check` passes. - Full local `pnpm build` and `pnpm -r typecheck` pass. - The policy follow-up updates an old assertion that required the removed readiness job. All 37 source-proof/release-workflow tests pass locally. - Full local `pnpm test:run` did not complete successfully while the Mac ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of generated Cargo output from this isolated worktree with `cargo clean`. GitHub CI passed on the final head: 52 successful checks and 2 optional skips. - Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`: **5/5**, successful current-head check, zero review threads. - September 23 refresh: the unchanged PR head merges cleanly with current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an isolated temporary worktree, all 377 workflow tests and 29 release/preview tests pass on the combined tree. `git diff --cached --check` passes. - Refreshed Actionlint workflow validation passes with ShellCheck disabled. Full Actionlint reports the same 10 existing ShellCheck diagnostics as master, with no added diagnostics. No source changes or new PR commits were needed. - The full local build/typecheck and current-head Linux CI results above remain the verification for the unchanged PR head. They were not rerun for this metadata-only refresh. No image publication or tenant deployment was initiated for this refresh. ## Risks **Deployment prerequisite satisfied (September 23):** The combined cleanup release is deployed to staging and production, and production Support is verified. Active managed-fleet automation uses standard-image composition. Explicit immutable previews remain supported by the retained preview publisher. This PR is ready for maintainer review; keep auto-merge disabled and wait for explicit merge authorization. - A consumer still selecting `Cloud deployable v1` will stop advancing at the last legacy-ready commit. Confirm active automatic consumers use the standard-image composition contract before merge. - Legacy cloud release-channel aliases stop advancing. Standard self-hosted aliases continue. - This PR deletes no registry images, cache tags, migrators, credentials, or runner infrastructure. Existing immutable releases remain usable for rollback. - Explicit legacy previews remain for commit-specific operator deployments. Retiring that compatibility path requires a separate consumer migration. - These changes affect CI publication, not database schema or application behavior. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used repository inspection, reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4cc387f907 |
fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.
## Linked Issues or Issue Description
Refs: #13455, #13454, #13192
**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.
**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.
**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.
**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.
## What Changed
- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.
## Verification
- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
|
||
|
|
44dde2dec4 |
ci: reuse dependency caches without per-PR uploads (#13300)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud releases wait for source verification before deployment.
> - That verification reuses compiled Rust dependencies to finish
sooner.
> - PR jobs save large pnpm stores under separate merge refs and
different lockfile keys.
> - Those copies compete with master build caches for the repository's
10 GB cache limit.
> - This PR makes PR dependency caches restore-only and reuses
master-compatible keys.
> - A separate pin update will activate the reviewed workflow.
## Linked Issues or Issue Description
**What happened?**
PR merge refs accumulated roughly 700 MB copies of the same pnpm store.
Master Rust caches disappeared, and Cloud readiness run
[34656098157](https://github.com/paperclipai/paperclip/actions/runs/34656098157)
rebuilt dependencies after cache misses. The repository currently has a
10 GB limit. GitHub rejected a request for 50 GB; that setting needs
separate organization/billing access.
**Expected behavior**
PR jobs should reuse downloaded packages without evicting post-merge
compilation caches through duplicate uploads.
**Steps to reproduce**
1. Run several PRs while the checked-in lockfile needs policy
regeneration.
2. Compare the setup-node keys in PR jobs and master jobs.
3. List Actions caches by ref, key, and archive size. The PR keys repeat
across merge refs.
**Paperclip version or commit**
|
||
|
|
37d7dfb0e3 |
ci: allow dependency changes in cloud eval verification (#13286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployment requires source verification for the exact merged commit. > - Contributor PRs leave lockfile updates to a separate bot PR. > - Most release checks can refresh an outdated lockfile while installing dependencies. > - Two Runner checks still require a frozen lockfile and fail after dependency changes. > - This PR gives those checks the same install policy as the other release checks. > - A valid dependency change can become deployable without waiting for another merge. ## Linked Issues or Issue Description Refs #13257. The dependency change in #13256 exposed this gap. The separate lockfile update is #13279. Related #12115 addresses the bot PR check trigger; this PR fixes exact-source cloud verification itself. **What happened?** [Cloud readiness for |
||
|
|
19c76bfc3f |
fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - CI installs dependencies before it verifies and builds cloud artifacts. > - Install jobs share a pnpm package-store cache with lockfile refresh. > - Lockfile refresh resolves versions without downloading packages. > - That job saved an empty cache before full install jobs could save theirs. > - This PR prevents lockfile refresh from publishing that empty entry. > - Full install jobs can then populate the cache and reuse dependencies. ## Linked Issues or Issue Description Refs #13259 for the related cloud verification cache work. No duplicate empty-cache fix was found. **What happened?** Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm cache at 18:58:08 UTC on September 10. Full install jobs still restore that empty entry. The cache API reports 216 bytes for master and about 703 MB for populated entries with the same key and cache version in PR scopes. **Expected behavior** A job that installs dependencies should populate the shared package-store cache. **Steps to reproduce** 1. Run lockfile refresh with a new lockfile cache key. 2. Its resolution-only command leaves the package store empty. 3. The Node action saves the empty archive before a full install finishes. 4. Later jobs report a cache hit but download packages again. **Paperclip version or commit** Observed on master |
||
|
|
bc68312327 |
ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.
## Linked Issues or Issue Description
Refs #13243.
**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.
**Current behavior**
For merge
|
||
|
|
a23ae894a5 |
ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud images become deployable only after source verification passes. > - The typecheck job builds the native Runner binary through the server package. > - Fresh runners repeatedly compile Rust dependencies for that binary. > - This PR caches those dependencies for exact-source master verification. > - Workspace code and every typecheck still rebuild or run as before. ## Linked Issues or Issue Description Refs #13243 and #13257. **What existing behavior does this improve?** The typecheck portion of post-merge cloud source verification. **Current behavior** The typecheck job has no Rust dependency cache. An observed release build in this job took 4m 13s, including dependency compilation. **Proposed behavior** Restore dependency build outputs for canonical master pushes with the exact source SHA. Use a separate cache key from the Runner verification job, which builds other profiles. **Reason and benefit** A warm cache should remove roughly 2–3 minutes of dependency compilation from this job. Overall deployment gains depend on the remaining critical path. The first cache population still compiles from scratch. **Breaking changes** None. All checks remain enabled. Non-master callers compile without restoring or saving this cache. ## What Changed - Select the pinned Rust toolchain before the typecheck cache lookup. - Reuse the existing pinned Rust cache action with a typecheck-specific key. - Exclude workspace crates and installed cargo executables. - Test restore/save trust boundaries and document cache behavior. ## Verification - All workflow script tests pass locally, including nine new cache trust/contract cases. - actionlint passes for release-verify.yml. - Full local typecheck and build pass on the same source base (167s and 206s); `pnpm test:run` is still running and is recorded with #13257. This PR changes only the workflow, cache guard tests, and documentation. - All 32 current-head checks are successful or intentionally skipped, including the complete Linux test matrix, build, and Greptile 5/5 with no unresolved findings. Verify cache population and subsequent restore on actual master runs. ## Risks - The first run and any toolchain/dependency invalidation compile from scratch. - Cache restore/save overhead reduces the benefit for small dependency graphs. - Disable the cache step to roll back; the existing uncached build remains valid. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. Exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — focused change tests pass; full-suite local permission failures are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0b7ba4194 |
ci: route approved master cloud builds to AWS (#13243)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images built from master commits. > - Cloud image builds share GitHub-hosted capacity with other workflows. > - The organization already operates AWS runners through RunsOn Fleet. > - This pull request allows approved master builds to use a dedicated cloud Fleet. > - The benefit is separate build capacity with a quick operator rollback. ## Linked Issues or Issue Description Refs #13189, #13192. **What existing behavior does this improve?** Placement of the Docker cloud build after a master merge. **Current behavior** Every Docker cloud build uses a GitHub-hosted runner. Busy periods delay the job. **Proposed behavior** An operator variable enables the approved cloud Fleet for canonical master pushes and manual master builds. Other events, refs, and repositories use GitHub-hosted runners. ## What Changed - Add a guarded AWS runner selector to the Docker cloud job. - Keep the existing image cache, verification, and publication steps. - Test the selector against master, branch, tag, PR, fork, and disabled contexts. - Document provisioning requirements, placement checks, and rollback. ## Verification - 29 focused Node tests pass for routing, readiness, and disk handling. - The full workflow-script Node suite passes. - `pnpm -r typecheck` passes locally. - The pinned PR routing regression suite passes. The first live PR run assigned 21 jobs to the approved AWS PR group. AWS then reclaimed 16 Spot instances. The failed run is being repeated on GitHub-hosted runners while the Fleet moves to On-Demand. - Actionlint passes with existing shellcheck findings excluded (SC2012, SC2016, SC2129). - `git diff --check` passes. - Greptile reports 5/5 on commit `764d505a41dd2023751c3f361906fa9ea35bf0c6`, with no review threads. - All 30 current-head CI checks pass, including typecheck, build, all server/workspace test shards, Runner verification, and browser tests. Two Storybook checks are intentionally skipped for this change. Run: https://github.com/paperclipai/paperclip/actions/runs/34630550799 - The broader local test/build sequence is still running. This Mac has reported failures in unchanged application suites; their complete Linux CI shards pass. Local targeted workflow tests and typecheck pass. - Both On-Demand Fleets are deployed and healthy. Live master cloud-build verification follows the merge. ## Risks - Missing Fleet capacity or runner-group authorization can leave an AWS job queued. Disable `AWS_CLOUD_BUILDS_ENABLED` and rerun the workflow to use GitHub-hosted capacity. - The runner group must restrict access to this repository and the master version of `docker-cloud.yml`. - Docker needs more disk space than the PR Fleet. Provision 120 GiB disks and retain the free-space check. - This changes image build placement only. Source verification and migrator publication remain separate prerequisites. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4fde92107e |
fix(ci): reuse cloud source verification for npm canaries (#13233)
Reuse the exact master source-verification result before npm canary publication, removing a duplicate verification matrix while preserving fail-closed release checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
974949a39b |
ci: spread cloud server verification across ten runners (#13227)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud waits for source verification before deploying a new image. > - The slowest server verification job spends about ten minutes running tests. > - Each job uses one test worker to preserve test isolation. > - This pull request distributes those suites across ten standard hosted runners. > - The benefit is a shorter verification path with the same test coverage. ## Linked Issues or Issue Description **Current behavior** In [readiness run 34572340764](https://github.com/paperclipai/paperclip/actions/runs/34572340764), the slowest server job ran for 638 seconds. Test execution used 594 seconds. This held readiness behind the image job. **Proposed behavior** Use ten general server jobs in the reusable release verification workflow. Keep the three chat jobs and every existing prerequisite. The complete partition test verifies that no server suite is omitted or duplicated. **Reason and benefit** Reduce merge-to-deployable time on the existing runner type. The next longest prerequisite was Runner verification at 526 seconds, so the initial expected total gain is about two minutes rather than a halving of readiness time. Measure actual queue and execution time before claiming a result. Related: #13198 introduced the separate chat lane. #12577 refreshes duration estimates; this change leaves that manifest alone. ## What Changed - Increase the general server matrix from five jobs to ten. - Verify the ten-way partition covers the complete server suite when combined with the chat lane. - Document runner demand and the unchanged local and PR grouping. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/__tests__/run-vitest-stable-shard.test.mjs`: 29 passed. - `actionlint .github/workflows/release-verify.yml`: passed. - Full local `pnpm -r typecheck` and `pnpm build`: passed. - All latest-head GitHub CI checks passed, including the complete Linux test partition, build, typecheck, and browser gates. Greptile: 5/5 with zero open findings. - [Ten-shard timing probe](https://github.com/paperclipai/paperclip/actions/runs/34606772388): all 16 jobs passed; slowest server job 6m 23s versus 10m 38s in the earlier five-shard sample. This compares the server lane, not total readiness, and is not a controlled same-source A/B. - The full local `pnpm test:run` is also running. It has reproduced previously observed macOS-only failures in unchanged skill-cache and native-session suites; the corresponding Linux CI suites passed. Final local results will be attached separately. No affected-workflow test failed. ## Risks Five additional concurrent jobs per release verification run increase runner demand and repeated setup work. Queueing can offset the gain. Test workers, timeouts, permissions, and readiness requirements stay unchanged. Revert the matrix and its partition test to restore the previous split. ## Model Used OpenAI GPT-6 / Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected workflow tests locally and they pass; full-suite macOS limitations are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
932c8bec56 |
fix(ci): bake the managed runtime identity into cloud images (#13210)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed deployments start from the image built by the Cloud workflow. > - The managed runtime requests user and group 1001. > - The image currently builds the node user as 1000. > - Startup must remap that user, which can walk a large mounted home directory. > - This pull request uses the existing Docker build arguments to bake user and group 1001 into Cloud images. > - Matching the runtime identity removes that startup work and helps avoid health-check retries. ## Linked Issues or Issue Description Refs #13208, #1923, and #7861. Searched open and closed PRs for the Cloud UID change. The older #7861 addresses build context and volume ownership repair. This change uses the existing identity arguments in the Cloud workflow and preserves ownership repair. **What happened?** A measured rollout had a container log `Updating node UID to 1001` after startup. The container stayed at this step for at least 2 minutes 55 seconds before rollback stopped it. The baked node identity was 1000, while the managed runtime requested 1001. A health check timed out and the target required a second deployment attempt. **Expected behavior** Cloud images should already have the managed runtime identity. A matching image should skip user and group remapping. Fresh or mismatched volumes must still receive ownership repair. **Steps to reproduce** 1. Build the current Cloud image with its default build arguments. 2. Start it with `USER_UID=1001`, `USER_GID=1001`, and a populated home volume. 3. Observe the startup user remap before the application starts. **Paperclip version or commit** `fc06f7f05f42c675be71ff0927b6334405d520ed` **Deployment mode** Docker on managed hosts. ## What Changed - Pass `USER_UID=1001` and `USER_GID=1001` to the Cloud image build. - Check the pushed digest's baked identity before the entrypoint can repair it. Then check the normal entrypoint's effective identity and writable home before publishing the verified full-SHA tag. - Add a workflow regression and two entrypoint cases for a matching Cloud identity, including a mismatched volume. - Document the runtime identity and the first-build cache cost. ## Verification - Focused workflow and artifact tests: 27 passed. - Entrypoint tests: 11 passed. Actionlint passed. Full local `pnpm -r typecheck` passed. Full local `pnpm build` passed. The manual [Cloud image build](https://github.com/paperclipai/paperclip/actions/runs/34575473213) passed on the exact PR head. It checked Sentry, baked and effective identity, writable home, orphan reaping, and full-SHA publication. The new identity check took one second. All 30 PR checks passed; the Storybook workflow was intentionally skipped. Greptile reviewed commit `114d408f637a0b53e2e2b1339c263779b1e4ae54` at 5/5 with no findings or open threads. - The full local suite for the same application source was already run in #13205. Its macOS general-server phase had 10,471 passes and 70 failures in seven unchanged files. Those failures included missing Runner fixtures, filesystem errors, timeouts, a port conflict, and a load-count mismatch. After configuring Cargo and rebuilding fixtures, 37 of 38 native tests passed; one unchanged native-resume assertion still failed. Linux PR CI passed. This change adds entrypoint tests and does not change application code. ## Risks - The first build must rebuild layers that depend on the base image identity. Later builds can reuse them. - A future managed runtime identity change must update these build arguments and checks together. - The Dockerfile's self-hosted defaults remain 1000. Runtime overrides and mounted-volume ownership repair remain supported. - The observed startup delay supports this change, but fleet timing also includes provider startup, image pull, canary order, and retries. No fixed end-to-end gain is claimed before a live rollout. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and tool use. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused workflow tests; full-suite limitations are listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fc06f7f05f |
fix(ci): isolate chaos verification by caller workflow (#13208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments require verified artifacts for the merged source commit. > - Cloud readiness and the npm release independently run the same source checks. > - Their shared chaos workflow used only the source ref as its concurrency key. > - One caller could cancel the other caller's required job for the same commit. > - This pull request scopes that key to the caller workflow and source ref. > - Both callers can finish their checks without blocking deployment readiness. ## Linked Issues or Issue Description Refs #13192 and #13205. Searched for related open issues and PRs; no duplicate fix was found. **What happened?** The master push for `398d304e15739d1ee6105633bd8a0e42c929d33f` started Cloud readiness and Release together. GitHub cancelled the Cloud readiness chaos job before it acquired a runner. Its annotation reported a higher-priority waiting request for the same concurrency group. The required readiness gate cannot pass after that cancellation. **Expected behavior** Cloud readiness and Release must each finish source verification for the same SHA. Standalone chaos evals must also have a separate group. **Steps to reproduce** Merge a commit to master while the npm release queue is empty. Both callers reach the reusable chaos workflow with the same source SHA. See [the cancelled job](https://github.com/paperclipai/paperclip/actions/runs/34569569760/job/103168603926). **Paperclip version or commit** `398d304e15739d1ee6105633bd8a0e42c929d33f`. **Deployment mode** GitHub Actions on master. ## What Changed - Add the caller workflow name to the chaos workflow concurrency group. Retain source isolation and cancellation of duplicate calls within the same workflow. - Add a regression test that evaluates the group for Cloud readiness, Release, and standalone evals at the same source SHA. - Document the concurrency boundary in the readiness runbook. ## Verification - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 26 tests. - The new regression test fails against the previous concurrency key and passes with this fix. - `actionlint -shellcheck= -pyflakes= .github/workflows/runner-chaos-evals.yml .github/workflows/release-verify.yml .github/workflows/cloud-readiness.yml` passed. - `git diff --check` passed. - The full local typecheck passed for the same application source in #13205. Its macOS general-server test phase had 10,471 passes and 70 failures in seven unchanged application test files: missing Cargo/Runner test binaries, filesystem permissions, timeouts, a port conflict, and a load-test count mismatch. Linux CI test checks passed. The full local build passed with Cargo on PATH. This PR changes workflow configuration, its test, and documentation only. - All CI checks pass on the final head, including typecheck, tests, browser suites, build, and canary dry run. Greptile is 5/5 with no open findings. After merge, verify both callers' chaos jobs complete for the same master SHA and record the resulting readiness time. ## Risks - Two callers may now run chaos tests at the same time. This uses two existing GitHub runners, which is the intended cost of independent verification. - Renaming a caller changes its concurrency group. The fixed prefix keeps this child group separate from caller-level concurrency groups. - The readiness gate continues to require every verification prerequisite. No gate is bypassed. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (26 focused workflow/artifact tests) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
398d304e15 |
docs: measure cloud deployment through target health (#13205)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Hosted deployments need a verified image and migrator for the same source commit. > - The cloud readiness workflow certifies those inputs before a deployment consumer acts. > - Its completion time does not show when a tenant runs the new commit. > - This pull request documents each milestone from merge through target health and fleet completion. > - Operators can use the evidence to find the slow stage and measure a complete deployment. ## Linked Issues or Issue Description Refs #13192, #13188, and #13189. Searched related issues and PRs; no duplicate timing documentation change was found. **Issue type** Missing documentation. **Where is the issue?** `doc/cloud-build-readiness.md`, Timing and rollout. **What's wrong?** The timing instructions stop at the readiness job. That omits consumer queues, artifact resolution, and target deployment. An image can be ready while the tenant still runs an older commit. **Suggested fix** Record separate merge, image, readiness, canary health, and fleet completion timestamps for the same full source SHA. Keep preparation-only runs out of deployment results. ## What Changed - Define the evidence needed for each merge-to-deployment milestone. - Explain how consumer queues can hide upstream build gains. - Require target source identity as well as health, and report exclusions, retries, cache state, and queue conditions. ## Verification - `git diff --check` passed. - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 25 tests. - Cross-checked the readiness identity and artifact prerequisites against the current workflows and consumer contract. - Full local `pnpm -r typecheck` passed using the session's installed Rust toolchain. The full local test suite and subsequent build are still running. - All CI checks pass and Greptile is 5/5 on the exact head, with no unresolved findings. This changes one documentation file and adds no runtime behavior. ## Risks - Low risk: documentation only. Timing must still use trusted run evidence and the actual target commit. A single measured run is not a latency guarantee. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (25 focused workflow/artifact tests; full checks pending) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d56be3f3fc |
fix(ci): verify deployable cloud artifacts independently (#13192)
Verify source, build the cloud image, and wait for exact-source migrator packages concurrently. Emit Cloud deployable v1 only when every prerequisite succeeds for the merged full SHA. Co-Authored-By: Paperclip <noreply@paperclip.ing> |