mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/plugin-task-execution
16
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d23cb6962 |
ci: pin a checkout-independent Rust cache path so PR lanes hit master's cache (#14394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip ships a native Runner binary, written in Rust, and seven CI lanes build it on every pull request > - `Canary Dry Run` is the slowest check on every green PR run, and most of its time is `cargo build --release` on third-party crates > - Master saves a Rust dependency cache for these lanes, but every PR lane logs `No cache found` and compiles every crate from zero > - The cache key matches, but GitHub also compares a hash of the absolute cache paths, and the master writer (RunsOn fleet, `/home/runner/_work/...`) and the PR readers (GitHub-hosted, `/home/runner/work/...`) hash different paths > - This pull request gives both sides a checkout-independent workspace path, so the hashes match and the PR lanes restore master's cache > - The benefit is about 2.5 minutes less wall clock per PR run and about 18 fewer runner-minutes per run ## Linked Issues or Issue Description No public issue exists for this problem. The description below follows the enhancement template. Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs #13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the workspace path, so none of them fixes this miss. **What existing behavior does this improve?** The `Swatinem/rust-cache` restore step in the PR workflow lanes that build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release Registry`, and the four `Verify Paperclip Runner` lanes. **Subsystem affected** CI workflows under `.github/workflows/`, their guard tests under `.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`. **Current behavior** Every PR lane logs `No cache found` although master holds an entry with the exact key. Run 36424309181 computed `v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master holds a 678 MB entry with that key. GitHub matches a cache entry on the key and on a version hash of the absolute paths in the cache. The master writer runs on the RunsOn fleet, where the checkout is `/home/runner/_work/paperclip/paperclip`. The PR readers run on GitHub-hosted `ubuntu-latest`, where the checkout is `/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…` is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The `work` paths hash to `1656e9ee…`. The key can never match, so each lane compiles every third-party crate again. **Proposed behavior** The writer and the readers pass the same checkout-independent path to `rust-cache`. Both runner layouts then produce the same version hash, and the PR lanes restore master's cache. **Reason and benefit** `Canary Dry Run` takes 533s on a green run. 251s of that is dependency compilation that a warm cache removes. Seven lanes pay this cost in every PR run. **Breaking changes** None. This change affects CI only. ## What Changed - Add a `Pin the Runner Rust workspace path` step before `rust-cache` in the master writer (`release-verify.yml`, typecheck and runner lanes) and in all four PR readers (`pr-trusted.yml`). The step creates the symlink `$HOME/paperclip-runner-rust` → `$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache` resolves the input with `path.resolve`, which does not follow symlinks, so both runner layouts now produce the same cache paths and the same version hash. `$HOME` is `/home/runner` on both images, which is why the `~/.cargo` paths already agreed. - Bump the shared keys `release-runner-v1` → `release-runner-v2` and `release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable entries are then visibly orphaned instead of sharing a key with the new ones. - Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`, and `typecheck-rust-cache`. They now require the pin step in both workflows with identical text, placed before the cache step, and they reject a `workspaces:` value that resolves under the checkout. The `pr-runner-rust-cache` test checks all four PR reader jobs and fails if a `rust-cache` step appears in a PR job that is not in its reader list. - Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the `release-runner-v2` key and to explain the pinned workspace path. ### Expected savings once merged Measured from run 36424309181. "Removed" is the dependency-compile time that a warm restore removes, minus about 18s to restore the 680 MB entry. The fleet writer's own restore shows this cost. | Lane | Today | Removed | Expected | |---|---|---|---| | Canary Dry Run | 533s | ~150s | ~380s | | Typecheck + Release Registry | 462s | ~155s | ~305s | | Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s | | Verify Paperclip Runner (rust) | 400s | ~175s | ~225s | | Build | 348s | ~130s | ~220s | | Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s | | Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s | - Wall clock per PR run: about 533s → about 385s. That is about 2.5 minutes faster to a green check set. `Canary Dry Run` stays the longest check. The rest is the non-cargo work in `release.sh` (standalone package builds ~30s, publish-payload preview ~73s). - Runner time: about 18 runner-minutes saved per PR run across the seven lanes. - The first master push after merge compiles from zero once in the fleet writer (about 4 extra minutes on that one run) and saves the v2 entry. Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key still gives a partial hit, as before. ## Verification - Run the guard tests for the three cache lanes: `node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/release-runner-cache.test.mjs .github/scripts/tests/typecheck-rust-cache.test.mjs` Result: 21 pass, 0 fail. - Run the full guard suite: `node --test '.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3 failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox temp-file ENOENT and fail the same way on the unmodified branch. - Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`. Result: 14 pass. - Local archive test: create a tar from the `_work` layout through the symlink (relative `../../../paperclip-runner-rust/target` entries, `tar -P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract it on the `work` layout. The files land in the real target directory and the symlink stays intact. - After merge, open any GitHub-hosted PR run and confirm that the seven Rust lanes log `Restored from cache key ...release-runner-v2...` in place of `No cache found`. ## Risks - Low risk. The change touches CI workflows, their tests, and one doc page. No product code changes. - If the pin step fails, `rust-cache` reports a miss and the lane compiles from zero, as it does today. The build does not break. - Both runner layouts sit four levels under `/home/runner`, so the relative `../../../` archive entries line up. The existing `~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this property. A future runner image with a different `$HOME` depth would miss the cache but would not fail the job. - `rm -rf "$pinned"` acts on the symlink itself (no trailing slash), never on the checkout behind it. It only matters on a reused runner. - Squash-merge note: the branch carries commits by `Bender (Fable)`. Add `Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the squash body to keep that authorship. ## Model Used - Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache listings, and PR operations. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change and contains no internal ticket id. The agent execution workspace fixed this branch name, so I cannot rename it. Squash-merge drops the branch name. - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Bender (Fable) <bender-fable@paperclip.local> |
||
|
|
0ce7df2648 |
ci: keep Cloud readiness markers out of the builder queue (#13330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud consumes versioned readiness markers for each merged source commit. > - Those markers can be published only after the required checks and artifacts pass. > - The marker jobs currently wait for the same AWS runner capacity as builds and tests. > - A busy builder pool can delay readiness after all required work has finished. > - This PR moves the small readiness jobs to GitHub-hosted runners while retaining every dependency gate. ## Linked Issues or Issue Description Refs #13326 and #13328. **What existing behavior does this improve?** Time from completed Cloud verification to a deployable marker, and AWS capacity occupied by artifact polling. **Current behavior** In [Cloud readiness run 34711557083](https://github.com/paperclipai/paperclip/actions/runs/34711557083), all builds and tests finished at 18:43:05 UTC. The source marker did not start until 18:44:21, and the deployable marker did not start until 18:44:37. Merge-to-deployable was 11m33s, although the prerequisite work finished in 9m44s. **Proposed behavior** Run the artifact wait and both versioned marker jobs on `ubuntu-latest`. Keep the compute jobs on the approved post-merge AWS fleet. **Reason and benefit** Avoid builder-pool queue delays after verification finishes. This also removes the long artifact-wait job from AWS capacity. Expected savings depend on queue depth: the observed run had over 90 seconds of avoidable marker waiting. The marker commands themselves take only seconds. **Breaking changes** Runner placement changes for three bookkeeping jobs. Marker names, exact-source artifact checks, required verification, and image verification stay the same. **Additional context** Searched the related runner and Cloud readiness work. This addresses queue time observed after the parallel verification change. ## What Changed - Place the artifact wait, source-verification marker, and deployable marker on GitHub-hosted runners. - Keep all existing job dependencies, source guards, permissions, and commands. - Extend routing regressions to enforce this placement and retain fail-closed readiness gates. - Document why readiness bookkeeping uses separate runner capacity. ## Verification - Passed 433 workflow, routing, source-verification, and Cloud readiness tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/cloud-readiness.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint` and `git diff --check`. - Passed all latest-head CI gates in [run 34712624340, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34712624340), including typecheck, build, browser, Runner, and all general/serialized tests. - Attempt 1 had one localhost readiness timeout in an unchanged test. A single targeted retry passed all 166 files (3,225 tests passed, one existing skip), including all seven tests in that file. No timeout, assertion, or application source was changed; the retry is documented in the PR comment. - Local full-suite verification is limited by local disk exhaustion; the focused checks above pass. - Fresh Greptile review is 5/5 with no unresolved findings. ## Risks - GitHub-hosted capacity can also queue, but these jobs no longer compete with AWS build/test demand. The change does not reserve instances or change box sizes. - Readiness must still fail if any prerequisite fails. The existing `needs` relationships and success-only execution are preserved and tested. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
42961b6ef1 |
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e970df4c5 |
ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
f9173782cd |
feat(release): add smoke-gated nightly channel and lane-separated Docker tags (#11006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem publishes the `paperclipai` npm package set and the Docker images on two lanes: canary on every master push, and stable on manual promotion > - There is no middle ground between those lanes. Users must track every merge or wait weeks for a stable. Docker `:latest` also tracks master, so Docker users have no stable image at all > - A calm prerelease lane needs to exist, and it must never ship a build that failed its checks > - This pull request adds the nightly channel: a scheduled job that selects the newest master commit with a green canary publish, runs the full release smoke suite against that exact published canary, and only then republishes it as the nightly. It also separates Docker tags by lane, so `:latest` finally means stable > - The benefit is that users can follow prereleases at a nightly cadence with a smoke-tested guarantee, and Docker users get real `:canary`, `:nightly`, and stable image tags ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** The project publishes only `canary` (every master push) and `latest` (manual stable). Users who want prereleases without per-merge churn have no option. Docker has a second problem: master builds overwrite `:latest`, and CI-published stables never produced Docker images, because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger in `docker.yml`. No stable-versioned image exists in ghcr today. **Proposed solution** Add a `nightly` channel. A scheduled job selects the newest canary-tagged master commit, smoke-tests that exact published canary, and republishes the same commit as `YYYY.MDD.P-nightly.N` under the `nightly` dist-tag. Separate Docker tags by lane (`:canary` for master, `:nightly` for nightly tags, `:latest` plus version tags for stable tags only), and have the release jobs dispatch `docker.yml` at the new tag so lane images actually build. **Alternatives considered** Moving the `nightly` dist-tag to the existing canary version without a republish. Rejected: the version string would say `canary` while the user is on nightly, which breaks at-a-glance lane identification in bug reports and `--version` output. ## What Changed - `scripts/release-lib.sh`: channel-parameterized `next_prerelease_version` and `prerelease_tag_name` helpers (canary helpers delegate to them), a `require_channel_tag_at_head` guard, and the no-provenance retry for Sigstore transparency-log duplicates now covers the `nightly` dist-tag as well as `canary` - `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry a `canary/v*` tag, publishes the full public package set as `YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source commit `nightly/vYYYY.MDD.P-nightly.N` - `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) — select candidate, smoke it via `release-smoke.yml`, publish on green under the existing `npm-canary` environment, push the tag, dispatch `docker.yml`. New `channel` dispatch input (default `stable`, so existing stable dispatches are unchanged) with `nightly_source_version` and `dry_run` support for forced runs. The stable path now also dispatches `docker.yml` at the new `v*` tag - `.github/workflows/docker.yml`: lane tag mapping for both image jobs — master pushes publish `:canary` and no longer move `:latest`; `nightly/v*` tags publish `:nightly`; only stable `v*` tags publish `:latest` and the versioned tags. New `workflow_dispatch` trigger for the release-job dispatches. Build-version stamping uses the exact nightly version on nightly tag builds - `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch choice list - `doc/CHANNELS.md` (new): user-facing guide to the channels - `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping table, and a nightly failure playbook - `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses `npm-canary` and needs no npm trusted-publisher changes - Tests: channel-parameterized version helper coverage in `scripts/release-registry-versions.test.mjs`, and nightly flow coverage (publish identity, notes not required, canary-tag guard) in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the release script suites: 68 pass, including 6 new tests. The only failure, `acpx-patch-packaging.test.mjs`, needs installed `node_modules` and fails identically on a pristine checkout of master in the same environment - `bash -n` on both shell scripts and YAML parse of all three workflows - Live fail-path check: `./scripts/release.sh nightly --print-version` from a master tip with no canary tag fails with `HEAD has no canary/v* tag` - Live success-path check: the same command from the `canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0` - Live selection check: the candidate-selection shell logic run against the real repository selects the commit of `canary/v2026.806.0-canary.7`, which matches the current npm `canary` dist-tag exactly - After merge: dispatch `release.yml` with `channel: nightly` and `dry_run: true` to preview, then a real forced run to validate end to end before the first scheduled run ## Risks - Docker `:latest` changes meaning from "latest master build" to "latest stable release". This is deliberate and will be announced. Users who want the old behavior pull `:canary`. Until the first stable release after this change, `:latest` stays at its current (master-built) image - The nightly is a rebuild of the same source commit, not the byte-identical canary artifact that was smoked. The lockfile pins dependencies, and the publish path's registry-visibility and clean-prefix install gates still run on the nightly artifacts - All npm publishing must stay inside `release.yml` because npm trusted publishing pins that workflow file per package. The nightly jobs were added to `release.yml` for exactly that reason; this constraint is now documented in `RELEASING.md` - The stable-lane Docker dispatch fails gracefully (a warning with manual instructions) when the source ref predates `docker.yml`'s `workflow_dispatch` trigger ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fc5a30805e |
feat(cli): add managed install, update, and service lifecycle (#10045)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies > - Operators need a predictable installation path that survives beyond an ephemeral `npx` process > - A durable installation needs an owned per-user payload store, stable command shim, safe shell integration, and supported service lifecycle > - Updates must preserve recoverability by backing up data, installing side-by-side, verifying the new payload, and retaining rollback state > - Bootstrap scripts and privileged service operations must fail closed across download, filesystem, ownership, and consent boundaries > - This pull request integrates managed install, update, rollback, service, uninstall, doctor, bootstrap-installer, and runtime-serving support into one workflow > - The benefit is a recoverable, inspectable, and documented installation lifecycle with explicit safety boundaries across Linux, macOS, containers, WSL, npm, npx, and source checkouts ## Linked Issues or Issue Description ### Problem Paperclip lacks a first-class durable installation and lifecycle workflow. Operators currently have to assemble npm/npx installation, PATH setup, background-service management, updates, rollback, diagnostics, and uninstall behavior themselves. That makes upgrades harder to recover, creates inconsistent behavior across platforms, and leaves shell/download/service trust boundaries without one documented implementation. ### Proposed Solution Add a managed per-user install store and stable shim, a verified shell bootstrap installer, service lifecycle commands, install-mode-aware update/rollback behavior, doctor checks, and documentation. Managed updates back up the database, install and smoke-test a side-by-side payload, atomically switch `current`, and retain prior payloads. The shell installer pins registry/download trust boundaries and requires explicit consent for non-interactive privileged actions. ### Alternatives Considered - Keep recommending `npx`: simple for evaluation, but ephemeral and unsuitable for stable services, atomic updates, or rollback. - Require global npm installation only: familiar, but cannot provide the owned side-by-side payload store and retained rollback semantics. - Split the capability across multiple PRs: rejected because install, update, service, uninstall, bootstrap, and serving behavior share contracts and security boundaries that need review together. ### Related Pull Requests - Supersedes #10042 and #10044 with one integrated final diff. - Incorporates and replaces the closed preparatory work in #10032 and #10034. ## What Changed - Added `paperclipai install`, `update`/`upgrade`, rollback, uninstall, service lifecycle, onboarding integration, and managed-install doctor checks. - Added a private managed payload store, verified manifest/marker ownership, exclusive mutation locks, atomic manifest/current/shim writes, retained previous payloads, and provenance validation. - Added npm and GitHub-ref install sources with exact target resolution, registry isolation, database backup, side-by-side verification, atomic activation, service restart coordination, and failure rollback. - Made managed-update backups report actionable service-start and `--no-backup` recovery guidance for unreachable databases, while clean never-onboarded instances skip an empty backup. - Added systemd user and launchd service definitions, status/health/log commands, single-instance coordination, stale-port recovery, and explicit sudo/lingering consent handling. - Added the `scripts/install.sh` bootstrap path with checked two-stage downloads, pinned public npm registry usage, platform checks, dry-run/non-interactive controls, and Docker fixtures. - Added embedded Postgres/native bootstrap integration, hot-restart/systemd-notify serving support, passive update notices, configuration contracts, README/CLI/install documentation, and focused regression tests. - Security re-review should explicitly re-verify: (1) `addManagedPathBlock`/`removeManagedPathBlock` reject symlinked or non-regular rc files, assert current-user ownership, preserve restrictive modes, and replace atomically; (2) managed shim replacement rejects unsafe parents, foreign-owned or multiply linked files, and uses checked atomic replacement; (3) the shell installer and sudo path preserve explicit consent and checked downloads; and (4) installed service/runtime serving remains bound to the validated managed shim and instance configuration. ## Verification - `bash -n scripts/install.sh scripts/clean-install-git.sh scripts/clean-install-npm.sh scripts/test-install-sh-docker.sh` - `pnpm exec vitest run cli/src/__tests__/install-store.test.ts cli/src/__tests__/install-command.test.ts cli/src/__tests__/managed-install-check.test.ts cli/src/__tests__/onboard-service.test.ts cli/src/__tests__/service-health-check.test.ts cli/src/__tests__/service-manager.test.ts cli/src/__tests__/update-command.test.ts cli/src/__tests__/update-notice.test.ts packages/db/src/embedded-postgres-native.test.ts` — 9 files, 66 tests passed - `pnpm --dir cli typecheck` - `pnpm --dir cli build` - Follow-up verification: `pnpm exec vitest run cli/src/__tests__/update-command.test.ts` (14/14), `pnpm --dir cli typecheck`, `pnpm --dir cli build`, and `pnpm --filter @paperclipai/server typecheck`. - `pnpm -r typecheck` - `pnpm build` - Full `pnpm test:run` exercised all suites; an injected static AWS credential changed one unrelated doctor expectation, which passed when those credentials were removed. A second run cleared that case and exposed stale pre-existing adapter-utils `dist` output; rebuilding `@paperclipai/adapter-utils` made the isolated test pass. The updated PR CI is the authoritative clean-workspace full-suite run. ## Risks - Installer/update code writes executable shims, symlinks, shell rc blocks, service definitions, and managed payloads; ownership, regular-file, symlink, hard-link, marker, and path-containment checks fail closed before destructive changes. - The bootstrap installer executes downloaded tooling; downloads are staged and checked before execution, npm traffic is pinned to the public registry, and non-interactive privileged behavior requires explicit consent. - Linux lingering may invoke `sudo`; the command is surfaced and confirmed before execution, and unsupported service managers fall back to foreground-run guidance. - Database migrations remain forward-only; payload rollback does not reverse migrations, so managed updates create a backup before activation unless explicitly disabled. - Service restart and runtime serving touch process/port ownership; lifecycle locks, health/version checks, and stable-shim service definitions reduce split-brain and stale-process risk. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex coding agents using GPT-5.5 and GPT-5.6-sol, with reasoning, repository/API access, shell execution, and test tooling. The runtime did not expose a reliable context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
29401b231b |
fix(ci): gate new release packages on npm bootstrap (#5146)
## Thinking Path > - Paperclip is a control plane for autonomous agent companies, so its release automation is part of the core operator trust boundary. > - The affected subsystem is npm/GitHub Actions release publishing for the public monorepo packages. > - The concrete failure was that a newly added package reached `master`, the canary workflow attempted its first publish, and npm trusted publishing was not yet bootstrapped for that package. > - That means the problem is not just one broken run; it is a missing pre-merge guard that lets release-ineligible packages land and only fail once `publish_canary` runs. > - This pull request makes release enrollment explicit, validates that enrollment in CI, and adds a PR-time bootstrap check against npm for changed release-enabled package manifests. > - The result is that we keep trusted publishing, avoid teaching CI to `npm adduser`, and move this class of failure from post-merge canary time to pre-merge review time. ## What Changed - Added `scripts/release-package-manifest.json` so release-managed public packages are explicitly enrolled instead of being inferred from every non-private workspace package. - Hardened `scripts/release-package-map.mjs` to validate the manifest before release workflows rewrite versions or assemble publish payloads. - Added `scripts/check-release-package-bootstrap.mjs` and wired it into `.github/workflows/pr.yml` so PRs that change a release-enabled package manifest fail if that package does not already exist on npm. - Added release-package manifest coverage tests to `scripts/release-package-map.test.mjs` and included them in `pnpm run test:release-registry`. - Wired manifest validation into `.github/workflows/release.yml` and documented the first-publish bootstrap policy in `doc/PUBLISHING.md` and `doc/RELEASE-AUTOMATION-SETUP.md`. ## Verification - `pnpm run test:release-registry` - `./scripts/release.sh canary --skip-verify --dry-run` - Confirmed the committed diff contains no obvious PII/secrets via targeted pattern scan before pushing. ## Risks - Low risk overall: this is CI/release-policy code, not product runtime logic. - The new PR bootstrap check depends on npm metadata availability, so a transient npm outage could block a PR that changes a release-enabled package manifest. - The manifest introduces a new source of truth that must stay aligned with public package additions, but that is intentional and now enforced. ## Model Used - OpenAI Codex via the `codex_local` Paperclip adapter; GPT-5-based coding agent with tool use, terminal execution, git, and GitHub CLI. Exact served model ID/context window are not exposed by the local runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
874fe5ec7d |
Publish @paperclipai/ui from release automation
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
528f836e71 | fix: use origin for github release creation in actions | ||
|
|
3e0e15394a | chore: switch release calver to mdd patch | ||
|
|
5cf841283a | fix: correct codeowners maintainer handle | ||
|
|
f1a0460105 | fix: reset lockfile changes before release publish | ||
|
|
4d8c988dab | fix: use one workflow for npm trusted publishing | ||
|
|
21c1235277 | chore: automate canary and stable releases |