mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
4e56afec11fe1ecd02f8ddf4d1e264d8c99c792d
293
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fa16f88d6b |
chore(lockfile): refresh pnpm-lock.yaml (#12771)
Use full dependency resolution in automated lockfile repair paths, add regression coverage, and refresh the stale Rollup snapshot. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: Codex <codex@openai.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
87d05e194b |
feat(work-products): add rich cards and run artifact inventory (#12717)
## Thinking Path > - Paperclip is the control plane for AI-agent companies. > - Agent outputs must remain visible after a run and easy to inspect from a task. > - The thread and artifact inventory need one consistent rich-card vocabulary. > - Run uploads also need durable artifact registration and producing-run context. > - Reviewers need deterministic examples for each rich-card kind and state. > - This pull request adds the shared presentation, registration, inventory, and Storybook review coverage. > - The benefit is a complete output path that reviewers can inspect without seeded data. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves work-product presentation in task threads and the task Artifacts tab. **Subsystem affected** The change affects shared work-product contracts, the runner diff path, server attachment and work-product services, GitHub metadata refresh, the React board UI, and Storybook. **Current behavior** The thread used generic cards. Some files uploaded by a run existed only as message attachments. The Artifacts tab showed a flat list without run context or filters. Storybook showed only one resting card per kind. **Proposed behavior** The thread uses rich cards for supported work-product types. Each run-produced file registers one attachment-backed artifact work product. The Artifacts tab groups outputs by run and supports filters. Storybook shows every kind and requested state, PR lifecycle states, stats variants, truncation, mobile layout, and message-tail media. **Reason and benefit** Users can identify outputs quickly. Reviewers can inspect all card permutations without creating task data. **Breaking changes** None. The metadata fields and automatic artifact registration are additive. Existing attachments and work products keep their current behavior. ## What Changed - Added a shared rich work-product card with kind-specific content and a compact inventory variant. - Added pull-request and commit diff metadata plus bounded GitHub state refresh. - Added media strips and typed file chips to message-tail attachments. - Registered each run-produced attachment as an artifact work product in the same server transaction. - Grouped task artifacts by run with agent and timestamp headings. - Added type and run filters, image thumbnails, compact cards, and a company Artifacts link. - Added a Storybook kind-by-state matrix with stats variants for all eight visual kinds. - Added PR open, draft, merged, and closed examples, long-title truncation, an exact 375-pixel viewport, and message-tail overflow coverage. - Closed reconciled runtime work products when the linked runtime stops or disappears, so the card shows `Stopped` instead of `Unhealthy`. ### Screenshots Before: one resting card per kind.  After: the kind and state matrix.  After: message-tail media at 375 pixels.  [Open the Storybook evidence viewer](https://pages.paperclip.ing/rich-work-product-storybook-20260902/). The earlier artifact inventory comparison remains available in the [artifact inventory viewer](https://pages.paperclip.ing/rich-artifacts-inventory-proof-20260902/). ## Verification - `pnpm --filter @paperclipai/ui typecheck` passed. - `pnpm check:token-gates` passed. - `pnpm build-storybook` passed. - `pnpm exec vitest run server/src/__tests__/work-product-runtime-reconciliation.test.ts` passed with 5 tests. - Chromium visual checks passed at desktop and 375-pixel widths. - All 30 latest-head GitHub checks passed. One unrelated annotation test was flaky and passed on its single retry. - Greptile passed at 5/5 with zero unresolved threads. ## Risks - Low risk. The Storybook change adds review fixtures only. The runtime fix changes read-time reconciliation without database writes. - The matrix is intentionally large so every permutation stays visible in one review surface. > I checked `ROADMAP.md`. This work does not duplicate planned core work. ## Model Used - OpenAI Codex with GPT-5 and GPT-5.6-sol across this pull request. Reasoning, tool use, and code execution were enabled. The context-window size is not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My public branch name describes the change and contains no internal task id - [x] I have run tests locally and the changed-path tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5716fe907e |
test(runner): add full-stack acceptance and eval gates (#12700)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner subsystem executes agent work across local and managed provider backends. > - The lower pull requests restore the task runtime, provider backends, and managed-provider control plane. > - The restored system needs repeatable full-stack checks before it can ship safely. > - Paid live checks also need clear access, cost, and secret controls. > - This pull request adds acceptance, live evaluation, chaos, and release gates for the restored runner stack. > - The benefit is measurable runner parity with safer release decisions. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. This change covers runner tests, release workflows, server contracts, and evaluation tools. **Problem or motivation** The runner stack did not have one complete acceptance surface for native Codex, ACPX, Claude Managed, and AWS AgentCore. Release checks could miss provider drift, task-view regressions, cost-policy errors, and destructive cleanup errors. **Proposed solution** Add a 57-cell full-stack catalog, a Daytona image, and opt-in paid workflows. Add live evaluation, chaos, cost-limit, redaction, and release contract checks. Add AWS AgentCore infrastructure and guarded provisioning tools. Keep the native runner experimental flag off by default. **Alternatives considered** We considered manual smoke tests only. They do not give repeatable evidence and they do not protect release branches. We also considered one large pull request. The stacked pull requests keep each review below the Greptile file limit. **Roadmap alignment** This work supports the shipped Cloud / Sandbox agents milestone and the shipped Agent evals & feedback milestone in `ROADMAP.md`. Related stack: - #12699 adds managed provider backends and lifecycle support. - #12691 adds qualified OpenCode and ACPX provider backends. - #12685 restores task runtime rendering and steering. ## What Changed - Add the runner full-stack harness with 57 catalog cells and 60 unit tests. - Add a Daytona runner image with digest-pinned base images and base-aware image-content checks. - Add guarded live evaluation and chaos workflows with a fixed 40-execution matrix; live and full-stack paid schedules now run only on Sundays or by manual dispatch. - Add in-flight reported-usage cost stops, post-turn cost caps, exact-threshold failure classification, secret redaction, retry classification, and actor authorization. - Reattach stream and hard-budget listeners before restart-recovery continuations so restored paid sessions cannot bypass in-flight interruption. - Preserve OpenCode usage and cost across tool-loop messages and turns while exposing an explicit current-run delta to durable accounting. - Keep PNG/WebM evidence in access-controlled artifacts only, reject SVG, and publish only pruned inert structured per-attempt evidence. - Add AWS AgentCore infrastructure, provisioning checks, and smoke tools; reject unsafe model identifiers, require exact stack ownership markers, and make failed-stack replacement explicit. - Add evaluation-session contracts and capability reports. - Add release workflow checks for immutable action pins, frozen dependency installs, exact weekly cron shape, paid-run guards, provider-secret isolation, and chaos test paths. - Reauthorize the original and triggering numeric actor IDs as the first step of every provider-secret job, including partial reruns, before checkout or provider access. - Give each full-stack matrix cell only its matching provider credential, expose Daytona only to Daytona cells, and disable shared dependency caches anywhere paid credentials or OIDC write access are present. - Protect the legacy manual E2E workflow with the same default-branch, allowlist, environment, and per-job authorization boundary. - Rotate live-eval candidates by week and retain 120 days of compatible history so the seven-week trend window remains viable. - Restore the root runner-acceptance commands and reconcile reported snapshots, raw receipts, and terminal usage without double counting or losing late usage. - Mark ACPX token deltas exact only when every budget field is present, keep cumulative cost/request authority separate, reject non-USD cost labeling, and include thought tokens in output-token budgets. - Keep `enableNativeRunner` off by default. The acceptance harness enables it only in its isolated test instance. ## Verification Passed locally: - `pnpm --filter @paperclipai/paperclip-runner typecheck` - `pnpm test:runner-acceptance:typecheck` - `pnpm test:runner-acceptance` (19 tests) - focused OpenCode proxy, driver, runnerd transport, live-session, and turn-stream tests (106 tests) - `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/live/clean-room-server.test.ts` (22 tests) - `pnpm test:e2e:runner:typecheck` - `pnpm test:e2e:runner:unit` (62 tests) - `node --test scripts/__tests__/release-verify-workflow.test.mjs` - `pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals` (22 tests) - `pnpm -r typecheck` - `pnpm build` - `node --test packages/paperclip-runner/scripts/aws-agentcore-provisioning.test.mjs` (6 tests) - `git diff --check` - `cargo test --manifest-path packages/paperclip-runner/runner/Cargo.toml -p paperclip-runner-core --lib --locked` (161 tests) - focused ACPX provider-event tests (10 tests) - The rebased PR changes 92 files. `pnpm-lock.yaml` is unchanged. I did not run paid live provider jobs or provision AWS resources. Those checks need credentials and can create cost. ## Risks The paid workflows can create provider cost. They require an allowlisted original and triggering actor, the protected `runner-e2e-paid` environment, explicit opt-in variables, and cost limits. The four provider credentials exist only in that master-only environment, which requires allowlisted reviewer approval and disables administrator bypass; repository and organization Actions scopes contain no copies. Provider usage arrives after a billable request, so the live guard cannot prevent one request from crossing a threshold. It interrupts immediately on the first reported threshold hit and permits no continuation. Visual evidence can contain secrets rendered as pixels. PNG/WebM remain only in access-controlled workflow artifacts; SVG and per-attempt XML are excluded, and S3/Pages receive a pruned structured dashboard. The AWS scripts can create cloud resources. They use explicit commands, least-privilege roles, KMS encryption, saved nonsecret metadata, and explicit teardown. This pull request does not enable the experimental native runner for existing instances. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5. The model used extended reasoning, tool use, code execution, and parallel subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
dda4dff645 |
fix(onboarding): restore browser launch and gate canaries (#12667)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The onboarding command starts the local server and opens the first-run wizard. > - Interactive onboarding stopped opening the browser by default. > - Organization creation could also succeed in the API while the wizard stayed on the name step. > - The npm canary workflow did not prove that the published package could complete this path. > - This pull request repairs the startup and organization transitions and adds an exact-version canary smoke gate. > - The benefit is a quickstart that works for users and is tested after each canary publish. ## Linked Issues or Issue Description Related: #12557 covers a separate final-route onboarding handoff. **What happened?** Interactive `paperclipai onboard` runs did not open the onboarding page. The organization API request could succeed while a same-company context update caused the wizard to stay on the organization step. The canary release lane did not test the exact published npm package through this path. **Expected behavior** Interactive onboarding must open the browser once. A successful organization request must advance to the first-agent step when the surrounding context adopts the same organization. Each published canary must install in a clean environment and reach the model connection step. **Steps to reproduce** 1. Run `npx paperclipai@canary onboard --data-dir "$(mktemp -d /tmp/paperclip-canary.XXXXXX)"` in an interactive terminal. 2. Enter an organization name while the company context refreshes from the create response. 3. Observe that the browser does not open or that the wizard can remain on the organization step after the API creates it. 4. Inspect the canary release lane and observe that no post-publish onboarding test runs against the exact npm version. **Paperclip version or commit** The issue reproduced with `2026.901.0-canary.8` and the source state before this pull request. **Deployment mode** Local trusted quickstart with embedded PostgreSQL. The install source can be npm or a source checkout. ## What Changed - Open the browser once for interactive foreground onboarding. - Preserve explicit browser opt-outs and restore the prior environment value after startup. - Accept a same-company context update after organization creation and reject a different-company takeover with an explicit error. - Export the exact canary version from the publish job. - Install and test that exact npm version in a clean Playwright smoke job through the "Connect a model" step. - Upload server logs, traces, screenshots, and the Playwright report when the canary smoke fails. - Document the interactive default and headless opt-outs. ## Verification - `pnpm exec vitest run cli/src/__tests__/onboard.test.ts ui/src/components/OnboardingWizard.step.test.tsx --reporter=dot` passes with 37 tests. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` passes with 9 tests. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/onboarding.spec.ts` passes with 2 tests. - `PAPERCLIPAI_VERSION=2026.901.0-canary.8 pnpm run test:canary-onboarding-smoke` passes against the published npm package. - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - A fresh interactive source run opens the browser and reaches "Connect a model" after organization and agent naming. ## Risks - Low risk. Automatic browser opening only applies to interactive foreground onboarding. - `PAPERCLIP_NO_BROWSER=1` and `PAPERCLIP_OPEN_ON_LISTEN=false` keep headless runs silent. - A different organization context still blocks the pending create transition. - The canary package is immutable before the smoke runs. A smoke failure leaves the package published but makes the release workflow red. - This change does not modify REST APIs, database schemas, or shared data types. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The runtime does not expose the exact deployment snapshot or context-window size. The model used reasoning, browser automation, repository tools, shell commands, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51ad751e0b |
feat(runner): integrate Codex native execution (#12616)
## Thinking Path > - Paperclip is the open source control plane for teams of AI agents. > - Agent runs currently use direct adapters and their established finalization paths. > - The new runner package needs one production integration before it can execute a real provider through the server. > - That integration must not change direct adapters or expose unsupported providers. > - The rollout must also preserve native runs that were already recorded when the feature flag changes. > - This pull request adds a default-off, Codex-only native execution path and its authority boundary. > - The benefit is a recoverable production vertical slice with explicit compatibility guards. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting server orchestration and adapter selection. **Problem or motivation** The runner package exists, but the server cannot yet start and recover a governed Codex run through it. A careless integration could also route existing direct adapters into the native runtime or lose cancellation and finalization state. **Proposed solution** Add a hidden `paperclip_runner` adapter for Codex. Keep it behind the default-off instance flag. Bind native execution, resume, cancellation, semantic tool authority, and finalization to the recorded company, issue, run, and coordinator identities. Leave every direct adapter on its existing path. **Alternatives considered** A multi-provider launch was rejected because only Codex has the complete production bridge in this series. Replacing direct adapter execution was rejected because the runner remains experimental. **Roadmap alignment** This work supports governed tool access, action attribution, and self-healing runs. It keeps the integration narrow and default-off. ## What Changed - Add the Codex-only native session executor and persisted resumption path. - Add run-scoped semantic tool projection, authorization, receipts, and idempotency. - Add audited native cancellation with durable issue and coordinator binding. - Add result fencing so a recorded result cannot reacquire the provider and run twice. - Reject fresh runner starts when the rollout flag is off while preserving recorded native recovery. - Keep direct adapters outside native status, cancellation, record creation, and finalization. - Add focused conformance, recovery, cancellation, status, portability, and compatibility coverage. ## Verification - GitHub Actions is the authoritative test environment for this large stack. - The PR policy and lightweight stack checks run while this is a middle PR. - The full required suite runs when this PR becomes the lowest unmerged or top PR. - Greptile will review this exact delta after the branch is pushed. ## Risks - The main risk is routing a legacy adapter into native execution. Runtime selection and heartbeat tests cover that boundary. - The next risk is stale or cross-company cancellation. Durable binding checks and transactional audit persistence cover it. - The adapter remains hidden and default-off. Only Codex is admitted. - There are no database migration, lockfile, or GitHub workflow changes in this PR. ## Stack 1. [Runner package, SDK, and developer tools](https://github.com/paperclipai/paperclip/pull/12608) 2. This PR: Codex production server integration 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617) ## Model Used OpenAI Codex with GPT-5, extended reasoning, repository tools, and parallel review agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
560e7e48b5 |
feat(runner): add SDK and developer tooling (#12608)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
39eafad47d |
test(e2e): shorten and split Smoke Lab coverage (#12506)
## Thinking Path > - Paperclip uses browser tests to protect critical operator flows. > - The trusted pull request workflow runs the E2E catalog on three existing runners. > - Smoke Lab was one 168-second spec, so the shard scheduler could not divide it. > - The spec also repeated service-start calls, page loads, and full-page screenshots. > - This pull request removes that repeated work and divides the scenario catalog into two independent specs. > - The benefit is a shorter Smoke Lab run and a balanced E2E lane without more AWS capacity. ## Linked Issues or Issue Description Refs: #10629 **What existing behavior does this improve?** This improves the trusted pull request E2E lane and its Smoke Lab Playwright coverage. **Current behavior** Smoke Lab is one indivisible 168-second CI spec. It starts services for every scenario, loads the same evidence page twice, and captures a full-page success screenshot for all 56 lifecycle steps. **Proposed behavior** Start Smoke Lab services once per spec. Keep the per-scenario fixture reset. Capture one representative success screenshot per scenario and keep every failure screenshot. Run P1–P4 and P5–P7 as separate specs so the existing duration-aware scheduler can put them on different runners. **Reason and benefit** The optimized lifecycle reduced local Smoke Lab wall time from 57.68 seconds to 37.04 seconds. This is a 35.8% reduction. The two halves also let the existing three runners target about 125, 124, and 124 seconds of recorded spec work instead of about 168, 125, and 124 seconds. **Breaking changes** None. The same seven scenarios and eight lifecycle steps still run. The result API still records every step. Successful non-connect steps no longer attach redundant screenshots. ## What Changed - Reused one Smoke Lab service start within each spec while retaining isolated fixture installation for every scenario. - Removed the duplicate catalog evidence navigation. - Reduced success screenshots from 56 to 7 while retaining screenshots for every failed step. - Split the shared lifecycle runner into P1–P4 and P5–P7 specs. - Mark each successful split result as partial and keep dashboard health amber until one run covers the full catalog. - Updated the duration manifest and contributor docs for the split. ## Verification - `pnpm -r typecheck` passed on Node.js 24.20.0. - `pnpm build` passed on Node.js 24.20.0. - `node --test scripts/__tests__/e2e-shard.test.mjs` passed 9 tests. - `pnpm exec vitest run ui/src/pages/tools/smoke-lab-matrix.test.ts` passed 8 tests. - Both split specs passed together on Node.js 24.20.0 after the review fixes: 2 passed in 35.7 seconds; shell wall time was 36.86 seconds. - The pre-change Smoke Lab baseline passed with a 57.68-second shell wall time. The optimized unsplit A/B run passed with a 37.04-second shell wall time. - The full local E2E catalog passed 44 tests and skipped 2 tests. One existing `pipelines-tutorial-flow.spec.ts` assertion failed again when run alone. - The broad local unit run reproduced failures in untouched workspace-runtime suites. Typecheck, build, shard tests, and all changed browser coverage pass. CI remains the authoritative full-suite result. ## Risks - The split duration weights use the measured local reduction and the previous 168-second CI weight. They should be refreshed after two real pull request runs. - Service state is shared within each half. Fixture installation still runs before every scenario to reset connection, policy, and catalog state. - Each half records passed execution with partial coverage. Dashboard health recognizes the partial flag and stays amber because no single runner covers the full catalog. A failed half still records failed/red. - Fewer success screenshots reduce redundant artifacts. Every scenario keeps its connect screenshot, and every failure still captures evidence. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context window are not exposed in this session. The model used agentic reasoning, code editing, shell execution, browser testing, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a34c615cc1 |
fix(release): skip lifecycle scripts for bundle staging (#12585)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release system builds the workspace before it prepares npm packages. > - Bundled packages then move the built files and runtime dependencies into a standalone staging directory. > - npm pack and npm publish still run package lifecycle scripts in that directory. > - The server `prepack` script requires the source workspace and cannot run from the standalone directory. > - This pull request disables lifecycle scripts only for npm operations on prepared bundled packages. > - The benefit is that release packaging uses the artifacts that the release already built. ## Linked Issues or Issue Description Refs #12584 **What happened?** The canary dry run passed bundled dependency installation and then failed while it packed `@paperclipai/server`. npm ran the server `prepack` script in the standalone staging directory. That script called `pnpm run prepare:ui-dist && pnpm run build`, which requires files from the source workspace. See the [failed canary dry-run job](https://github.com/paperclipai/paperclip/actions/runs/33399041712/job/99510762364). **Expected behavior** Bundled package packing and publishing must use the artifacts that the release already built. They must not run workspace-only package lifecycle scripts from the standalone staging directory. **Steps to reproduce** 1. Build the Paperclip workspace. 2. Prepare the bundled server package in a temporary directory. 3. Run npm pack from that directory. 4. Observe that npm runs the server `prepack` script outside the source workspace. **Paperclip version or commit** `08af15bd7629790a618e8787c11490c96a1b619a` ## What Changed - Add `--ignore-scripts` to npm pack for prepared bundled packages. - Add `--ignore-scripts` to both normal and no-provenance npm publish attempts for prepared bundled packages. - Update release helper tests to require this behavior. ## Verification - `node --test scripts/acpx-patch-packaging.test.mjs scripts/release-lib.test.mjs` (22 passed) - `pnpm test:release-registry` (98 passed) - `git diff --check` - The full test suite and build were not run locally. GitHub runs the canary dry run and full matrix. ## Risks - Low risk. The flag applies only to bundled packages that the release prepares after the workspace build. - Normal pnpm package publishing is unchanged. - Package lifecycle scripts remain in the published manifest. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The model used agentic reasoning, tool use, and code execution. The context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
08af15bd76 |
fix(release): omit dev dependencies from bundle staging (#12584)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release system publishes the server package with selected runtime dependencies inside its tarball. > - The staging step writes a temporary package manifest before it runs npm install. > - That manifest kept private development dependencies that npm still tried to resolve with `--omit=dev`. > - The canary release then stopped because the matching private runner version was not published yet. > - This pull request removes development dependencies from only the temporary install manifest. > - The benefit is that bundled package staging installs only the runtime dependencies that the tarball includes. ## Linked Issues or Issue Description Refs #12582 **What happened?** The canary release failed while it prepared `@paperclipai/server`. npm tried to resolve `@paperclipai/paperclip-runner@2026.831.0-canary.8` from the temporary staging manifest. The runner package was not published at that version, so npm returned `ETARGET`. See the [failed release job](https://github.com/paperclipai/paperclip/actions/runs/33395418107/job/99504474818). **Expected behavior** Bundled package staging must install only dependencies that the published tarball bundles. Private development dependencies must not affect the staging install. **Steps to reproduce** 1. Prepare a bundled package with a public bundled runtime dependency. 2. Add an unpublished package version to `devDependencies`. 3. Run `scripts/prepare-bundled-package.mjs`. 4. Observe that npm resolves the development dependency even when the command uses `--omit=dev`. **Paperclip version or commit** `5a988df600ebda30e446496862bf83c76d6d53d6` ## What Changed - Remove `devDependencies` from the temporary manifest used for bundled package installation. - Keep the final publish manifest unchanged. - Add unit and staging regression checks for the unpublished development dependency case. ## Verification - `pnpm install --frozen-lockfile` - `node --test scripts/acpx-patch-packaging.test.mjs` (12 passed) - `pnpm test:release-registry` (98 passed) - `git diff --check` - The full test suite and build were not run. This change has focused release-packaging coverage. ## Risks - Low risk. The change affects only the temporary manifest used to install bundled runtime dependencies. - The script restores the complete publish manifest before it creates the package tarball. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The model used agentic reasoning, tool use, and code execution. The context window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8610e7934e |
fix(release): bundle vendored runner ACPX runtime (#12582)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The npm release includes the Paperclip server and a vendored runner. > - The vendored runner imports ACPX when it starts a Codex agent. > - The server package did not include the ACPX version that the runner needs. > - A fresh canary install therefore stopped with `ERR_MODULE_NOT_FOUND` after onboarding. > - This pull request bundles the patched ACPX runtime with the server package. > - The benefit is that a fresh npm install can load the vendored runner. ## Linked Issues or Issue Description No public issue exists for this bug. A GitHub search found no duplicate or related pull request. **What happened?** A fresh `npx paperclipai@canary onboard` command completed onboarding. The server then failed to start. Node could not resolve `acpx` from the vendored Paperclip runner. **Expected behavior** The server should start after onboarding from a fresh npm cache and a temporary data directory. **Steps to reproduce** 1. Run `npx paperclipai@canary onboard --data-dir "$(mktemp -d /tmp/paperclip-canary.XXXXXX)"`. 2. Select Quickstart. 3. Start Paperclip. 4. Observe `ERR_MODULE_NOT_FOUND` for `acpx`. **Paperclip version or commit** `paperclipai@2026.831.0-canary.6` **Deployment mode** Other: local trusted Quickstart through `npx`. **Installation method** npm through `npx`. **Agent adapter(s) involved** Codex. **Database mode** Embedded PGlite. **Access context** Board operator during onboarding. **Node.js version** Node.js 26.4.0. **Operating system** macOS. **Relevant logs or output** ```shell Cannot find package 'acpx' imported from .../node_modules/@paperclipai/server/dist/vendor/paperclip-runner/drivers/acpx/codex-runtime-adapter.js ``` **Relevant config (if applicable)** No custom configuration was required. **Additional context** The published adapter utilities contain a nested `acpx@0.12.0`. Node cannot resolve that nested package from the sibling vendored runner. Installing `acpx@0.13.1` at the clean package root makes the failing runner import succeed. **Privacy checklist** The log excerpt contains no user path, token, company name, or other private value. ## What Changed - Added `acpx@0.13.1` as a bundled server runtime dependency. - Added a version-specific patch check for the ACPX versions used by the server and adapter utilities. - Added release-package coverage for the server ACPX bundle. ## Verification - `pnpm test:release-registry` passed 98 tests. - `node --test scripts/acpx-patch-packaging.test.mjs` passed 12 tests. - `pnpm exec vitest run server/src/__tests__/server-package-build-script.test.ts` passed 4 tests. - `node --test scripts/release-package-map.test.mjs` passed 12 tests. - `pnpm -r typecheck` passed. - A clean extracted server tarball contained the patched `acpx@0.13.1` runtime. - The previously failing vendored runner module imported from that clean tarball. - The repository-wide test suite was stopped before completion at the maintainer's request because it takes too long for this urgent packaging fix. ## Risks - Risk is low. - The server tarball grows because it now contains ACPX and its production dependencies. - The release stager now uses a version-specific marker to verify the ACPX patch. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. - The model used reasoning mode, tool use, and code execution. - The context window size was not disclosed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ca24bba3c |
feat(runner): pin the Codex ACPX runtime (#12400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The package-local host boundary is ready for a concrete ACP implementation, but the first production profile is Codex only. > - ACPX must not inherit the server process environment or choose an executable by pathname after admission. > - Codex must not re-enable ambient apps, memory, skills, MCP configuration, or instructions inside its isolated home. > - This pull request pins only the two required production packages and applies narrowly tested host patches. > - The benefit is a minimal dependency boundary that follows the repository's CI-owned lockfile process. ## Linked Issues or Issue Description **Agent or provider** Codex through `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. **Why this adapter is useful** The injected runtime host needs a concrete ACP session manager and the exact reviewed Codex ACP server. Upstream ACPX does not yet expose a host-owned spawn callback, and upstream Codex ACP does not yet apply Paperclip's isolated instruction, MCP, app, memory, and skill boundary. Both behaviors are required before the dependency can execute inside the runner. **How the agent is invoked** The next pull request will adapt these pinned packages to the private runtime host. ACPX receives a host-owned callback that consumes the already verified executable lease. Codex receives only the isolated environment, explicit base instructions, explicit MCP servers, and the skills rooted in its private `CODEX_HOME`. This pull request alone does not spawn either package or register an adapter. **Additional context** This pull request is stacked on #12399. It adds no Pi, Claude, AWS, SDK, lab, browser, or UI dependency. It intentionally does not commit `pnpm-lock.yaml`: the repository policy job regenerates a manifest-only PR lockfile artifact for downstream frozen installs, and the lockfile bot updates master separately. ## What Changed - Pin `acpx` to `0.13.1` and the Codex ACP server to `1.6.2` in the runner package. - Register both patches in the pnpm 9 root configuration and newer-pnpm workspace configuration. - Preserve the existing embedded-Postgres and ACPX 0.12 patch entries used by other packages. - Patch ACPX to evaluate an allowlisted environment at child-spawn time and keep spawn cwd out of provider-visible session identity. - Patch ACPX to accept a host-owned spawn callback with the resolved arguments and options, allowing the verified command lease to own execution. - Patch Codex ACP to retain runner-owned MCP server identity in permission requests. - Patch Codex ACP to pass explicit Paperclip base instructions on both start and resume. - In isolated mode, disable ambient apps, memory, and existing MCP configuration; load skills only from `CODEX_HOME`; and configure only requested servers. - Add a package contract test that enforces exact versions, Codex-only dependency scope, both pnpm patch registries, and every required patch hook. ## Verification - Both patch files dry-apply successfully to fresh published tarballs for `acpx@0.13.1` and `@agentclientprotocol/codex-acp@1.6.2`. - A local no-lockfile install applied both patches; their runtime markers and exact installed versions were inspected. - Runner TypeScript typecheck — passed against the patched packages. - Runner package tests — passed: 16 Node protocol/package tests and 426 Vitest tests. - `pnpm -r typecheck` — passed for all applicable workspaces. - `pnpm build` — passed, including runner binary, server, UI, and workspace packages. - `git diff --check` — passed. - The diff contains 6 files and does not change `pnpm-lock.yaml`, a GitHub workflow, server selection, or UI behavior. ## Risks The primary risk is drift between published package contents and checked-in compiled patches. Exact versions are pinned, both patches are exercised by package-contract gates, and CI performs the authoritative regenerated-lockfile frozen install. The spawn callback does not grant a new executable path: the following adapter must consume the opaque verified command lease. Codex isolation changes activate only when `PAPERCLIP_ACPX_ISOLATED_CONTEXT=1`, so existing direct Codex adapters are unaffected. ## Model Used OpenAI Codex with GPT-5 and repository tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked an existing public item or described the issue in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal task identifier - [x] I have run the affected tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have documented the dependency, patch, isolation, and lockfile boundaries - [ ] All applicable GitHub Actions are green - [ ] Greptile is 5/5 with every actionable comment resolved - [x] I will address all review findings before requesting merge |
||
|
|
a560b48d6d |
feat(apps): refine Postman and Shopify setup (#12357)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Provider catalogs must match each provider's current protocol and credential contract. > - Postman method labels and API-key placement were outdated. > - Shopify now offers a UCP commerce endpoint that needs a managed agent-profile argument. > - This pull request updates both providers and documents the complete connection-authoring workflow. > - The benefit is accurate setup, safer runtime defaults, and a repeatable provider review process. ## Linked Issues or Issue Description Refs #11965 This is stack 10 of 11. It depends on stack 9 and preserves the final catalog work recovered from #11965. Related: #5904 covers Shopify skill routing. This pull request covers the Apps connection contract instead. ## What Changed - Update Postman hosted MCP methods, capability choices, default selection, and bearer-token placement. - Add Shopify UCP commerce and Storefront compatibility methods with public-store prerequisites. - Inject the reviewed Shopify UCP agent profile at runtime and remove that managed field from user input schemas. - Classify Shopify checkout completion and cancellation as destructive actions. - Expand the connection authoring runbook from provider research through verification and pull request handoff. - Add focused shared, server, and UI coverage. - Make the approved-execution waiter phase-aware so slow preparation cannot consume the provider execution timeout and grace period. - Settle legacy pre-execute-on-approve requests and invocations as failed, clear their stale idempotency key, and allow a fresh governed approval instead of leaving work stuck in `executing`. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts ui/src/pages/apps/AppsConnect.test.tsx -t "Postman|Shopify|normalizeConnectionMethodConfig|classifyRisk"` (16 passed) - `pnpm exec vitest run server/src/services/approved-execution-wait.test.ts` (4 passed) - `pnpm exec vitest run server/src/__tests__/tool-gateway.test.ts -t "enforces policy, approvals, retries, rate limits, and company boundaries for connected remote MCP calls"` (1 passed) - `pnpm exec vitest run server/src/__tests__/tool-gateway-service.test.ts` (21 passed; includes legacy approval settlement and fresh-approval recovery) - `pnpm --filter @paperclipai/server typecheck` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` ## Risks - Shopify UCP calls now include a Paperclip-managed agent profile that overrides caller input at the same path. - Postman EU credentials now use the hosted MCP server's bearer-token contract instead of the general REST API header. - The catalog generator and checked-in definitions change together to prevent regeneration drift. - Approved execution preparation has an explicit two-minute bound; provider execution retains its own 65-second timeout and persistence grace starting from durable provider start. - Legacy approvals created before execute-on-approve are intentionally terminalized and must be requested again under the current signed contract. > I checked `ROADMAP.md`. This provider update does not duplicate planned core work. The related open Shopify PR addresses skill routing, not Apps connections. ## Model Used OpenAI Codex, GPT-5. The runtime exact model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fcb2e99e8f |
feat(apps): expand the self-serve connection catalog (#12344)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A useful app store needs accurate and selectable provider definitions. > - Local brand assets now cover the expanded provider set. > - Provider methods differ in transport, authentication, ownership, and required scope. > - This pull request expands the catalog and encodes those provider contracts. > - The benefit is a larger self-serve store with explicit setup choices. ## Linked Issues or Issue Description Refs #11965 This is stack 6 of 11. It depends on stack 5 and replaces another reviewable part of #11965. ## What Changed - Add and update provider definitions for the self-serve catalog. - Add Google Workspace connection methods and capability profiles. - Add catalog generation, ingestion, URL matching, and contract tests. - Update legacy key tests to use a provider that still uses header credentials. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 164 tests passed. - `pnpm build` ## Risks - An incorrect provider definition can offer the wrong setup method. - Contract tests verify transport, authentication, and provider URL behavior. - The change does not add a database migration. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Refs #` or (b) described the issue in this pull request - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6244e4cf32 |
feat(apps): add Composio and Gmail connectors (#12342)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - App connections need both direct providers and managed provider hubs. > - The grant layer now defines safe credential ownership. > - Composio needs parent and child connection lifecycle rules, and Gmail needs governed setup. > - This pull request adds both connector families on the grant foundation. > - The benefit is broader app access without weakening credential isolation. ## Linked Issues or Issue Description Refs #11965 This is stack 4 of 11. It depends on stack 3 and replaces another reviewable part of #11965. ## What Changed - Add Composio parent and child connection support. - Add Gmail connection setup and governance. - Preserve credential paths and remove duplicate binding declarations. - Cascade Composio pause and restore actions to child connections. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 164 tests passed. - `pnpm build` ## Risks - Parent lifecycle changes can affect every Composio child. - The service restores only children whose provider accounts remain active. - Credential binding paths are normalized before secret resolution. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b51112798f |
feat(apps): improve gateway and workspace connection UX (#12340)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - App connections must work in both the operator UI and agent tool gateway. > - The first stack layer adds secure remote connections. > - Operators still need clear setup, test, and recovery states. > - This pull request adds the gateway behavior and the workspace connection experience. > - The benefit is a connection flow that is easier to understand and recover. ## Linked Issues or Issue Description Refs #11965 This is stack 2 of 11. It depends on stack 1 and replaces another reviewable part of #11965. ## What Changed - Improve remote tool gateway connection behavior. - Add clearer app setup, test, and recovery states. - Add focused server and UI tests for the new paths. - Keep the diff isolated from later identity and catalog work. - Stabilize DNS-pinned remote HTTP protocol fixtures and the managed-runtime public-origin fixture for this independently tested layer. ## Verification - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` (150 passed) - `pnpm test:run` - `pnpm check:token-gates` - `pnpm build` ## Risks - Gateway errors now surface through new user-facing states. - A stale connection can require a new setup attempt. - The change does not add a database migration. - The injected HTTP transport and public URL are test-only fixtures; production DNS pinning and runtime behavior are unchanged. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cabc9146d0 |
feat(apps): add secure remote MCP and PostHog setup (#12339)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give those agents governed access to external tools. > - Remote MCP setup needs secure endpoint validation and durable credentials. > - PostHog needs both browser sign-in and personal API key setup paths. > - This pull request adds the shared remote MCP foundation and the PostHog definition. > - The benefit is a secure and reusable base for later app connection work. ## Linked Issues or Issue Description Refs #11965 This is stack 1 of 11. It replaces the first reviewable part of #11965. ## What Changed - Add guarded remote MCP setup and credential handling. - Add PostHog OAuth and API key connection methods. - Add focused server, shared contract, and UI coverage. - Keep the migration replay-safe and idempotent. - Give the late-close security regression the same 10-second CI headroom as the adjacent real-timer handshake test. - Synchronize fake-timer handshake tests at the exact ensure-session boundary so real filesystem setup cannot race the fake deadline. - Drive PTY overflow coverage only after listener registration so scheduling cannot reorder the test fixture. ## Verification - pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts server/src/__tests__/plugin-worker-manager.test.ts (220 passed; affected cases also passed five focused stress repetitions) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never leaks a sandbox-provided value from a late close rejection into logs or the result"` (1 passed) - `pnpm exec vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never promotes a late ensureSession resolution|closes a late-resolving real handle exactly once"` (2 passed) - `pnpm -r typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - `pnpm --filter @paperclipai/db check:migrations` - `pnpm build` ## Risks - Remote endpoint validation can reject configurations that previously passed without checks. - OAuth configuration errors can block setup until the operator corrects the provider settings. - The migration uses guarded statements so repeated execution is safe. - The test-only synchronization changes do not affect runtime behavior; they remove filesystem/fake-clock and listener-registration races observed under parallel CI load. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
39b8ee2960 |
ci: optimize checks for stacked pull requests (#12507)
## Thinking Path > - Paperclip uses GitHub Actions to protect changes before they enter `master`. > - GitHub evaluates every pull request in a native stack against the stack base. > - The current workflow therefore starts the complete CI matrix for every layer in a stack. > - A large stack can queue many copies of the same integrated verification and delay every pull request. > - GitHub provides stack position and base metadata so workflows can select merge-relevant layers. > - This pull request keeps policy and required check names on every layer, but runs full CI only for ordinary pull requests, the top layer, and the lowest unmerged layer. > - The benefit is much lower CI load without weakening the required-check contract. ## Linked Issues or Issue Description **What existing behavior does this improve?** The trusted pull request workflow currently runs every test, build, canary, and E2E lane for every pull request in a native stack. **Subsystem affected** GitHub Actions pull request verification. **Current behavior** A stack with 61 pull requests can start 61 complete CI matrices after a cascading rebase. **Proposed behavior** Run the always-on policy job and stable required-check aggregators for every layer. Run the complete verification matrix only for ordinary pull requests, the top stack layer, and the lowest unmerged stack layer. **Reason and benefit** The top layer verifies the integrated stack. The lowest unmerged layer verifies the current merge candidate. Middle layers keep branch-protection checks without consuming the complete runner matrix. **Breaking changes** Middle stack layers no longer run the complete CI matrix. Their `ci / verify` and `ci / e2e` checks still require the policy job to pass and require every expensive lane to be intentionally skipped. ## What Changed - Add a fail-safe stack scope decision to the trusted PR runner gate. - Run typecheck, general tests, build, serialized tests, canary, and E2E shards only for ordinary, top, and lowest-unmerged pull requests. - Preserve the required `ci / verify` and `ci / e2e` names on every layer. - Make the required aggregators distinguish valid middle-layer skips from failures or missing scope decisions. - Add regression coverage for ordinary, top, bottom, middle, and malformed stack metadata. ## Verification - `node --test scripts/__tests__/e2e-shard.test.mjs` — 11 tests passed. - `actionlint .github/workflows/pr-trusted.yml .github/workflows/pr.yml` — passed. - `git diff --check origin/master...HEAD` — passed. - The caller remains pinned to the current trusted workflow. A separate activation change must advance the immutable SHA after this pull request lands. ## Risks - Incorrect stack classification could skip important jobs. Missing or malformed stack metadata defaults to full CI. - Middle-layer required checks depend on the policy job and verify that all expensive jobs have the `skipped` result. - The reusable workflow change does not become active until the immutable caller SHA advances in a separate change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment suffix and context window are not exposed. The model used reasoning, repository tools, code execution, Git, and GitHub API access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
64b7dce0ad |
refactor(adapter-utils): replace the process-wide byte ledger with route-local byte bounds (#12465)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The adapter layer carries sandbox requests to host processes. > - The HTTP/2 bridge used one process-wide byte ledger for all routes. > - One busy route could exhaust that shared budget and move another route to file transport. > - This pull request gives each host retention site a fixed byte bound and limits concurrent HTTP/2 streams. > - The benefit is local protection: one route cannot consume the byte budget of another route. ## Linked Issues or Issue Description **What happened?** The HTTP/2 bridge used one aggregate byte ledger for retained bytes across all routes. A busy route could exhaust the shared budget and force an unrelated route to use file transport. **Expected behavior** Each route should protect its own retained bytes. A reset on one HTTP/2 stream should cancel only that stream's host forward. **Steps to reproduce** 1. Start the HTTP/2 bridge with multiple sandbox routes. 2. Send enough retained data through one route to reach the aggregate byte limit. 3. Send a request through a sibling route. 4. Observe that the sibling route can fall back to file transport because the first route used the shared ledger. **Paperclip version or commit** `47639e227e78e3c5e0dd1a3c0e2d792fe86895a3` **Deployment mode** Built from source with the adapter-utils and server test suites. ## What Changed - Bound each host retention site with a fixed local byte limit. - Limited concurrent live HTTP/2 streams with one built-in stream limit. - Bound each host forward and response-body read to its own HTTP/2 stream lifetime. - Removed the process-wide byte ledger, its environment override, its metrics, and its file-transport fallbacks. - Added tests for the stream limit, host body budget, and sibling-stream cancellation. ## Verification - Run `pnpm vitest run --project adapter-utils`. - Confirm that 996 adapter-utils tests pass. - Confirm that `test_live_forward_work_never_passes_the_stream_limit` passes. - Confirm that `test_the_host_body_budget_matches_the_stream_limit` passes. - Confirm that the sibling-stream cancellation test passes. - Run `pnpm tsc --noEmit`. - Confirm that all pull request checks pass. ## Risks The bridge no longer uses a process-wide byte ledger. A local bound or stream limit that is too low can reject or delay valid work. The tests cover the new limits and stream cancellation behavior. ## Model Used OpenAI GPT-5 Codex. Runtime model ID: GPT-5. The model used code execution and repository tools. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1da6b37fc5 |
fix(ci): regenerate stale stacked lockfiles (#12461)
## Thinking Path > - Paperclip uses trusted GitHub Actions workflows to verify every pull request > - Native stacked pull requests use another pull request branch as their base > - A parent layer can change a package manifest without committing `pnpm-lock.yaml` > - A child layer can inherit that manifest change without changing a manifest itself > - The current policy skips lockfile regeneration for that child and downstream frozen installs fail > - This pull request validates the complete merge tree and shares a regenerated lockfile only when needed > - The benefit is reliable stacked pull request verification without weakening the trusted workflow boundary ## Linked Issues or Issue Description **What happened?** A stacked child pull request inherited a package manifest change from its parent. The child did not change a manifest itself. The policy job skipped lockfile regeneration. Downstream jobs tried to restore an artifact that did not exist and then failed during frozen dependency installation. **Expected behavior** The policy job must validate the complete pull request merge tree. It must upload a regenerated lockfile when the checked-in lockfile is stale, including on a stacked child layer. **Steps to reproduce** 1. Create a parent pull request that changes `package.json` without committing `pnpm-lock.yaml`. 2. Create a child pull request on that branch without another manifest change. 3. Run the trusted pull request workflow for the child. 4. Observe that frozen dependency installation fails because no `pr-lockfile` artifact exists. **Paperclip version or commit** `f173ee09fa5c2ced7806bba47b54c3df853ab4df` **Deployment mode** GitHub Actions trusted pull request workflow. **Agent adapter(s) involved** Not adapter-specific. This is a core CI workflow bug. ## What Changed - Regenerate the lockfile from every checked-out merge tree. - Compare the generated lockfile with the checked-in copy before upload. - Download the artifact only when the policy job reports that it uploaded one. - Fail closed when a reported artifact is missing. - Add a workflow contract test for stacked lockfile handling. - Keep the caller pinned to the last merged trusted SHA; after this implementation merges, a separate activation PR will advance the immutable pin to its merge commit. ## Verification - `actionlint .github/workflows/pr-trusted.yml` - `node --test scripts/__tests__/e2e-shard.test.mjs` ## Risks - The policy job runs one lockfile-only install for every pull request. This can add a small amount of CI time. - A missing artifact now fails immediately when the policy job reports an upload. This is intentional because it exposes workflow corruption. - No runtime or product behavior changes. - The implementation/activation split is intentional: unmerged PR-authored workflow code must never execute on trusted runners. ## Model Used OpenAI Codex with model `gpt-5`, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
c916af0cc0 |
ci: call trusted PR workflow (#12439)
## Thinking Path > - Paperclip uses pull request CI to validate each proposed change > - The existing workflow defines every heavy job in a PR-controlled file > - A trusted reusable workflow now contains the synchronized CI definition > - The caller must use an immutable default-branch SHA > - This pull request replaces the duplicate job list with that pinned caller > - The benefit is automatic secure runner selection without workflow drift ## Linked Issues or Issue Description Refs #12436 Refs #12438 **What existing behavior does this improve?** This improves how the pull request workflow selects trusted CI capacity. **Subsystem affected** Cross-cutting CI automation. **Current behavior** The active workflow contains a duplicate list of all heavy jobs. It cannot use the administrator-controlled runner gate. **Proposed behavior** The active workflow calls the synchronized trusted workflow at an immutable SHA. The trusted workflow selects GitHub-hosted or isolated AWS capacity from the validated contributor identity. **Reason and benefit** The thin caller prevents pull request changes from replacing the external-runner security gate. It also keeps runner selection automatic. **Breaking changes** The check names gain the reusable workflow job prefix. AWS routing remains disabled until the canary starts. ## What Changed - Replaced the duplicated heavy CI job list with one reusable-workflow call. - Pinned the call to the reviewed default-branch commit. - Limited the caller token to actions, contents, and pull request read access. ## Verification - actionlint on both workflow files - Trusted-routing tests - Confirmed the pinned SHA contains the workflow and is an ancestor of master - Full AWS and GitHub runner-boundary verification with routing disabled ## Risks The check context names change when GitHub expands the reusable workflow. The rollout verifies the new aggregate contexts before branch rules change. The repository kill switch remains off during this pull request. ## Model Used OpenAI Codex with GPT-5, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked a related public PR or described the issue with the matching template fields - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
dbf052577d |
Follow the current onboarding arc in the release smoke (#12423)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A release is gated by the release smoke: it installs the published
`paperclipai` artifact into a Docker container and drives the sign-in →
onboarding → first-agent path with Playwright
> - That suite runs only from the release pipeline, never on a pull
request, so it sees the UI only after the UI has already changed
> - The onboarding wizard was rebuilt into the agent arc. The "Name your
organization" step, the "Start Onboarding" launcher, and the agent role
picker are all gone
> - The spec still waited for those, so it failed on its first assertion
and blocked every nightly and beta release
> - The failure was also hard to read. The workflow uploaded no
container logs, because it learned the container's name only after the
harness succeeded, and the harness ran the container with `--rm` and
deleted it before anything read it
> - This pull request rewrites the spec to follow the current arc, and
repairs the log capture at both ends
> - The benefit is that nightly and beta releases are unblocked, and the
next failure arrives with the logs attached
## Linked Issues or Issue Description
No existing issue. Describing it inline, following
`.github/ISSUE_TEMPLATE/bug_report.yml`.
Refs #12274 (removed the company-naming step from the wizard).
Refs #12135 (the previous alignment of this spec, before #12274).
Refs #12316 (open; also edits `scripts/docker-onboard-smoke.sh`, in the
bootstrap helpers rather than the container lifecycle, so the two
changes do
not overlap. Whichever lands second should rebase and re-run).
**What happened?**
The release smoke fails.
`tests/release-smoke/docker-auth-onboarding.spec.ts`
never gets past its first wait:
```
✘ tests/release-smoke/docker-auth-onboarding.spec.ts:43:3 › Docker authenticated onboarding smoke › logs in, completes onboarding, and hires the lead agent
Error: expect(locator).toBeVisible() failed — element(s) not found (timeout 20000ms)
> 33 | await expect(wizardHeading.or(startButton)).toBeVisible({ timeout: 20_000 });
```
The spec waits for an `h3` reading "Name your organization" or a
"Start Onboarding" button. Neither exists. #12274 removed the
company-naming
step; the string now survives only in a code comment and in
`ui/src/components/OnboardingWizard.step.test.tsx`, which asserts it is
*absent*. The steps after the first wait are stale too: the CTA on step
1 is
"Continue" and not "Next", the organization input's placeholder changed,
and
the agent step's `#onboarding-agent-role` picker is gone, so every
onboarding
hire is filed under the neutral `general` role.
The suite runs only from the release pipeline, so nothing on a pull
request
saw the drift. Both `smoke_nightly` and `smoke_beta` call the same
reusable
workflow, so every nightly and every beta was blocked.
The failure also arrived without diagnostics. The job's "Capture Docker
logs"
step is `if: always()`, but it is guarded on `SMOKE_CONTAINER_NAME`,
which the
"Launch Docker smoke harness" step writes to `$GITHUB_ENV` only *after*
the
harness returns. On any failure before that the guard is false, the step
does
nothing, and the upload reports "No files were found". Below that,
`scripts/docker-onboard-smoke.sh` starts the container with
`docker run -d --rm`, so the `docker stop` in its EXIT trap deletes the
container and its logs together — and a container that crashes on its
own is
removed the instant its process exits.
**Expected behavior**
The spec walks the onboarding arc the app actually presents, and proves
the
company is created, the lead agent is hired, and the first task is
seeded and
dispatched. When the smoke fails, the run's artifact carries the
container's
logs.
**Steps to reproduce**
1. Run the Release Smoke workflow against a published artifact that
carries
#12274, or run it locally:
`PAPERCLIPAI_VERSION=2026.828.0-canary.3 SMOKE_DETACH=true
./scripts/docker-onboard-smoke.sh`
2. Run `pnpm run test:release-smoke` against that container.
3. The single spec fails at `openOnboarding()` after 20 seconds.
4. In CI, open the run's `release-smoke` artifact. It has no
`docker-onboard-smoke.log`.
**Paperclip version or commit**
`2026.828.0-canary.3` (commit
|
||
|
|
7895f7f2b0 |
Install the declared Sentry server package into the hosted image (#12330)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip supports opt-in Sentry error monitoring for server and browser errors. > - The hosted image must include the server package when an operator sets SENTRY_DSN. > - The server package is an optional peer in the source tree, so the image did not include it. > - This pull request installs the declared server package in the hosted image and checks the result. > - The benefit is a hosted tenant can send server errors without a manual package install. ## Linked Issues or Issue Description No public issue exists for this change. **What happened?** The hosted image did not include the declared @sentry/node server package. A hosted tenant could set SENTRY_DSN, but the server could not load the package from the image. **Expected behavior** The hosted image must include the exact @sentry/node version from server/package.json. The self-hosted image must remain without this optional package. **Steps to reproduce** 1. Build or pull the hosted image. 2. Resolve @sentry/node from the server package path. 3. Compare its version with server/package.json. 4. Confirm that the tsx loader path still resolves. **Paperclip version or commit** Commit b6ff556a33ebdbe764b7f495951cd59009776608. **Deployment mode** Docker hosted image. ## What Changed - Add a cloud-server-deps Docker stage that installs the declared @sentry/node version in isolation. - Copy the isolated package into the cloud image without changing the production image. - Add a probe that checks the tsx loader and the resolved Sentry version. - Run the probe after the hosted image push in the Docker workflow. - Add server tests and update the observability documentation. ## Verification - Run `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/cloud-image-sentry.test.ts`. - Confirm that the changed test passes in CI. - Confirm that all pull request checks pass. - Note that the Docker workflow does not run for pull requests. It runs after a push to master, for configured tags, or after manual dispatch. ## Risks - Low risk. The production image body stays unchanged. - The cloud image adds the declared Sentry package and a small dependency tree. - The workflow probe fails if the image loses the tsx loader or resolves a different Sentry version. ## Model Used OpenAI GPT-5; exact model version supplied by the execution service; tool use and code execution; context window not specified. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ed1f51f75 |
fix(scripts): silence pnpm DEP0169 at provisioning install call sites (#12228)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Worktree provisioning prepares the dependencies and tools that these agents need. > - The pinned pnpm version calls the deprecated `url.parse()` function during each install. > - Node.js 24 reports this call as `DeprecationWarning [DEP0169]`. > - The provisioning scripts run more than one install, so the warning repeats in each run. > - This pull request disables only `DEP0169` at each affected pnpm install call site. > - The benefit is a clear provisioning log while other deprecation warnings remain visible. ## Linked Issues or Issue Description **What happened?** The worktree provisioning scripts printed `DeprecationWarning [DEP0169]` during each pnpm install. The warning came from pnpm 9.15.4 and its `toNerfDart` call to `url.parse()`. **Expected behavior** The provisioning scripts should hide this known warning from the pinned pnpm version. They should keep other deprecation warnings visible. **Steps to reproduce** 1. Use Node.js 24 with pnpm 9.15.4. 2. Run worktree provisioning with a base-workspace repair or dependency install. 3. Observe the repeated `DeprecationWarning [DEP0169]` output. **Paperclip version or commit** Commit `5cd41b1a9996713efdfdc62373da8045664c7f30`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific (core bug). **Database mode** Not database-related. ## What Changed - Add `--disable-warning=DEP0169` to each affected pnpm install call site. - Append the flag to `NODE_OPTIONS` so the scripts keep existing values. - Add comments that name the source of the warning and the removal condition. - Add a regression test for all affected scripts and warning codes. ## Verification - `bash -n scripts/provision-worktree.sh` passes. - `bash -n scripts/provision-worktree-runtime.sh` passes. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` passes with 15 tests. - GitHub Actions must pass all required checks before merge. ## Risks This change has low risk. It changes warning output only for `DEP0169`. It does not overwrite existing `NODE_OPTIONS` values. Revert commit `5cd41b1a9996713efdfdc62373da8045664c7f30` to restore the prior output. ## Model Used OpenAI Codex, GPT-5. The model used tool calls and code execution. The runtime did not expose a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6524d2b67f |
fix(dev): honor --data-dir isolation (#12193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The development runner starts the local API, UI, and embedded PostgreSQL services. > - Developers need separate data roots when they run more than one local checkout. > - The runner forwarded `--data-dir` to a child process that did not use it. > - The migration check and server therefore continued to use the default Paperclip home. > - This pull request applies the data root before worktree setup and migration checks. > - The benefit is an isolated database and state root for each requested development run. ## Linked Issues or Issue Description Refs #7466 ## What Changed - Parse and consume `--data-dir`, `--data-dir=<path>`, and `-d` in the development runner. - Set isolated default home, config, and context paths before worktree setup and migration checks. - Keep explicit config and context paths unchanged. - Include the normalized data root in the local service identity. - Let `dev:list` and `dev:stop` select the matching isolated service registry. - Keep explicit option environments independent of ambient process instance values. - Add regression tests and development documentation. ## Verification - `PAPERCLIP_INSTANCE_ID=ambient-test-instance pnpm exec vitest run server/src/__tests__/dev-runner-options.test.ts` passes with 8 tests. - `pnpm --filter @paperclipai/server typecheck` passes. - `pnpm --filter @paperclipai/adapter-utils build` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm dev:list --data-dir ./tmp/dev-service-review-fixture` selects the isolated registry. - A live run of `pnpm dev --data-dir ./tmp/data-dir-pr-smoke` became healthy on port 3101 while another checkout used port 3100. - The live run used `./tmp/data-dir-pr-smoke/instances/default/db` on a separate PostgreSQL port. - The latest-head Linux CI matrix passes, including build, typecheck, canary, all general and serialized test shards, and all e2e shards. - `pnpm test:run` was attempted on macOS. Current `master` has unrelated workspace path failures because `/tmp` resolves to `/private/tmp`. The focused regression suite passes, and the full Linux matrix is green. ## Risks - Risk is low. The change only affects development runs that pass `--data-dir` and matching service-management commands. - Explicit `PAPERCLIP_CONFIG` and `PAPERCLIP_CONTEXT` values still take priority. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol` produced this change. The run used tool-enabled reasoning and code execution. The runtime did not expose its context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5db8ce3c44 |
fix(docker): make tini PID 1 in the server image so adopted orphans are reaped (#12137)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runs execute inside the server container, and they spawn many
short-lived descendants: git, the adapter CLI, esbuild, sh
> - The server image sets `ENTRYPOINT ["docker-entrypoint.sh"]`, and
that entrypoint ends in `exec`, so node becomes PID 1
> - Node reaps only the children it spawned itself. It installs no
`SIGCHLD`/`waitpid` handler for orphans that the kernel re-parents onto
PID 1, so those orphans stay as zombies forever
> - Zombies accumulate monotonically. When the cgroup pid limit is
reached, every `fork()` in the container fails and the instance is dead
> - This pull request installs `tini` and makes it PID 1 in front of the
existing entrypoint, adds a behavioural test that proves reaping, and
adds a `pids_limit` backstop to both compose files
> - The benefit is that a long-running container no longer degrades into
total fork failure, and a future regression is caught by CI instead of
by an outage
Depends-on: none — this change is self-contained in the image build and
its tests, and it touches no other in-flight branch
## Linked Issues or Issue Description
No public GitHub issue exists for this defect. It was found on a live
long-running instance. Description follows the bug report template.
**What happened?**
The server container ran for 22 hours and reached 2039 of 2048 pids in
its cgroup. Of 1760 processes, 1731 were zombies, and all 1731 had PID 1
as their parent. PID 1 was `node --import
./server/node_modules/tsx/dist/loader.mjs server/dist/index.js`. Zombies
accrued at about 79 per hour and were never reaped. The oldest zombie
was 20.8 hours old against a container uptime of 22.0 hours, so nothing
had been reaped since boot. Once the pid limit was reached, `git` and
`gh` failed with `pthread_create failed: Resource temporarily
unavailable`.
**Expected behavior**
PID 1 reaps orphaned processes that the kernel re-parents onto it. The
pid count of a long-running container stays flat instead of growing
without bound.
**Steps to reproduce**
1. Start the server image without `docker run --init` and without `init:
true`.
2. Run agent work that spawns descendants which outlive their immediate
parent.
3. Read `/sys/fs/cgroup/pids.current` and count processes in `Z` state
over several hours.
4. The zombie count grows monotonically and every zombie has PPID 1.
**Relevant logs or output**
```
cgroup pids.current / pids.max : 2039 / 2048
total processes : 1760
zombies : 1731 (98.4%)
parent of every zombie : PID 1 (1731/1731)
PID 1 cmdline : node --import .../tsx/dist/loader.mjs server/dist/index.js
container uptime : 22.0 h
oldest zombie : 20.8 h median: 14.4 h
zombie names : git 717, claude 280, MainThread 167, sleep 141,
esbuild 138, postgres 76, sh 65, sccache 50
```
**Additional context**
The fix pattern is already in this repository.
`docker/agent-runtime/Dockerfile.base` installs `tini` and sets
`ENTRYPOINT ["/usr/bin/tini", "--"]`. It was never applied to the server
image.
## What Changed
- `Dockerfile`: install `tini` in the `base` stage and set `ENTRYPOINT
["/usr/bin/tini", "--", "docker-entrypoint.sh"]`. The entrypoint stays
in the exec chain, so UID/GID remapping, `gosu`, and graceful shutdown
are unchanged.
- `scripts/assert-orphan-reaping.sh` (new): a behavioural probe. It
spawns a leader that forks a grandchild, exits the leader, and asserts
that the orphaned grandchild leaves `Z` state instead of persisting. It
fails closed if the grandchild is not re-parented onto PID 1, so a pass
cannot mean the check ran too early.
- `.github/workflows/docker.yml`: run that probe against the pushed
image after the publish step. The publish step is multi-arch with `push:
true`, so nothing is loaded into the runner daemon and the pushed tag is
the only thing to test. The cloud variant is `FROM production` and
inherits the same `ENTRYPOINT`.
- `scripts/docker-build-test.sh`: run the same probe against a local
build.
- `docker/docker-compose.yml` and
`docker/docker-compose.quickstart.yml`: add `pids_limit: 2048` as a
backstop, so a future leak dies visibly at its own ceiling instead of
starving the host of pids.
- `server/src/__tests__/container-init-reaping.test.ts` (new): 13
assertions that guard the configuration the probe depends on.
No per-orchestrator init lever was added. The image owning PID 1 covers
compose, plain `docker run`, the quadlet units, and the ECS task
definition in one place. Adding `init: true` in compose or
`initProcessEnabled` on the ECS task would nest a second init around
`tini`, and `tini` then warns on every boot that it is not PID 1. The
new test asserts the absence of both levers across all three manifests,
so the decision survives the next edit.
## Verification
| Check | Result |
|---|---|
| `scripts/assert-orphan-reaping.sh` against a real init | Grandchild
re-parented to PPID 1, then reaped. Exit 0. |
| Same probe forced against a genuine zombie | Reports `Z` and fails.
The failure branch is not vacuous. |
| Config guard against the pre-fix files | Exactly the 3 relevant
assertions turn red. |
| Config guard with `tini` removed from `apt-get` but the comments kept
| Red. It checks the install, not a mention of the name. |
| `cd server && npx vitest run
src/__tests__/container-init-reaping.test.ts` | 13 passed |
| `npx tsc --noEmit -p server` | Clean |
| `node scripts/check-docker-deps-stage.mjs` | PASS |
| `node --test scripts/release-verify-workflow.test.mjs` | 8 passed |
Not verified locally: no container runtime is available in the authoring
environment, so the probe has not run against a build of this image. The
new `docker.yml` step runs it against the pushed image on this PR.
## Risks
Low risk, but it is an image and entrypoint change, so it affects
deployments.
- `tini` adds one small package to the `base` stage.
`docker/agent-runtime/Dockerfile.base` already installs it from the same
Debian archive.
- Signal handling changes shape: `tini` receives `SIGTERM` and forwards
it to the entrypoint, which `exec`s node. `tini` forwards signals to its
direct child by default, and the exec chain keeps node as that child, so
graceful shutdown is preserved. A reviewer should confirm this on a real
stop.
- `pids_limit: 2048` is new for compose users. A deployment that
legitimately needs more than 2048 processes would now hit the ceiling.
The measured steady state on a busy instance was under 400.
- If a deployment already passes `--init` or `init: true`, `tini` runs
under another init and prints a warning that it is not PID 1. Reaping
still works because the outer init handles it. The compose files in this
repository do not set `init: true`.
## Model Used
Claude Opus 5 (`claude-opus-5`), extended thinking, with tool use and
code execution in an agent harness.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local issues or links
- [x] My branch name describes the change and contains no internal
ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: zannis <1011451+zannis@users.noreply.github.com>
|
||
|
|
ffff1fe6e3 |
feat(runner): define package API and verification boundary (#12129)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package now has protocol, transport, provider, catalog, and authorization foundations. > - Its first upstream package boundary should expose only the implemented runtime and test-helper surfaces. > - Rust correctness belongs in the repository existing build verification, without introducing a parallel release process. > - Direct package creation must build the files declared by the package manifest. > - This pull request defines the minimal package API and verifies the optimized runner binaries in the existing PR and release Build jobs. > - The benefit is a production-ready runner package boundary with minimal build-process change. ## Linked Issues or Issue Description Refs #11962 This pull request replaces one bounded part of the archived large runner change. It follows the package-local authorization change in #12126. ## What Changed - Export only `@paperclipai/paperclip-runner` and `@paperclipai/paperclip-runner/testing`. - Keep Node-only fixture loading and semantic conformance helpers out of the runtime root. - Add a provider-neutral semantic conformance kit with stable JSON comparison and fail-closed input checks. - Keep deferred SDK, eval, browser, React, lab, and command surfaces private. - Pin the runner Rust toolchain to 1.97.1 with the minimal profile and `rustfmt`. - Run the Rust workspace tests in release mode. - Launch the optimized `paperclip-runnerd` and fake-harness binaries in process-level integration coverage. - Add one `pnpm --filter @paperclipai/paperclip-runner check:all` step to each existing PR and release Build job. - Make the existing server `prepack` lifecycle run its existing build after it prepares UI assets. - Document that no production adapter starts runnerd yet. This revision adds no standalone GitHub Actions job. It adds no server runner dependency or runner vendoring. It adds no Docker bootstrap or clean-consumer harness. It does not change `pnpm-lock.yaml`. ## Verification - `pnpm --filter @paperclipai/paperclip-runner check:all` - 66 TypeScript tests - 8 protocol contract tests - 56 Rust unit and integration tests - Release-mode integration coverage launches the optimized runnerd and fake-harness binaries. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/server-package-build-script.test.ts` (2 tests) - Clean `pnpm pack` from `server/` rebuilt the server and produced both `package/dist/index.js` and `package/dist/index.d.ts`. - `node --test scripts/__tests__/release-verify-workflow.test.mjs` (8 tests) - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `git diff --check` - No `pnpm-lock.yaml` diff. - The diff changes 12 files. ## Risks The runner adds Rust work to the existing Build jobs. These jobs can take longer on a cold cache. The pinned toolchain makes contributor and CI behavior reproducible. Cargo tests use `--release` to verify optimized executables. The server prepack lifecycle now performs the build that its published entry points require. This can make direct server packing slower. This pull request does not wire runnerd into the server. It does not select runnerd for any adapter. Existing application execution and finalization paths remain unchanged. ## Model Used OpenAI Codex with GPT-5. Agentic coding mode used repository tools, code execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0a01444514 |
test(release-smoke): cover the background-service leg of onboarding (#12151)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release pipeline gates each nightly and beta on a smoke suite that onboards the published npm artifact and drives the golden path > - That smoke runs onboarding inside a Docker container, and containers have no service manager, so the background-service leg of onboarding has zero automated coverage > - v2026.824.0 shipped a service install that crash-looped on a missing shim, and every smoke check stayed green (#12148 fixed the defect itself) > - This pull request adds a `smoke_service` job that runs the same published artifact directly on the runner VM's systemd and requires the installed service to end up serving > - The benefit is that a release with a broken service install can no longer pass the release smoke suite ## Linked Issues or Issue Description Refs #12148 — the fix for the defect this coverage gap let through. The gap: the release smoke runs `onboard` with `--yes` inside Docker, which both skips the service prompt and lacks systemd, so no CI job ever executed `manager.install()` against a real service manager. ## What Changed - New `scripts/service-onboard-smoke.sh`: onboards the published artifact with `--yes --install-service` on a systemd host, then fails unless the managed shim exists and is executable, `paperclipai.service` is active, and `/api/health` answers. A health response while the unit is not active also fails, because that is the signature of something other than the service serving. The script refuses to run over an existing managed install unless `SMOKE_FORCE=true`, and cleans up after itself by default so it is safe to run locally. - New `smoke_service` job in `.github/workflows/release-smoke.yml`: starts a user systemd session on the hosted runner (`loginctl enable-linger` + exported `XDG_RUNTIME_DIR`/`DBUS_SESSION_BUS_ADDRESS`), runs the script against `inputs.paperclip_version`, and uploads `systemctl status` + journal output as diagnostics. - No `release.yml` changes needed: `smoke_nightly` and `smoke_beta` call this reusable workflow, and a `workflow_call` result aggregates all jobs, so the new job gates nightly promotion automatically. ## Verification - `bash -n scripts/service-onboard-smoke.sh` passes and the workflow YAML parses. - End-to-end: dispatched this branch's Release Smoke workflow against the published canary that contains #12148; the `smoke_service` job onboards, installs the service, and verifies the service serves health. (Run link in PR comments.) - Negative case: the same assertions fail against v2026.824.0 — reproduced in a systemd container during the #12148 investigation: shim missing, unit in a 203/EXEC restart loop. ## Risks - Low risk to the product: no application code changes. - Pipeline risk: a flaky user-session setup on the hosted runner would block nightly promotion. Mitigated by validating the job end-to-end from this branch before merge, a 30-minute job timeout, and diagnostics uploaded on every run. - The service leg only covers systemd. launchd (macOS) still has no CI coverage; a macOS runner job is a possible follow-up. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended thinking, agentic tool use via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
890ab9acfe |
feat(release): thorough notes skeletons — nest each PR's summary at creation (#12124)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release workflow drafts the upcoming stable's notes skeleton the moment a beta publishes > - That skeleton was a bare list of commit subjects, so the notes only reached the shipped stable's depth after a later authoring pass during the soak > - Stable release notes are consistently verbose and thorough; the initial draft should start that way too > - This pull request nests each referenced PR's own summary under its subject line at creation time, and states the density bar in the authoring skill > - The benefit is a thorough raw document from day one of the soak, with no LLM tokens in Actions ## Linked Issues or Issue Description **What existing behavior does this improve?** The `draft_stable_notes` skeleton generated at beta publish (`scripts/draft-stable-notes.sh`). **Current behavior** The skeleton groups bare commit subjects by conventional-commit type. All substance arrives later, when a maintainer or agent rewrites it — reviewed maintainer feedback: stable notes are a lot more verbose, and the initial beta notes should be consistent with that. **Proposed behavior** Each subject that references a PR carries that PR's own summary nested beneath it — the PR template's "What Changed" bullets, else the first prose lines — fetched best-effort via `gh` and skipped silently when unavailable. The release-changelog skill now states the density bar explicitly: the beta-keyed draft ships verbatim as the stable's notes and is written at the previous stable's depth from the first pass. **Reason and benefit** The notes author starts from a thorough raw document instead of a commit list, and beta-time notes match the verbosity the stable will ship with. ## What Changed - `scripts/draft-stable-notes.sh`: `enrich_pr` nests PR summaries under subjects; best-effort (`gh` failure or `DRAFT_NOTES_SKIP_PR_ENRICHMENT=1` degrades to today's output); pipefail-safe when a "What Changed" section has no bullets. - `.github/workflows/release.yml`: the `draft_stable_notes` step gets `GH_TOKEN` so `gh` can read PR bodies. - `.agents/skills/release-changelog/SKILL.md`: "write at full stable depth from the first pass" guideline. - `scripts/draft-stable-notes.test.mjs`: three new tests — enrichment rendering via a fake `gh`, silent degradation without one, and the sparse-body case that previously killed the script under `set -o pipefail`. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 11 pass. - Live run against the real repository for the current beta (`2026.818.0-beta.1`, 172 commits): exit 0, 439 nested summary lines; spot-checked entries carry the correct PRs' What Changed bullets. - `bash -n` on the script; `release.yml` re-parsed as YAML. ## Risks - Low: the publish path is untouched; enrichment is read-only `gh` calls in the post-publish draft job and degrades to the current skeleton on any failure. Roughly one API call per commit in the range (~170 today) — well inside the token's rate budget, adds a couple of minutes to a job with a 10-minute timeout. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0dfa0fb988 |
ci: refresh general-server shard duration manifest (#12075)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its slowest check sets the feedback time for all contributors > - The general-server test lane splits its vitest suites across five runners with a duration-weighted partition (`scripts/general-server-shard.mjs`) > - The partition reads a duration manifest that was sampled on 2026-08-04, when the lane had 279 suites and 946s of serial time > - The lane has since grown to 405 suites and 1274s; 126 suites had no recorded duration and one suite grew from 37s to 123s > - The stale weights made the partition uneven: in the fully green actions run 32708351172, "General tests (server (2/5))" ran 364s and was the slowest check in the whole run, while sibling shards ran 292-330s > - This pull request refreshes the manifest with per-suite durations measured from that same run > - The benefit is a level five-shard split (255s ±1s of predicted suite time per shard), which removes ~50s from the slowest PR check ## Linked Issues or Issue Description **Describe the current behavior** In the fully green PR actions run [32708351172](https://github.com/paperclipai/paperclip/actions/runs/32708351172) (2026-08-24), the check "General tests (server (2/5))" completed in 364s. Its test step ran 315s while sibling shards ran 241-276s. It was the slowest check in the run. **Describe the improvement** The duration manifest `scripts/general-server-shard-durations.json` is stale. It holds 279 suites sampled on 2026-08-04, but the lane now has 405 suites. The 126 unknown suites fall back to the median weight (~1.3s), and `server/src/__tests__/workspace-runtime.test.ts` grew from 37.4s to 123.3s. The partition therefore predicts a level split but produces an uneven one. Refreshing the manifest restores the level split without any code change. **Expected impact** All five server shards level at ~255s of predicted suite time (~310s job time). The slowest PR check drops from 364s to about 317s, so the PR critical path improves by roughly 50s. ## What Changed - Regenerated `scripts/general-server-shard-durations.json` from actions run 32708351172 (2026-08-24): 405 suites, 1274s total serial time (was 279 suites, 946s from 2026-08-04) - Updated the `$comment` field to name the new sample run and date - No code changes; the partition logic in `scripts/general-server-shard.mjs` is untouched ## Verification - Parsed all five "General tests (server (n/5))" job logs from run 32708351172 with the consecutive-completion-timestamp method described in the manifest `$comment`; asserted that the parsed suite set equals the exact file list that `run-vitest-stable.mjs` collects (405/405, no misses, no extras) - Ran `node scripts/run-vitest-stable.mjs --mode general --group general-server --shard-index N --shard-count 5 --dry-run` for N=0..4 with the new manifest: each shard predicts 255s (±1s) of suite time, and the five shards form a complete, non-overlapping cover of all 405 suites - Ran `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs`: 13/13 pass ## Risks - Low risk. The change is data-only. Wrong weights cannot break correctness: the partition always covers every suite exactly once, so the worst case of a bad weight is an uneven shard, which is the current state. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable (no code change; existing partition tests pass) - [x] I have updated relevant documentation to reflect my changes (manifest `$comment` updated) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #11528 (balanced the serialized server shards by recorded duration), #10923 (split serialized tests into five shards), #11156 (split workspaces-a into two shards). Co-authored-by: Claude <noreply@paperclip.ing> |
||
|
|
633e102971 |
fix: verify issue-update writes instead of inferring success (#12051)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents report task state to the control plane with `PATCH
/api/issues/{id}` at the end of each heartbeat
> - On remote sandbox targets those writes cross a relay that can fail
at the connection level
> - An agent that pipes its status curl through `head` cannot see that
failure; the write is lost but the run reports success
> - The issue then stays `in_progress` with no disposition, and the
missing-disposition recovery must repair it
> - This pull request makes the issue-update helper verify every write,
and it teaches the shared skill to require verified writes
> - The benefit is that a lost status write becomes a visible, retried
failure instead of a silent success
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
A sandboxed heartbeat run answered its issue in a comment. It then sent
`PATCH /api/issues/{id}` with `status: done` through `curl -sf ... |
head -c 400`. The relay dropped the connection. The `-f` flag suppressed
the error output, and the pipe replaced curl's exit code with the exit
code of `head`. The agent saw empty output and exit 0. It reported the
write as an "empty 2xx" success and exited. The issue stayed
`in_progress`, and the successful-run recovery had to close it in a
corrective run.
**Expected behavior**
A status write that does not reach the server must surface as a failure.
The helper script must retry transient failures. It must exit non-zero
when the write is unconfirmed. Skill guidance must forbid write patterns
that hide failures.
**Steps to reproduce**
1. Point `PAPERCLIP_API_URL` at an endpoint that drops connections
intermittently.
2. Finalize an issue with `curl -sf -X PATCH
"$PAPERCLIP_API_URL/api/issues/$ID" -d '{"status":"done"}' | head -c
400`.
3. Observe exit code 0 with empty output while the server never received
the PATCH.
## What Changed
- `scripts/paperclip-issue-update.sh` now captures `%{http_code}`,
retries a retryable failure (connection-level, 429, 5xx) once — two
attempts total, which matches the shared bounded-write-retry rule —
rejects an empty 2xx body, and confirms the response echoes the
requested status before it exits 0.
- Failure output states plainly that the write was NOT saved, so the
calling agent reports it accurately.
- `skills/paperclip/SKILL.md` Step 8 adds a required "Verify writes —
never infer them" rule: a successful PATCH always returns the updated
issue JSON, disposition writes must never run through `head`/`tail`
pipelines, and an unconfirmed write must be reported as FAILED.
- `server/src/__tests__/paperclip-skill-utils.test.ts` pins the new
skill rule; a new `paperclip-issue-update-helper.test.ts` exercises the
helper's behavior end-to-end.
## Verification
- `bash -n scripts/paperclip-issue-update.sh`
- `server/src/__tests__/paperclip-issue-update-helper.test.ts` runs the
helper end-to-end against a local HTTP server: confirmed-echo success
(exit 0), empty 2xx (exit 1), wrong echoed status (exit 1), 422 reject
(exit 1, exactly one request), 503 then success (two requests),
connection refused (two attempts, then exit 1 with a "NOT saved"
report).
- `npx vitest run
server/src/__tests__/paperclip-issue-update-helper.test.ts
server/src/__tests__/paperclip-skill-utils.test.ts
server/src/__tests__/cli-invocation-safety.test.ts` — 50 passed.
## Risks
- Low risk. The success-path output is unchanged (the updated issue
JSON).
- The helper now exits non-zero on unconfirmed writes. Callers that
previously missed silent failures now see explicit errors. That is the
intended behavior change.
- The single retry re-sends the PATCH after a retryable failure. If the
first request committed and only its response was lost, an attached
comment can post twice. The duplicate is visible and benign; the prior
behavior lost the write silently.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
dc5b070709 |
fix(runtime): guard empty Bash 3.2 array expansion (#11891)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in isolated worktrees with a separate Paperclip runtime. > - Runtime provisioning uses a Bash script on macOS hosts. > - macOS ships Bash 3.2, where an empty array expansion fails under `set -u`. > - The source-config argument array is empty when the base workspace already has a config. > - This pull request guards that expansion and tests the normal base-config path on Bash 3.2. > - The benefit is that managed worktree provisioning no longer fails before database seeding. ## Linked Issues or Issue Description No public GitHub issue exists for this problem. PR #11752 added the conditional source-config argument that exposed the failure. **What happened?** `scripts/provision-worktree-runtime.sh` expands an empty `source_config_args` array while `set -u` is active. Bash 3.2 reports `source_config_args[@]: unbound variable` and stops provisioning when the registered base workspace already has `.paperclip/config.json`. **Expected behavior** Runtime provisioning must call `worktree ensure-seeded` without a source override when the base workspace config exists. It must work with the Bash 3.2 version that macOS supplies. **Steps to reproduce** 1. Use macOS system Bash 3.2. 2. Create a base workspace with `.paperclip/config.json`. 3. Run `scripts/provision-worktree-runtime.sh` with `set -u` active in the script. 4. Observe the unbound-variable error before `worktree ensure-seeded` runs. **Paperclip version or commit** Reproduced on `origin/master` before this change. **Deployment mode** Local managed worktree runtime on macOS. ## What Changed - Guard all three optional source-config array expansions with Bash 3.2-compatible parameter expansion. - Add a regression test that uses the base-config path and verifies that no `--from-config` argument is sent. - Document the Bash 3.2 compatibility requirement in the runtime script. ## Verification - `/bin/bash -n scripts/provision-worktree-runtime.sh` - `node --test --test-name-pattern='runtime provisioning invokes ensure-seeded once|runtime provisioning omits the source override|runtime provisioning guards every optional source-config expansion' scripts/__tests__/provision-worktree-self-heal.test.mjs` - `git diff --check` ## Risks Low risk. The change only affects expansion of an optional two-element CLI argument array. The regression tests cover both the empty and non-empty paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model ID `gpt-5`. The runtime did not expose the context-window size. Reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
417336f8be |
fix(workspaces): attach PR preparation to existing branches (#11703)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Execution workspaces isolate an agent task from the primary checkout. > - Pull request preparation can need a branch that already contains completed work. > - The workspace policy could not require an exact existing branch. > - Workspace cleanup also treated worktree creation as branch ownership. > - This pull request adds an exact existing-branch policy and separate branch ownership metadata. > - The benefit is safe pull request preparation that preserves every existing commit and operator-owned branch. ## Linked Issues or Issue Description **What happened?** A pull request preparation run could not pin its execution workspace to an exact existing branch. Workspace reuse and cleanup could also confuse worktree creation with branch ownership. **Expected behavior** The run must attach only to the requested branch in an isolated Git worktree. It must fail if the branch is missing, busy, or inconsistent. Cleanup must not delete a branch that Paperclip does not own. **Steps to reproduce** 1. Create a branch that contains completed work. 2. Configure a pull request preparation task to use that branch. 3. Start the task and observe that the prior policy cannot require the exact branch. **Paperclip version or commit** This behavior reproduces on the base revision before this pull request. **Deployment mode** Local development with isolated Git worktrees. ## What Changed - Add `existingBranch` to the execution workspace policy and shared validation contracts. - Require `existingBranch` to use an isolated Git worktree and reject conflicting branch templates. - Attach to the exact branch without creating, renaming, resetting, or deleting it. - Track branch ownership separately from worktree creation and use that ownership during cleanup. - Return HTTP 422 for invalid existing-branch settings on every issue-producing route. - Add a bounded repair script for existing pull request preparation tasks. - Add focused policy, route, heartbeat, runtime, and ready-comment tests. - Document the exact-branch behavior and safety rules. ## Verification - `pnpm exec vitest run server/src/__tests__/execution-workspace-policy.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/issue-existing-branch-validation-status.test.ts server/src/__tests__/workspace-runtime.test.ts server/src/services/workspace-runtime-exposure.test.ts server/src/services/workspace-runtime-ready-comment.test.ts` passed 335 tests. - `pnpm -r typecheck` passed for all workspace projects. - `pnpm test:run` passed 4,431 tests. Two unrelated embedded-Postgres setup hooks timed out under aggregate load. Their isolated rerun passed 74 tests. - `pnpm build` passed for all workspace projects. - The two review regressions passed with 139 unrelated tests skipped. - All latest-head CI gates passed after one unrelated timing-sensitive test passed on rerun. - Greptile scored the latest head 5/5 with no unresolved review threads. ## Risks - Invalid workspace settings now return HTTP 422 instead of a generic validation response. - The exact branch must already exist and must not be checked out by another worktree. - The new policy fails closed when it cannot prove branch identity or ownership. - This change has no database migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family assisted with this change. The runtime did not expose its exact deployment ID or context window. The agent used high-reasoning mode, repository tools, shell execution, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
233c12f029 |
feat: add kimi-local adapter for Kimi Code CLI (CLI + ACP engines) (#9967)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local agent adapters (`claude_local`, `gemini_local`, `grok_local`, …) are the integration surface that lets Paperclip run coding CLIs on the host machine > - The Kimi Code CLI (`kimi`, Moonshot AI) has a documented non-interactive mode, `kimi -p --output-format stream-json` with session resume via `kimi -r`, but Paperclip has no built-in adapter for it > - So Kimi users (especially Kimi membership / OAuth subscribers) cannot onboard their CLI to Paperclip agent teams > - This pull request adds a complete built-in `kimi_local` adapter (both execution engines, session management, instructions + skills delivery, thinking-effort control, environment test, UI and CLI modules, docs) following the established `gemini_local`/`grok_local` package pattern > - Kimi Code ships an ACP server (`kimi acp`), so the adapter runs on Paperclip's shared acpx engine by default (streaming transcript with live tool status, like `claude_local`/`gemini_local`) and falls back to a headless CLI lane (`kimi -p --output-format stream-json`) when ACP prerequisites are unavailable > - The benefit is that Kimi Code becomes a first-class Paperclip agent lane: selectable in the UI, resumable across heartbeats, with the same operating context (instruction bundle, skills, effort) and streaming transcript the other local adapters get ## Linked Issues or Issue Description - Supersedes #9880 (same branch; expanded from the CLI-only lane into a complete adapter with the default ACP engine lane, control-plane skill install, and live transcript wiring) - Refs #9879 (adapter request for Kimi Code CLI, filed with this PR) - Refs #163 (original Kimi support request) Duplicate/related prior PRs, per the dedup search (both appear stale: no updates or maintainer review since May 2026, and both target an older Kimi CLI interface; calling them out for reviewer context per CONTRIBUTING.md): - Refs #6276 (`feat: add kimi-local adapter`): targets an older array-based content format (`{type: think}`/`{type: text}` blocks), not the current documented stream-json schema - Refs #5202 (`feat(adapter): add Kimi CLI local adapter with Wire protocol support`): builds on a `--wire` JSON-RPC interface that current Kimi Code CLI (0.27.0) no longer documents; the current documented headless interface is `-p --output-format stream-json` This PR is a fresh implementation against current master and the currently documented/verified Kimi CLI behavior (see Verification). Happy to fold in anything useful from the earlier attempts if a reviewer prefers. ## What Changed - **New adapter package** `packages/adapters/kimi-local` (`@paperclipai/adapter-kimi-local`), modeled on `gemini-local`/`grok-local`: - `src/server/execute.ts`: spawns `kimi -p <prompt> --output-format stream-json` (argv array, no shell), `-m <model>` only when configured, `-r <sessionId>` when the stored session cwd matches the run cwd, automatic fresh-session retry on unrecoverable-session errors, headless-safe env (`CI=1`, `NO_COLOR=1`, `KIMI_CODE_NO_AUTO_UPDATE=1`, `TERM=dumb`; user-configured values win), full remote (ssh/sandbox) execution lane with runtime install via `@moonshot-ai/kimi-code` - **Instruction bundle delivery**: the prompt path directive now names the sibling instruction files (`./HEARTBEAT.md`, `./SOUL.md`, `./TOOLS.md`) alongside the prepended entry file, and local runs pass `--add-dir <instructions-dir>` so Kimi can actually open them (matching `claude_local`). Without this, only the entry file reached Kimi and agents improvised the operating workflow that `HEARTBEAT.md` documents - **Thinking effort**: a configured `effort` is forwarded as the `KIMI_MODEL_THINKING_EFFORT` operational override (Kimi has no per-invocation effort flag). It is only sent for models that advertise `support_efforts` (currently `kimi-code/k3`) to avoid provider rejections, and `medium` maps to `high` since Kimi has no medium tier (`low`/`high`/`max` pass through) - **Skills delivery**: desired Paperclip skills are delivered via Kimi's `--skills-dir` flag from a dedicated per-run directory (a local snapshot, or the synced snapshot on remote targets), so skills load reliably and in isolation. Paperclip never overwrites the shared `$KIMI_CODE_HOME/skills` home, so skills installed by the operator or other agents are left intact. `--skills-dir` is only passed when at least one skill is desired, so unconfigured agents keep Kimi's default skill discovery - **Live run status**: the adapter now forwards each streamed stream-json line to `onEvent` (assistant `content` as an assistant snippet, `tool_calls` as tool-name events), which drives the issue-thread activity indicator (`currentToolName` / `lastAssistantSnippet` / `lastEventAt`). Previously the adapter only wrote the raw run log, so the issue thread showed a stale "no output for N s" line with no tool or reasoning context while Kimi worked. Tool results are omitted so the last meaningful "Using X" / snippet is not overwritten by a generic label - `src/server/parse.ts`: parses the verified Kimi stream-json event shapes (`assistant` text, `assistant.tool_calls` with JSON-string arguments, `tool` results, trailing `meta.session.resume_hint` for session-id capture) plus failure classifiers (`kimi_auth_required`, transient network, unrecoverable session). A signaled exit (null exit code, not a timeout) is now reported as a failure rather than coalesced to success, and the error message names the terminating signal - `src/server/skills.ts`: lists/syncs Paperclip skills for the adapter's skill-management surface - `src/server/test.ts`: environment test covering CLI resolution + `kimi --version`, cwd check, auth detection (OAuth credential dirs, keyed `[providers.*]` in config.toml, or the `KIMI_MODEL_NAME` + `KIMI_MODEL_API_KEY` env pair), and a live hello probe - `src/ui/` (stdout-line parser for transcripts, config builder) and `src/cli/` (stream event formatter) modules - Root metadata: three managed model aliases (`kimi-code/kimi-for-coding`, `kimi-code/kimi-for-coding-highspeed`, `kimi-code/k3`), effort-capable-model metadata (`EFFORT_CAPABLE_MODELS`, effort mapping helpers), `agentConfigurationDoc` - Tests: 101 tests across parse, execute (args building, resume gating, retry, auth error code, timeout, signaled-exit failure, effort forwarding/gating/mapping, `--add-dir` instructions directive, `--skills-dir` gating, `onEvent` runtime-event forwarding), ACP engine (engine resolution, acpx config build, node-version gate), ACP transcript delegation, environment test, UI parse/build-config - **ACP engine lane (default)** (`src/server/acp.ts` + shared `adapter-utils/acpx-engine`): Kimi Code ships an ACP server (`kimi acp`), so `kimi_local` now runs on Paperclip's shared acpx engine by default, matching `claude_local`/`codex_local`/`gemini_local`. The issue-thread transcript streams live (assistant text deltas, tool calls with a `pending`->`completed` status lifecycle) instead of the CLI lane's bursty complete-message output. Registered `kimi_local -> "kimi"` in `ACPX_ADAPTER_AGENT_IDS` and resolved the built-in agent command to `kimi acp`; `execute.ts` dispatches to the ACP executor first with an automatic CLI fallback when ACP prerequisites fail (`engine=acp` requires ACP, `engine=cli` pins the headless lane); `index.ts` falls back to the shared acpx session codec; the UI/CLI delegate `acpx.*` events to the shared acpx transcript parser and event formatter. The headless CLI lane (above) remains as the fallback - **Registration** (one entry each, mirroring existing adapters): server adapter registry + `BUILTIN_ADAPTER_TYPES`, `AGENT_ADAPTER_TYPES` (shared), UI adapter registry + display registry (`Kimi Code`, Moon icon) + capabilities defaults, CLI adapter registry, `Dockerfile` (package copy + `npm install --global @moonshot-ai/kimi-code@latest`), `vitest.config.ts` workspace, `scripts/release-package-manifest.json` - **Behavioral sets** mirroring `gemini_local` (Kimi resumes sessions the same way): `GIT_SENSITIVE_LOCAL_ADAPTER_TYPES`, `SESSIONED_LOCAL_ADAPTERS` (heartbeat + recovery), `REMOTE_MANAGED_ADAPTERS`, ssh/sandbox execution-target allow-lists, `ADAPTER_DEFAULT_RULES_BY_TYPE` (`timeoutSec: 0`, `graceSec: 15`), and `LEGACY_SESSIONED_ADAPTER_TYPES` + `ADAPTER_SESSION_MANAGEMENT` in adapter-utils - **UI touch-points**: New Agent default-model branch, AgentConfigForm command map (`kimi_local: "kimi"`) + model defaults + a Kimi-specific thinking-effort option list (`Low`/`High`/`Max`, reflecting Kimi's tiers rather than borrowing Claude's), OnboardingWizard (command map, model default, `kimi login` / `KIMI_MODEL_NAME + KIMI_MODEL_API_KEY` auth hints, manual-debug command line), InviteLanding enabled adapters - **Control-plane skill install** (`cli/src/commands/client/agent.ts`): `paperclipai agent local-cli` seeded the Paperclip control-plane skills into `~/.codex/skills` and `~/.claude/skills` so Codex/Claude agents auto-discover the API reference every run. Kimi had no equivalent target, so `kimi_local` agents began each session without the control-plane skill and rediscovered routes (e.g. the company-scoped `POST /api/companies/{companyId}/issues`) by trial and error. Added `~/.kimi-code/skills` (honoring `KIMI_CODE_HOME`) as a third install target for parity. Independent of the per-run `--skills-dir` delivery, which only applies to explicitly configured skills. - **Docs**: `docs/adapters/kimi-local.md` (prerequisites, auth options, config fields including `effort`, session resume, instruction bundle, skills delivery, control-plane skill install) + a row in `docs/adapters/overview.md` Out of scope (deliberately): model profiles, built-in agent `allowedAdapterTypes` additions. ## Verification\n\nCurrent-master rebase verification (OpenAI Codex, 2026-08-03): 13 focused files / 231 tests pass; adapter-utils, server, UI, CLI, and Kimi adapter typechecks pass; full repository build and UI token gates pass. The branch is conflict-free against master at head `1249df117c5e12e5771b9a570a6340866450619e`.\n\nAutomated (all from repo root, pnpm 9.15.4, Node 22): - `vitest run packages/adapters/kimi-local`: 89/89 pass (includes coverage for the instruction `--add-dir` directive, effort forwarding/gating/mapping, `--skills-dir` gating, the signaled-exit failure path, and `onEvent` runtime-event forwarding with cross-chunk line buffering) - `vitest run server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/services/heartbeat-stop-metadata.test.ts ui/src/adapters/adapter-display-registry.test.ts`: 37/37 pass - `vitest run cli/src/__tests__/skills.test.ts`: 13/13 pass (the control-plane skill install target follows the existing Codex/Claude install path, whose symlink logic is unchanged) - `vitest run packages/shared`: 307/307 pass; `vitest run packages/adapter-utils`: pass except one pre-existing, unrelated failure (`mcp-isolation.integration.test.ts` requires Claude CLI ≥ 2.1.207; host has 2.1.185, fails identically on unmodified master) - `pnpm --filter @paperclipai/adapter-kimi-local typecheck|build`, plus typecheck of `server`, `ui`, `cli`, `adapter-utils`: all clean - `pnpm install --frozen-lockfile`: passes (the PR diff itself contains no lockfile changes, per repo policy; verified against a locally regenerated lockfile) - `node scripts/check-no-git-push.mjs` and `node scripts/check-forbidden-tokens.mjs`: pass - CI note: the `policy` job's release-bootstrap step is expected to stay red until a maintainer bootstraps the first npm publish of `@paperclipai/adapter-kimi-local`; see the CI Note for Maintainers comment. All other contributor-actionable checks are green. Manual end-to-end (real Kimi CLI 0.27.0, OAuth login, dev server on an isolated instance): 1. Server `GET /api/adapters` lists `kimi_local` as builtin with correct capability flags; models endpoint returns the three Kimi models 2. `POST .../adapters/kimi_local/test-environment`: all checks pass, including a live `kimi -p` hello probe 3. Created a `kimi_local` agent and invoked two heartbeats: run 1 spawned `kimi -p ... --output-format stream-json`, Kimi used its `Read` tool, produced the expected answer, and the session id was captured from the `session.resume_hint` meta event; run 2 resumed the **same** Kimi session (`sessionIdBefore == sessionIdAfter`) via `-r` 4. UI: adapter appears in the New Agent dropdown; selecting it shows the Kimi command placeholder, the three models, and the Kimi config fields; the run transcript renders Kimi tool calls via the adapter's stdout parser The instruction-bundle, thinking-effort, and `--skills-dir` changes landed after the manual run above. They are covered by the unit tests listed under Automated, and the Kimi CLI flags they rely on (`--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were confirmed against the installed Kimi Code CLI 0.27.0 (`kimi --help`, config-file thinking-effort docs). Screenshots (assets branch on the fork, not part of the diff):       ## Risks - Low risk to existing behavior: the change is additive, one new workspace package plus single-entry registrations alongside existing adapters; no existing adapter code paths are modified. - The adapter invokes the locally installed `kimi` CLI; like other local adapters, run behavior depends on the host's Kimi version. The parser is written against the documented/verified 0.27.0 stream-json schema and degrades gracefully (malformed lines are skipped, failures surface as run errors). - `--skills-dir` overrides Kimi's auto-discovery of user and project skills for the run. This is intentional (paperclip-managed agents get a reproducible, isolated skill set), and it is only passed when at least one Paperclip skill is desired, so unconfigured agents keep default discovery. - Thinking effort is only forwarded to models that advertise `support_efforts` (currently `kimi-code/k3`); `EFFORT_CAPABLE_MODELS` must be extended when more Kimi models gain support, otherwise a configured effort is silently ignored for them. - `Dockerfile` now installs `@moonshot-ai/kimi-code@latest` globally alongside the other agent CLIs, so image size increases slightly. - Maintainer action needed for the npm bootstrap gate: the `policy` job's release-bootstrap step fails until the first npm publish of `@paperclipai/adapter-kimi-local` (the gate from #5146 that every new adapter package has passed through). Enrollment with `publishFromCi: true` is required by the manifest validator (dropping the entry, `false`, or `private` are all rejected), so this is intentionally left to a maintainer. Remaining CI lanes are expected to run once it is done. ## Model Used\n\n- **Current-master rebase, conflict adaptation, and registry-parity coverage:** OpenAI, **GPT-5 Codex** (Codex agent; exact serving model ID and context-window size were not exposed to the runtime), with repository, shell, Git, and GitHub tooling. It preserved Hawik’s commit authorship, reconciled ACPX and environment-capability changes, added current registry tests, and ran the verification above.\n- **Adapter implementation and initial review:** Moonshot AI, **Kimi K3 Coding** (latest), via **Kimi Code CLI v0.27.0** (`kimi-code/k3` alias, 1M-token context window, thinking mode, agentic tool use). The CLI agent explored the repo, wrote the adapter implementation (delegated to a coder sub-agent of the same model), ran tests, and drafted the first version of this PR body. A second model-driven review pass (read-only, same model) audited the diff for security/correctness before submission; its findings (shell-quoting hardening, auth-detection false positive, session-compaction registration, test gaps) were fixed and are included. - **Harness-context fixes and review responses:** Anthropic, **Claude Opus 4.8** (`claude-opus-4-8`) via Claude Code. Diagnosed from run logs that Kimi received only the entry instructions file (not the `HEARTBEAT.md`/`SOUL.md`/`TOOLS.md` bundle) and that `effort` was never wired, then implemented the instruction `--add-dir` delivery, `KIMI_MODEL_THINKING_EFFORT` forwarding, and `--skills-dir` skill delivery, added the accompanying tests and docs, and addressed the automated review comments (preserving external skills on remote sync, treating a signaled exit as a failure). Also extended the `paperclipai agent local-cli` installer to seed the control-plane skills into `~/.kimi-code/skills` for Codex/Claude parity, wired `onEvent` runtime events so the issue-thread activity indicator reflects Kimi's tool and reasoning output live, and built the ACP engine lane (`kimi acp` via the shared acpx engine, default) so the transcript streams with live tool status like the other ACP adapters. The Kimi CLI flags, subcommand, and env var relied on here were verified against the installed Kimi Code CLI 0.27.0. - All CLI behaviors claimed here (`-p`, `--output-format stream-json`, `-r` resume, event shapes, `--add-dir`, `--skills-dir`, `KIMI_MODEL_THINKING_EFFORT`) were verified empirically against the installed Kimi CLI, not assumed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green *(only the release-bootstrap step remains red, pending the maintainer npm publish described in Risks)* - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(will address all Greptile comments as they arrive)* - [x] I will address all Greptile and reviewer comments before requesting merge --- ## Maintainer Addendum (2026-08-20) The shared acpx-engine and issue-chat changes (run-summary segmentation, placeholder tool-event coalescing, `ISSUE_CHAT_TRANSCRIPT_MAX_VISIBLE_ENTRIES` 30 → 400, live-reasoning UI) have been **extracted to #11761** so the cross-adapter behavior changes review and revert independently — both commits there preserve @hawikk's authorship. This PR is now the kimi-specific adapter only (60 files, +3,793/−8, essentially pure addition); the only shared-engine touch left is the `kimi acp` command resolution. `publishFromCi` is `true` — the package name is bootstrapped on npm. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Devin Foley <devin@paperclip.ing> |
||
|
|
a9d1f740f0 |
fix(workspaces): seed managed worktrees when the base checkout has no config (#11752)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents do that work in isolated git worktrees, and a managed
worktree runs its own Paperclip instance with a cloned database
> - That clone needs a seed source, and the source must come from
server-owned registration, never from state the workspace itself can
rewrite
> - The seed-source resolver requires the registered base project
workspace to hold its own `.paperclip/config.json`
> - A managed project workspace is a plain `git clone`, and no code
writes that file into it
> - Every isolated worktree provision, deferred seed, and workspace
repair therefore fails on a managed checkout
> - This pull request lets a named source supply the config when the
base checkout has none
> - The benefit is that managed worktrees provision again, and the seed
source stays server-owned
## Linked Issues or Issue Description
No public GitHub issue exists for this problem. It is described below.
**What happened?**
Agent runs that need an isolated worktree fail during provisioning. The
provision command exits with this error (paths redacted):
```
Execution workspace provision command "bash ./scripts/provision-worktree.sh" failed:
Registered base project workspace has no canonical Paperclip config:
<instance-home>/instances/default/projects/<company-id>/<project-id>/<repo>/.paperclip/config.json
```
`resolveRegisteredWorktreeSeedSource` sets `registeredConfigPath` to
`<baseCwd>/.paperclip/config.json` whenever the caller names a
registered base workspace. It then requires that file to exist.
`scripts/provision-worktree.sh` applies the same rule.
A managed project workspace never has that file.
`materializeManagedProjectWorkspace` creates it with `git clone` and a
rename, so the checkout holds repository content only. The control plane
keeps its config at `<home>/instances/<id>/config.json` instead.
The failure reaches three paths: worktree provisioning, deferred seeding
through `worktree ensure-seeded`, and workspace repair.
The behavior changed in #11671. That pull request replaced a fallback
chain with a single hard requirement. Fixture code in
`scripts/__tests__/provision-worktree-self-heal.test.mjs` writes a
config into the fake base workspace, so tests kept passing.
**Expected behavior**
A managed worktree provisions and seeds from the registered source. The
seed manifest still never selects that source.
**Steps to reproduce**
1. Register the Paperclip repository as a project with a `repoUrl`, so
the server materializes a managed checkout.
2. Assign an issue to an agent whose workspace strategy is
`git_worktree`.
3. Watch the workspace operation log for the provision command.
4. The command exits non-zero with the error above.
**Paperclip version or commit**
Reproduced on `master` at
|
||
|
|
b5a3a863c3 |
feat(release): bootstrap new npm packages with a placeholder publish (#11757)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its release pipeline publishes a set of npm packages from CI with npm trusted publishing (GitHub OIDC), gated by `scripts/release-package-manifest.json` > - A brand-new package name cannot be published by CI directly: the PR bootstrap gate requires the name to resolve on npm, and a trusted-publisher rule can only be configured after the package page exists > - The current bootstrap helper closes that gap by building the package locally and publishing its real output from a maintainer machine — before the PR that adds the package has passed CI or review > - This pull request replaces that flow: the helper now publishes a minimal deprecated placeholder at version `0.0.0` that only reserves the name, so every real version ships from CI > - The benefit is that unreviewed build output never reaches npm, and the bootstrap runs from any checkout (including `master`, before the new package's PR merges) with no local build ## Linked Issues or Issue Description **What existing behavior does this improve?** The one-time npm bootstrap for a brand-new release package (`pnpm run release:bootstrap-package`). **Current behavior** The helper builds the target package locally and publishes the real build output from a maintainer machine. That content has not passed repository CI or review at publish time. The helper also requires the new package to exist in the local workspace, so it must run from the (unmerged) PR branch that adds the package. **Proposed behavior** The helper publishes a three-file placeholder at version `0.0.0` (manifest, README, and an `index.js` that throws a descriptive error), waits for the registry to show the package, then deprecates it. The PR bootstrap gate (`scripts/check-release-package-bootstrap.mjs`) only requires the name to resolve on the registry, so the placeholder satisfies it. The first real calver release from CI supersedes the placeholder, and a stable release moves `latest` off it — the same `latest` window that existed under the old flow, but containing an explicit inert stub instead of unreviewed code. **Reason and benefit** Real package content only ever reaches npm from CI, after review and merge. The bootstrap becomes safer (scope guard refuses names outside `@paperclipai/`, already-published names are rejected) and simpler (no local build, no workspace state, runs from any checkout). **Breaking changes** None at runtime. The helper's CLI surface changes: it now takes a package name only (no directory selector) and drops `--skip-build`. `doc/PUBLISHING.md` is updated to match. ## What Changed - `scripts/bootstrap-npm-package.mjs`: replaced the build-and-publish flow with a placeholder publish — stages `package.json` + `README.md` + throwing `index.js` at version `0.0.0` in a temp directory, previews with `npm publish --dry-run`, and publishes only with `--publish`. One-time passwords are prompted interactively (never passed as arguments, since they are single-use and would land in shell history), with re-prompt on a rejected or expired code. After publishing, the helper polls the registry until the package is visible (a first publish can lag by minutes; verified live at ~5 minutes), requiring two consecutive sightings before prompting for a second code and deprecating the placeholder so accidental installs warn loudly; on timeout or failure it prints the exact manual `npm deprecate` command. Added an `@paperclipai/`-scope guard and a fail-fast error when `--publish` runs without an interactive terminal. Removed the workspace-plan dependency so it runs from any checkout. - `scripts/bootstrap-npm-package.test.mjs`: rewrote for the new interface — argument parsing, scope validation, the generated placeholder files (manifest shape, throwing entry point, README), the OTP re-prompt loop, and the registry poll (consecutive-sighting requirement, timeout, transient-error tolerance) via injected fakes. - `doc/PUBLISHING.md`: rewrote the "One-time bootstrap sequence for a new package" section for the placeholder flow, including the `latest` dist-tag window and the trusted-publishing setup ordering (placeholder publish → trusted publisher rule → `"publishFromCi": true`). - `.github/scripts/check-pr-release-bootstrap.mjs` (+ test, + wiring in `run-quality-gates.mjs`): new informational commitperclip notice on PRs that need this bootstrap. It fires when the PR newly release-enables a package that is missing from npm, or adds an unpublished `publishFromCi: false` package that published packages declare a `workspace:*` dependency on, and names the exact maintainer command — so contributors know the red `policy` check is not theirs to fix. It never fails the gate (the `policy` job remains the enforcer), only looks up scope-validated names on the registry, and stays quiet on registry errors. ## Verification - `node --test scripts/bootstrap-npm-package.test.mjs`: 13/13 pass - `node --test .github/scripts/tests/*.test.mjs`: 147/147 pass (10 new for the PR notice) - `pnpm run test:release-registry`: 82/82 pass - Replayed the new PR notice against a real historical PR's live API data (files, manifest at base and head refs): with the registry in its pre-bootstrap state it produces the exact maintainer instruction; with the package bootstrapped it stays silent - Full live end-to-end run: the flow bootstrapped `@paperclipai/adapter-kimi-local` for real — dry-run preview (634-byte, 3-file tarball), publish, registry visibility after ~5 minutes of propagation lag, deprecation confirmed via `npm view ... deprecated` - Guards verified live: an already-published name is rejected, an out-of-scope name (`left-pad`) is rejected, unknown options (including the removed `--otp`) are rejected, and `--publish` in a non-interactive shell fails fast before any network call ## Risks - The `latest` dist-tag points at the deprecated `0.0.0` placeholder until the first stable release supersedes it. This window also existed under the old flow (which parked `latest` at a locally built version); internal consumers are unaffected because release version rewrites pin exact calver versions. - The registry poll caps at ~10 minutes. If propagation is slower than that, the helper prints the exact `npm deprecate ... --otp <code>` command to run manually once `npm view` resolves. - The helper no longer validates the name against the workspace release plan, so a typo within the `@paperclipai/` scope would reserve a wrong name. The dry-run preview shows the exact name before any publish. ## Model Used - Anthropic, **Claude Fable 5** (`claude-fable-5`) via Claude Code, with repository, shell, and Git tooling. It analyzed the existing bootstrap flow and the release scripts (`release-package-map.mjs`, `check-release-package-bootstrap.mjs`, `release.sh` dist-tag handling), wrote the replacement script and tests, updated the documentation, and ran the verification above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bd059a073d |
fix(workspaces): make managed runtimes reliable across restarts (#11740)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Execution workspaces need isolated databases, ports, and runtime services > - Concurrent workspaces could reuse ports or lose service ownership after a restart > - A markerless worktree also needed seed recovery, but normal markerless instances still needed to boot > - This pull request makes seed, port, and service ownership state explicit and recoverable > - It also checks live process and listener identity before it reclaims shared resources > - The benefit is reliable workspace startup, restart, adoption, and concurrent provisioning ## Linked Issues or Issue Description **What happened?** Managed workspaces could lose runtime service ownership after a control-plane restart. Concurrent worktrees could also reuse a port when their parent paths differed. A seed recovery change made every markerless instance resolve a worktree seed source, so normal instances without a source could not start. **Expected behavior** Paperclip must preserve healthy managed services across restarts. It must reserve unique ports across worktree parents. It must provision a registered markerless worktree, but it must skip seed work for a normal markerless instance. **Steps to reproduce** 1. Start two managed worktrees under different parent paths at the same time. 2. Restart the control plane while a managed service stays alive. 3. Start Paperclip with a config that has no seed markers and no registered worktree source. 4. Observe duplicate port selection, lost service adoption, or a seed-source startup error. **Paperclip version or commit** Current `master` plus the workspace runtime reliability changes in this pull request. **Deployment mode** Local development with managed execution workspaces and embedded Postgres. ## What Changed - Added a shared port registry with lease heartbeats, process identity checks, and live listener probes. - Reserved worktree ports across custom parent paths and repaired duplicate legacy assignments. - Preserved and adopted healthy managed services across control-plane restarts. - Reconciled guest bind modes and verified listener ownership before termination or reuse. - Provisioned registered markerless worktree databases and kept normal markerless instance startup as a no-op. - Added CLI, shared, server, and shell regression tests for seed, port, listener, restart, and adoption behavior. - Updated the worktree development documentation. ## Verification - `pnpm exec vitest run cli/src/__tests__/worktree.test.ts --reporter=verbose` — 63 tests passed. - `pnpm exec vitest run packages/shared/src/worktree-port-registry.test.ts --reporter=verbose` — 5 tests passed. - Focused runtime Vitest set — 199 tests passed across 37 suites. - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 10 tests passed. - `git diff --check` passed. ## Risks - Port reservation now depends on lease and process identity data. The fallback listener probe prevents early reclamation when process metadata is incomplete. - Runtime adoption is stricter about bind and owner identity. The tests cover healthy adoption, stale records, PID reuse, and unrelated listeners. - Markerless seed detection now separates registered worktrees from normal instances. The tests cover both paths. - There are no database schema migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with the `gpt-5` model family. The serving snapshot and context-window size are not exposed. The agent used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
a2bf936f9a |
feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktree services run isolated Paperclip instances with cloned databases. > - A reachable service was reported as ready even when its database, runtime identity, or login path was not usable. > - The first candidate added verified database seeding and managed repair in #11665. > - This pull request consolidates that candidate with signed login handoff and a complete readiness contract. > - Post-QA fixes close five defects in repair identity, repair responses, UI retry, seed journal handling, and seed-source trust. > - The benefit is a workspace that either opens safely or reports one accurate recovery action. ## Linked Issues or Issue Description No public GitHub issue exists for this work, so the problem is described here. **What happened** Managed workspace URLs could return HTTP 200 and report ready while login failed. QA also found cases where repair used the wrong instance identity, returned a generic error, left the UI stuck, rejected a safe journal lag, or trusted a mutable workspace manifest. **Expected behavior** Opening a ready workspace signs the board user in to the correct isolated instance. Provisioning and repair use a registered source and report a structured recovery state. **Actual behavior** Entry depended on a password copied into the clone. Several failure paths could publish stale readiness, hide the repair precondition, or trust state that the workspace could modify. **Additional context** This pull request includes the commits first published in #11665. That pull request keeps the original base head for review history. This consolidated pull request is the merge candidate. Related open readiness work includes #11575 and #11621. ## What Changed - Adds a short-lived, signed, single-use login ticket. It binds the user, workspace, instance, and runtime origin. - Exchanges the ticket through Better Auth. It creates the session and cookie through the supported adapter path. - Adds protected workspace readiness fields for the database, clone data, login handoff, seed phase, and runtime identity. - Fails readiness closed when the guest has no company or execution-workspace binding. - Binds ticket issuance to the exact cloned user and active company membership selected for the handoff. - Verifies every current active board identity through the exact-user handoff before publication or reuse. - Gates managed runtime publication on the readiness contract and the recorded worktree instance identity. - Refreshes runtime work products from the live runtime row after a port change. - Adds one workspace access card with ready, degraded, repairing, and failed states. - Uses the runtime response identity for repair. It returns structured repair precondition errors. - Lets a valid source journal lag converge during provisioning. - Binds seed and repair manifests to a source registered outside the agent-writable worktree. - Clears recovered UI errors so a successful retry can open the workspace. - Makes runtime tests register canonical sources and avoid ports owned by live host listeners. - Keeps Vitest on source suites when compiled `dist` trees exist. - Isolates CLI and adapter tests from ambient AWS and runtime API environment variables. - Preserves a 404 response for cross-company workspace ID lookups before runtime authorization. - Makes concurrent single-flight coverage independent of path-canonicalization scheduling order. ## Verification The following checks passed on the integrated head: ```sh pnpm -r typecheck pnpm build pnpm check:token-gates pnpm --filter @paperclipai/db check:migrations ``` - The server source lane passed 420 files and 4,953 tests. Five tests were skipped. - The CLI lane passed 57 files and 385 tests. - The database lane passed 26 files and 97 tests. - The shared package passed 58 files and 506 tests. - The adapter utility lane passed 640 tests. Four tests were skipped. - The Claude adapter passed 220 tests. One test was skipped. - The Codex adapter passed 323 tests. - The OpenClaw adapter passed 13 tests. - The OpenCode adapter passed 42 tests. - The plugin SDK passed 45 tests. - The workspace runtime suite passed 124 tests. - The caller-scoped readiness and handoff suite passed 52 tests. - The workspace provisioning shell suite passed 7 tests. - The runtime exposure suite passed 17 tests while live host mappings occupied fixed test ports. - `git diff --check` passed and the worktree is clean. The serialized route lane will run in GitHub CI with its normal shards. No deployment or active-workspace migration was performed. ## Risks - This is a medium-risk authentication and runtime-readiness change. - The login ticket uses exact origin, workspace, instance, and user binding. It has a short expiry and a one-time nonce. - Runtime publication is stricter. A real readiness, identity, per-user handoff, or control-plane database disagreement now blocks publication. - This pull request supersedes #11665 as the merge candidate. Close #11665 after this pull request merges. - No new database migration is included. The lockfile and workflow files are unchanged. - Deployment and active-workspace migration are intentionally outside this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool use, and code execution produced the main candidate. OpenAI GPT-5 (`gpt-5`) through Codex, with agentic reasoning, tool use, and code execution, integrated the post-QA fixes and hardened the test gates. The Codex context-window size was not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e434829889 |
fix(release): draft-notes baseline survives candidate-cut stables (#11647)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The release workflow drafts the upcoming stable's notes skeleton at
beta publish, ranging from the last stable to the beta's source commit
> - Stables cut from candidate branches leave their tag off master's
lineage, so the ancestor-only baseline rule skips them and falls back a
whole release too far
> - The first live draft did exactly that: it ranged from v2026.722.0
instead of the just-shipped v2026.817.0 and re-included everything
already released
> - This pull request replaces the baseline rule with a merge-base walk
over stable tags and adds a candidate-branch regression fixture
> - The benefit is a correct draft range for every stable lineage:
master-cut, candidate-cut, and promoted-older-source
## Linked Issues or Issue Description
**What happened?**
The first `draft_stable_notes` run (for beta `2026.818.0-beta.1`)
generated `releases/beta/v2026.818.0-beta.1.md` with the range
`v2026.722.0..664052f8e` — 311-commits-worth of already-shipped
v2026.817.0 content re-included.
**Expected behavior**
The draft ranges from the point the shipped stable's content diverges
from the beta source. For v2026.817.0 (cut from
`candidate/release-2026.817.0`) that is the promoted source commit
`8f7b8b3fd`, giving the 172 commits of genuinely-new work.
**Steps to reproduce**
Ship a stable from a candidate branch (tag lands off master's lineage),
then publish a beta from master and read the generated skeleton's range
line.
**Paperclip version or commit**
master at `664052f8e`.
Related (not duplicates): #11567 introduced the generator; its
ancestor-only rule was itself a review fix for the newest-by-version
rule, and this PR is the second iteration with the lineage case the
first fix missed.
## What Changed
- `scripts/draft-stable-notes.sh`: the range start comes from walking
stable tags newest-first and taking the first whose merge-base with the
beta source is a proper ancestor of the source. Candidate-cut stables
resolve to the promoted commit, master-lineage stables to the tag
itself, and tags containing the source are skipped (an older-source
promotion cannot produce an empty range). The skeleton header prints a
runnable short-sha range with the stable tag as a labeled baseline.
- `scripts/draft-stable-notes.test.mjs`: new regression fixture with the
stable tag on an unmerged candidate branch; the existing older-source
and fallback fixtures still pass unchanged.
## Verification
- `node --test scripts/draft-stable-notes.test.mjs` — 8 pass, including
the new fixture.
- Regenerated the live `2026.818.0-beta.1` draft against the real
repository: range start resolves to `v2026.817.0 (merge-base
|
||
|
|
664052f8ea |
feat(release): draft stable notes at beta publish, read them from master at promotion (#11567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system promotes builds canary → nightly → beta → stable, and stable releases publish a GitHub Release from `releases/vYYYY.MDD.P.md` > - The stable lane requires that notes file to exist inside the promoted source commit, but the file is named for the promotion date, which is unknown when the source commit is created > - A promoted beta can therefore never pass the notes check: every happy-path stable is forced through the candidate-branch fix path, with a soak-gate justification, for a notes-only change > - This pull request drafts the notes automatically when the beta is published and lets the stable promotion read them from `master` > - The benefit is a walkable stable happy path: the soak gate stays exact, notes get a real review window during the soak, and the justification path returns to its real purpose (cherry-picked fixes) ## Linked Issues or Issue Description **What existing behavior does this improve?** The stable promotion path in the release channel system (`release.yml`, `scripts/release.sh`). **Current behavior** `release.sh stable` requires `releases/vYYYY.MDD.P.md` in the checked-out source tree, and `publish_stable` checks out the exact promoted SHA. The soak gate requires a `beta/v*` tag to point at that same SHA. No commit can satisfy both for a promoted beta, so a stable promotion must cut a candidate branch with a notes-only commit and bypass the soak gate with a written justification. Release notes are also written at promotion time, under time pressure, with no review window. **Proposed behavior** When a beta publishes, a `draft_stable_notes` job generates a grouped notes skeleton at `releases/beta/v<beta-version>.md` and pushes it to a machine-owned branch; a human opens the PR and edits it during the 3-day soak. The stable preflight resolves notes before the `npm-stable` approval gate: source-tree notes first (the candidate fix path, unchanged), then the merged beta-keyed file on `master`; it fails early with the missing path named when neither exists. After the stable ships, a canonicalization job pushes a branch that moves the file to `releases/vYYYY.MDD.P.md`. Related (not duplicates): #11006 and #11008 introduced the nightly and beta lanes this builds on; older changelog PRs (for example #10669) authored notes manually at promotion time, which is the flow this replaces. **Reason and benefit** The happy path becomes: promote the exact soaked SHA, no justification, notes reviewed during the soak instead of written at the gate. The `releases/vYYYY.MDD.P.md` invariant still holds durably via the canonicalization PR. ## What Changed - `scripts/release.sh`: new `--notes-file PATH` (stable only) overrides where the pre-publish notes check looks, so notes can live outside the source checkout without dirtying the worktree. - `scripts/create-github-release.sh`: same `--notes-file` override for the GitHub Release body. - `scripts/draft-stable-notes.sh` (new): deterministic skeleton generator — commit subjects from the newest stable tag (falling back to the previous beta, then full history) to the beta's source commit, grouped into Features / Fixes / Other. - `.github/workflows/release.yml`: - `draft_stable_notes` job after `publish_beta`: runs the generator and force-pushes `release-notes/v<beta-version>`; the job summary links the compare page. It recreates the beta tag locally if the tag push was rejected (the known workflows-permission case), so drafting is not blocked on manual tag recovery. - `preflight_stable`: computes the target stable version (`release.sh stable --print-version`) and resolves the notes source (`source_tree` → `master_beta` → fail early / warn on dry run); new outputs. - `publish_stable`: materializes `master`-side notes into `RUNNER_TEMP` and passes `--notes-file` to both scripts; outputs the published stable version. - `canonicalize_stable_notes` job: pushes the `git mv` branch after a stable that used `master`-side notes. - `doc/RELEASING.md`, `doc/RELEASE-CHECKLIST.md`: document the drafted-notes flow, the preflight resolution order, and the canonicalization step; the LLM changelog flow now targets the draft branch during the soak. - `.agents/skills/release-changelog/SKILL.md`, `.agents/skills/release-changelog-discord-message/SKILL.md`: the notes-authoring skills now describe this flow — range ends at the beta source commit (not `HEAD`), the file is beta-keyed on the `release-notes/v<beta-version>` branch (seeded with `scripts/draft-stable-notes.sh` for betas that predate the automation), and the canonicalization link caveat is called out for announcements. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 6 tests, temp git-repo fixtures: grouping, stable-tag range, previous-beta and full-history fallbacks, default output path, malformed version, missing tag. - `node --test scripts/release-lib.test.mjs` — unchanged suite still green. - `bash -n` on both changed shell scripts; `release.yml` re-parsed as YAML. - `./scripts/release.sh stable --print-version` unchanged (prints the next stable version); `--notes-file` on a non-stable channel fails with a clear error. - Not exercised end-to-end: the new workflow jobs need a real beta publish to run. The first beta after merge is the live test; the draft job is additive and cannot affect the publish result (it runs after `publish_beta` completes). ## Risks - Low risk to publishing itself: `--notes-file` defaults preserve today's behavior everywhere; the draft and canonicalization jobs are additive and run after the publishes succeed. - The preflight now fails a real stable run when no notes are found. That is the intended fail-early behavior (it previously failed later, inside `publish_stable`, after the `npm-stable` approval). - `draft_stable_notes` force-pushes only the machine-owned `release-notes/v<beta-version>` branch; a beta re-cut regenerates it cleanly. - The stable version computed at preflight could differ from the published one if a run crosses UTC midnight between the two jobs; the materialized notes are passed by path, so the publish still succeeds, and the canonicalization job uses the actually-published version. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
49217aadf0 |
refactor: balance serialized server shards by recorded suite duration (#11528)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ec984502a |
fix(release): stop smoke_beta silently skipping on promote-mode betas (#11582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system re-smokes every published beta as post-publish verification (`smoke_beta`) > - The candidate-branch beta lane (#11209) added `verify_beta_candidate` to `publish_beta`'s needs; that job is skipped on every normal promote-mode beta > - `smoke_beta`'s condition has no status-check function, so GitHub attaches an implicit `success()` that evaluates the needs chain transitively — a skipped ancestor makes it false > - This pull request makes the condition explicit so promote-mode betas smoke again, and pins the shape in the workflow wiring test > - The benefit is that the post-publish beta gate actually runs instead of silently skipping ## Linked Issues or Issue Description **What happened?** Beta `2026.818.0-beta.0` (run 32082007439) published successfully, but its post-publish `smoke_beta` job was skipped. No configuration or input asked for that: the run was a plain `channel: beta` dispatch with `dry_run` at its default `false`, and the same expression `!inputs.dry_run` evaluated true inside `publish_beta`'s own steps (the Docker dispatch step ran). **Expected behavior** Every non-dry-run beta publish is followed by the release smoke suite against the exact published version, as documented in `doc/RELEASING.md` and `doc/RELEASE-CHECKLIST.md`. **Steps to reproduce** Dispatch `release.yml` with `channel: beta` promoting a nightly (promote mode). `verify_beta_candidate` is skipped by design; `publish_beta` runs through its explicit `!cancelled()` condition; `smoke_beta` then skips because its implicit `success()` sees the skipped ancestor in the transitive needs chain (actions/runner#2205 semantics). The beta published on 2026-08-11 predated #11209, so this never surfaced before. **Paperclip version or commit** master at `43ab441f0` (workflow file, current head). Related (not duplicates): #11209 introduced the candidate lane whose skipped job triggers this; #11208 covers the adjacent tag-push failure playbooks. ## What Changed - `smoke_beta`'s condition becomes `!cancelled() && needs.publish_beta.result == 'success' && !inputs.dry_run` — an explicit status-check function suppresses the implicit `success()`, and the result check keeps the dependency on a successful publish. - A comment above the job records why the explicit form is load-bearing. - `scripts/__tests__/release-verify-workflow.test.mjs` pins the new shape so the implicit form cannot silently return. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 8 pass, including the new assertion. - `release.yml` re-parsed as YAML. - The exact skip is visible on run 32082007439 (`smoke_beta: skipped` after `publish_beta: success`); the coverage gap for that beta was closed manually by dispatching `release-smoke.yml` with `paperclip_version: beta` (run 32084880767). - Not exercised end-to-end: the corrected condition needs the next real promote-mode beta to demonstrate; the expression change is minimal and the semantics are the documented actions/runner behavior. ## Risks - Low risk: condition-only change on one job plus a test. Dry runs still skip the smoke (`!inputs.dry_run` retained). Candidate-mode betas, where `verify_beta_candidate` actually runs, behave as before. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
4c349fe6b7 |
feat(runtime): managed Tailscale HTTPS lifecycle, durable runtime leases, and bounded control recovery (#11525)
<!-- Simplified Technical English (ASD-STE100). --> > **Stacked pull request.** This targets #11524. Merge #11524 first. Review only the second commit, `feat(runtime): managed Tailscale HTTPS lifecycle...`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services, so an agent's branch can be previewed while the agent works > - The previous pull request added the host broker, the shared contract, and the database columns, but no code used them > - A managed runtime can only be exposed over HTTPS if it holds a stable loopback port pair for the whole life of the service. The current control path cannot promise this: two controls can race the same execution workspace, a stranded control can stay `running` forever, and a start can adopt a port it does not own > - This pull request adds the HTTPS lifecycle and the control-path hardening that the lifecycle depends on > - The benefit is that a managed preview becomes reachable from another device, and a managed control now always reaches a terminal state ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, workspace operations, the execution workspace routes, and the workspace runtime UI. **Problem or motivation** A managed runtime service is reachable only on loopback, so a preview cannot be opened from a phone or a second computer. Exposing it safely needs an exclusively held port pair. Three existing gaps block that. Overlapping controls can race the same workspace. A control whose owner dies stays `running` and blocks the lane forever. Port allocation does not confirm that the process holding a port is the process Paperclip spawned. **Proposed solution** Add the exposure lifecycle on top of the broker from #11524: reserve before spawn, expose after readiness, validate the public URL, and remove on stop. In the same change, make managed controls mutually exclusive per workspace, give each control a durable issue-owned lease and a terminal state, and verify port ownership before use. **Alternatives considered** - Add HTTPS exposure without the control hardening. This was rejected because a raced or stranded control makes exposure point at the wrong process. - Guard the lane with an in-memory lock only. This was rejected because the lock does not survive a server restart, so the lane can be lost or double-claimed. - Trust the requested bind address. This was rejected because a checkout that predates managed HTTPS overwrites `PAPERCLIP_BIND` from its own `--bind` argument, and then binds the wildcard address. **Roadmap alignment** This completes the managed workspace runtime capability that already exists. It adds no new product surface beyond the HTTPS link. **Additional context** This is the second of three pull requests. The third adds central mediation of leased port pairs. ## What Changed Exposure lifecycle: - Add the server-side broker client and the exposure lifecycle manager. The manager reserves the mapping before spawn, exposes after backend readiness, validates the public URL, and removes the mapping on stop. - Default managed worktree runtimes to `tailscale_https`, read exposure intent from legacy `expose` blocks, and backfill runtimes that are still HTTP-only. - Verify listener ownership for the app port and its Vite HMR companion before the broker is asked to expose anything. An unrelated listener on either port fails the start closed. - Force the loopback bind through argv instead of environment hints. Leave a non-Paperclip service's `--bind` argument alone. - Probe loopback for readiness instead of the public URL, and give Vite HMR its own loopback-bound server in middleware mode. - Preserve operator-declared Serve mappings across the managed lifecycle, so cleanup never removes a mapping that Paperclip did not create. - Name which listener predicate denied an expose, so an operator can act on the message. Control-path hardening: - Make `start`, `stop`, `restart`, and job `run` mutually exclusive per execution workspace. An overlap gets `409 workspace_runtime_control_in_progress`, and authorization is still checked first. - Take a durable exclusivity lease on the execution workspace, owned by the controlling issue. A different issue gets `409 workspace_runtime_lease_conflict` before any operation is recorded. Board and operator actions bypass the lease. - Give every control a terminal state. Each control stamps its owning process and pid, heartbeats while it runs, and has a wall-clock ceiling. Recovery of a stranded control uses a compare-and-swap on `updated_at`, so a live owner is never stolen. - Bound readiness probes, verify allocated port ownership on POSIX and Windows, harden sibling port allocation, and reconcile desired runtimes on server startup. - Surface exposure state and bounded runtime errors in the workspace runtime UI. - Record the new behavior in `doc/DEVELOPING.md`. ## Verification Focused checks, all run on this branch: - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, exactly the count on `master`. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. - `pnpm --filter @paperclipai/ui typecheck` — clean. - Server suites, 177 tests pass across 9 files: `workspace-runtime.test.ts`, `workspace-runtime-leases.test.ts`, `workspace-runtime-control-recovery.test.ts`, `execution-workspace-runtime-control-conflict.test.ts`, `execution-workspace-runtime-lease-route.test.ts`, `workspace-operations-reconciliation.test.ts`, `workspace-runtime-start-terminality.test.ts`, `app-hmr-port.test.ts`, and `workspace-runtime-ready-comment.test.ts`. - Exposure unit suites, 77 tests pass: `src/services/runtime-exposure/` and `workspace-runtime-exposure-backfill.test.ts`. - UI: `WorkspaceRuntimeControls.test.tsx` and `WorkspaceServiceControlBar.test.tsx` — 34 tests pass. **One suite is red on the development host and is expected to be green in CI.** `server/src/services/workspace-runtime-exposure.test.ts` has 10 failures on the machine used to write this branch. The cause is host contamination, not the code. That machine already runs an HTTPS canary that holds ports 42000, 42001, 52000, and 52001 on a tailnet address. The suite allocates from the same range, so the new listener-ownership check correctly reports: ``` listener_ownership_mismatch — port 42000 is bound to 100.123.243.20, 127.0.0.1, fd7a:115c:a1e0:0:0:0:dd3a:f314 ... instead of loopback only ``` A CI runner has no listener on those ports, so the check sees loopback only and the suite passes. Please confirm this from the CI result on this pull request rather than from a local run on a host that already exposes a managed runtime. This is a real weakness of the current test fixture, and the third pull request in the series removes it by allocating the pair through a central mediator instead of a stubbed availability check. `workspace-runtime-https-live-exercise.test.ts` needs a live `tailscale` host and was not run locally. ## Risks - This is the behavior-bearing pull request of the three, so it carries the most risk. - Two new `409` responses appear on managed control routes. A caller that assumed a control always starts must handle a conflict. Board and operator actions are deliberately exempt, so an agent lease cannot lock an operator out. - Managed worktree runtimes now default to `tailscale_https`. If the host has no working broker, the start fails closed and reports the exposure failure instead of silently serving plain HTTP. This is intended, and it is the reason the failure message names the denying predicate. - Startup reconciliation touches persisted runtime rows. It is scoped to desired state and does not resurrect a service that never came up. - The lease has a 30-minute time to live and explicit release paths, so a crashed owner cannot hold a lane forever. - No migration runs in this pull request. The tables and columns land in #11524. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass, with the one host-contaminated suite explained above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d2665ff6b4 |
fix(ui): align the mobile task chat composer with the thread (#11296)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task thread is where people read work and guide agents. > - The mobile composer should use the same content width as the thread. > - The composer kept the desktop 80% width on mobile, so its edges did not align with the thread. > - Long assignee-aware placeholder text could also clip inside the mobile editor. > - Some extracted style tokens used legacy HSL wrappers around complete semantic colors, which made those declarations invalid. > - This pull request makes the composer full width on mobile, preserves the narrower desktop layout, wraps the placeholder, and repairs the invalid color compositions. > - The benefit is a stable mobile composer that aligns with the task thread and keeps its intended visual styles. ## Linked Issues or Issue Description Related work: Refs #11263. **What happened?** At mobile widths, the task chat composer used the same 80% width as the desktop composer. Its horizontal edges did not align with the full task thread. A long assignee-aware placeholder could clip on one line. The composer's extracted shadow also used a legacy `hsl(var(...))` wrapper around complete semantic color values, so the browser could reject the declaration. **Expected behavior** The composer must match the task thread width on mobile. It must stay narrower on larger screens. Long placeholder text must wrap inside the editor. Semantic color tokens must form valid shadows and gradients. **Steps to reproduce** 1. Open a task with the chat-style thread on a mobile viewport. 2. Compare the composer edges with the task thread edges. 3. Select an assignee whose placeholder text wraps to two lines. 4. Inspect the computed composer shadow and the extracted semantic color styles. **Paperclip version or commit** The change is based on `dc6fcd1ff1` from `master`. **Deployment mode** Local build from source. The behavior also applies to packaged web builds. ## What Changed - Made the task chat composer full width below the medium breakpoint and kept the 80% desktop width. - Matched the composer dock padding to the task thread padding. - Allowed long composer placeholders to wrap and reserved enough mobile editor height for two lines. - Replaced invalid legacy HSL wrappers around full semantic colors in extracted shadows, gradients, and approval styles. - Added a token gate that prevents legacy `hsl(var(--token))` wrappers from returning. - Added focused regression tests for responsive width, padding, placeholder wrapping, mobile height, and semantic shadow validity. ## Verification - `pnpm check:token-gates` — all four gates pass. - `pnpm --filter @paperclipai/ui exec vitest run src/components/TaskChatThread.test.tsx src/components/task-chat/TaskChatComposer.test.tsx src/components/task-chat/TaskChatComposerStyles.test.ts` — 37 tests pass. - `pnpm --filter @paperclipai/ui typecheck` — passes. - `pnpm --filter @paperclipai/ui build` — passes. The build prints existing CSS optimizer and bundle-size warnings. ## Risks - Low risk. The width change is limited to the mobile breakpoint. The desktop 80% layout remains in place. - The semantic token fixes can affect shadows and gradients that were previously invalid. The new gate prevents the invalid wrapper pattern from returning. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5.6-sol`. The context-window size is not exposed in this environment. The model used reasoning, repository tools, code execution, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a8d118a779 |
Prefer public base URL for generated invite links (#7619)
Fixes #7623 ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company invites are part of the access subsystem and must produce URLs that recipients can open from outside the host machine. > - Paperclip already has public/auth base URL configuration for deployments behind a public hostname, Tailscale, or a reverse proxy. > - Invite URL composition was still deriving its origin from the incoming request host, so loopback-bound servers emitted `http://127.0.0.1:3100/invite/...`. > - A loopback invite URL is not shareable with a remote human or agent, even when the token itself is valid. > - This pull request makes invite URL builders prefer the configured public base URL and keep the existing request-host fallback when it is unset. > - The benefit is that copied invite links use the reachable deployment origin without changing local-only behavior. ## Linked Issues or Issue Description Fixes #7623 No duplicate or related PRs/issues were found in a GitHub search for invite URL, loopback, public base URL, and `authPublicBaseUrl` terms. ## What Changed - Added base URL resolution in `server/src/routes/access.ts` that strips trailing slashes and prefers configured `authPublicBaseUrl` over the request-derived host. - Threaded `authPublicBaseUrl` through invite summary, invite onboarding manifest, onboarding text, access routes, `createApp`, and server startup wiring. - Added `server/src/__tests__/invite-url-public-base-url.test.ts` covering configured public-base precedence, unset fallback behavior, and trailing-slash normalization. - Registered the invite public-base URL test in the serialized Vitest server runner. ## Verification ```bash pnpm install --frozen-lockfile pnpm exec vitest run server/src/__tests__/invite-url-public-base-url.test.ts pnpm run test:run:serialized ``` Local results from the rebased PR branch: - `pnpm install --frozen-lockfile` exited 0. - Targeted invite URL test exited 0: 1 file, 3 tests passed. - Serialized server suite exited 0: 106 serialized suites completed; the new invite URL test passed inside that runner. Manual check after deployment: set `PAPERCLIP_AUTH_PUBLIC_BASE_URL` or equivalent public base URL config, create a company invite, and confirm the returned/copied invite URL uses that public origin instead of `127.0.0.1`. ## Risks Low risk. The new public base URL parameter is optional and falls back to existing request-derived behavior when unset. The main operational risk is misconfigured public base URL input; the implementation only trims trailing slashes and otherwise trusts the configured origin. ## Model Used - Original implementation: Anthropic `claude-sonnet-4-6`, 200k context, tool use and test execution. - Conflict repair and verification: OpenAI Codex GPT-5.5, coding agent with shell, git, GitHub CLI, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Coder (Claude) <coder-claude@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip Coder (Claude) <lad-agent@paperclip.ing> |
||
|
|
71e9d6bb0b |
feat(release): candidate-branch beta builds and the release checklist (#11209)
> Follow-up to #11208 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channels promote artifacts along canary → nightly → beta → stable, with the happy path being promotion of an existing build > - When one or two targeted fixes are needed before a beta or stable, the only options today are waiting for the next nightly or absorbing a whole day of unrelated master changes > - The channel model was designed with an escape hatch for exactly this: short-lived candidate branches carrying only cherry-picked fixes > - This pull request implements candidate-branch beta builds with full verification, documents the stable fix path through the soak-justification gate, and adds the release captain's checklist > - The benefit is that a surgical fix can ship forward without either delay or blast radius, with its provenance recorded ## Linked Issues or Issue Description Refs #11008 — completes the fix-path half of the channel model introduced there. **Subsystem affected** Release automation: `scripts/release.sh`, `.github/workflows/release.yml`, `doc/RELEASING.md`, new `doc/RELEASE-CHECKLIST.md`, tests. **Problem or motivation** Beta promotion only accepts commits that already shipped as a nightly, and stable promotion expects a soaked beta. There is no supported way to ship one or two cherry-picked fixes between lanes: an urgent fix must wait for the nightly cycle or pull in every unrelated master change from the day. The original channel design called for candidate branches to cover this, and they were deferred from the initial implementation. **Proposed solution** Candidate-branch beta builds: cut `candidate/beta-<target>` from a nightly's source commit, cherry-pick the fixes, and dispatch `channel: beta` with the new `candidate_branch` input. Selection enforces the naming convention, rejects heads that already shipped as a beta or predate the candidate tooling, and records the cherry-picked commits in the job summary. Because candidate heads never went through a canary or nightly, publication is gated on a full `release-verify` run (promoted nightlies keep skipping re-verification). The stable fix path (`candidate/release-<target>` as `source_ref`) works through the existing soak gate: the justification requirement is the deliberate, recorded trade-off for shipping unsoaked bits, and is now documented as such. ## What Changed - `scripts/release.sh`: `--from-candidate` flag (beta only) waives the shipped-a-nightly requirement while keeping the duplicate-beta guard - `.github/workflows/release.yml`: `candidate_branch` dispatch input; candidate mode in `select_beta` (naming validation, duplicate and tooling-era rejection, cherry-pick recording); new `verify_beta_candidate` job gating candidate publishes on full verification - `doc/RELEASING.md`: beta fix-path and stable fix-path sections - `doc/RELEASE-CHECKLIST.md` (new): the release captain's checklist for all four lanes as built - Tests: dry-run fixture coverage for `--from-candidate` (waives the nightly guard, keeps the duplicate guard, rejected outside beta) and wiring tests for candidate validation plus the verification gate ## Verification - `node --test` on the four affected suites: 42 pass in total (17 + 25 across the two runs), including the 5 new tests - `bash -n` on `release.sh`; YAML parse of the workflow - After merge: exercise the path end to end the first time a real cherry-picked beta is needed — dispatch with a `candidate/beta-*` branch and confirm the summary records the picks and verification runs ## Risks - Candidate builds bypass the smoke-tested-nightly provenance by design; the compensating controls are full verification before publish, the post-publish beta smoke, the human `npm-beta` gate, and recorded cherry-picks - The stable fix path rides the existing justification mechanism rather than adding a second bypass — one recorded escape hatch, not two ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2da6a248c3 |
fix(release): surface recovery commands when a lane tag push is rejected (#11208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's promotion lanes publish to npm, then push a lane tag and dispatch the Docker image build at that tag > - The first nightly of the beta-tooling merge published to npm and then died at the tag push: GITHUB_TOKEN may not create refs pointing at workflow-modifying commits from dispatch or scheduled runs > - The failure was a bare `remote rejected` with no guidance, leaving the release half-finished (npm live, no tag, no images) until an operator reverse-engineered the recovery > - This pull request makes every lane's tag push degrade into exact recovery instructions in the job summary > - The benefit is that a rare platform-permission rejection becomes a two-minute runbook operation instead of a forensic exercise ## Linked Issues or Issue Description Refs #11008 — the incident occurred promoting that change's own merge commit, the first workflow-modifying commit to flow through the lanes it introduced. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31445344811 published `2026.811.0-nightly.0` to npm, then failed pushing `nightly/v2026.811.0-nightly.0`: `refusing to allow a GitHub App to create or update workflow .github/workflows/release.yml without workflows permission`. The tagged commit modifies workflow files, and GITHUB_TOKEN may not create refs pointing at such commits from dispatch or scheduled runs (push-event runs are exempt, which is why the canary tag on the same commit succeeded). The job failed with no explanation and the Docker dispatch never ran. **Proposed solution** Wrap the nightly, beta, and stable tag pushes: on rejection, write the exact recovery commands into the job summary — create and push the tag with maintainer credentials, dispatch `docker.yml` at the tag, and for stable also run `create-github-release.sh` — then fail the job. Document the cause and recovery in the failure playbooks and pin the three recovery blocks with a wiring test. ## What Changed - `.github/workflows/release.yml`: recovery-summary wrappers on the nightly, beta, and stable tag-push steps - `doc/RELEASING.md`: failure-playbook entry for the workflows-permission rejection - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting all three lanes carry the recovery summary ## Verification - Wiring tests: 6 pass; YAML parse of the workflow - The recovery commands are exactly the ones used to resolve the real incident (tag push + `docker.yml` dispatch for `nightly/v2026.811.0-nightly.0`) ## Risks - Low. The happy path is unchanged (a successful push skips the wrapper); the failure path trades a bare error for actionable output and still fails the job, since the release state is genuinely incomplete ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |