mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
66575fe519db7320147aece94fa66e15eba375c1
3529
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
66575fe519 |
fix(paperclip-page): scope uploader credentials to the publish helper (#10894)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents publish static pages with the paperclip-page skill and its `publish.sh` helper > - The skill docs told operators to bind the page-uploader IAM keys as the global `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` > - Static env keys have precedence over `AWS_PROFILE` in the AWS CLI and in all AWS SDKs > - Because of this, each agent run lost the host role identity and lost access to Secrets Manager and other AWS services > - This pull request adds namespaced credential variables that apply only to the helper's own `aws` calls > - The benefit is a stable host AWS identity in agent runs, with no change to page publishing ## Linked Issues or Issue Description No public GitHub issue exists. Description of the problem: **What happened?** Agent runs on a host with `AWS_PROFILE` set lost access to AWS Secrets Manager. The failures looked intermittent. The cause is deterministic: the page-uploader keys were bound as global `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` in agent run environments. These static keys shadow `AWS_PROFILE`. Each process in the agent run then used the S3-upload-only uploader identity. **Expected behavior** The page-uploader credentials apply only to the page publish helper. All other processes keep the host identity from `AWS_PROFILE`. **Steps to reproduce** 1. Set `AWS_PROFILE` to a role with Secrets Manager access. 2. Export `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` for an IAM user without that access. 3. Run `aws sts get-caller-identity`. The identity is the IAM user, not the role. 4. Run `aws secretsmanager list-secrets`. The call fails with `AccessDeniedException`. ## What Changed - `publish.sh` reads `PAPERCLIP_PAGE_AWS_ACCESS_KEY_ID` and `PAPERCLIP_PAGE_AWS_SECRET_ACCESS_KEY`, with optional `PAPERCLIP_PAGE_AWS_SESSION_TOKEN`. - The helper applies these values only to its own `aws` invocations. It clears ambient `AWS_PROFILE` and `AWS_SESSION_TOKEN` for those calls. - Credential precedence is: namespaced key pair, then `PAPERCLIP_PAGE_AWS_PROFILE`, then the ambient credential chain. Existing global-name bindings continue to work during migration. - Validation: the key pair must be set together. The pair plus `PAPERCLIP_PAGE_AWS_PROFILE` is an error. A session token without the pair is an error. - `SKILL.md` and `README.md` now instruct operators to bind the secrets under the namespaced names and explain the shadowing hazard. ## Verification - Run `node --test .agents/skills/paperclip-page/scripts/publish.test.mjs`. All 11 tests pass. - New tests cover: the incomplete key pair, the pair-plus-profile conflict, the token-without-pair error, and a fake-`aws` environment capture that proves the helper's calls see the page keys while `AWS_PROFILE` and `AWS_SESSION_TOKEN` stay unset. - Run `bash -n .agents/skills/paperclip-page/scripts/publish.sh` for a syntax check. ## Risks - Low risk. The change is contained in one skill helper and its documents. - The ambient credential chain remains the fallback, so current deployments do not break before operators rebind the secrets. - Operators must rebind the two page secrets to the namespaced names to get the benefit. The README documents this. ## Model Used Claude Fable 5 (`claude-fable-5`), Anthropic. Context window: 1,000,000 tokens (128K max output). Agentic coding session with extended thinking and tool use (Claude Code harness). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d648becb90 |
refactor(ci): split workspaces-a into two Vitest native shards
Split the slow workspaces-a CI lane into two Vitest native shards and keep release verification in parity. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.811.0-canary.3 nightly/v2026.811.0-nightly.1 |
||
|
|
6601014898 |
fix(release): reject promotion sources that predate their channel tooling (#11197)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem promotes builds along canary → nightly → beta → stable, and each publish job checks out the promotion's source commit and runs that tree's release tooling > - The first beta dispatch failed with `unexpected argument: beta`: the selected nightly's source predated the beta channel, so its `release.sh` did not know the argument > - The failure was clean (argument parsing, nothing published) but cryptic, and the same trap waits for any promotion of a source older than its target channel's tooling > - This pull request makes the selection jobs reject such sources with an actionable error and documents the property > - The benefit is that a bootstrapping or old-source promotion fails in seconds with instructions, instead of mid-publish with a parser error ## Linked Issues or Issue Description Refs #11008 — the guard hardens the beta promotion flow introduced there, after its first dispatch surfaced the gap described below. **Subsystem affected** Release automation: `.github/workflows/release.yml`, `doc/RELEASING.md`, workflow wiring tests. **Problem or motivation** Run 31444045044 (first beta dispatch) failed in `publish_beta` with `unexpected argument: beta`. Promotions deliberately build from the pinned source commit, which means they also run that commit's `scripts/release.sh` — and a source that predates the target channel's introduction cannot publish it. Nothing guards this today; the error surfaces deep in the publish job with no explanation. **Proposed solution** Guard at selection time: `select_nightly` requires the source canary's `release.sh` to know the nightly channel, and `select_beta` requires the source nightly's `release.sh` to know the beta channel. Each guard literally matches the channel case arm and fails closed with a clear message naming the remedy (promote a newer source). Document the tooling-era property in `RELEASING.md` and pin the guards with a wiring test. ## What Changed - `.github/workflows/release.yml`: tooling-era guards in `select_nightly` and `select_beta` - `doc/RELEASING.md`: documents that promotions run the source commit's release tooling - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test pinning both guards ## Verification - Wiring tests: 5 pass - Guard expressions exercised against real commits: accepts the beta-capable merge commit of the beta-channel change, rejects a pre-beta commit - YAML parse of the workflow - After merge: the next beta dispatch selects a beta-capable nightly and passes the guard ## Risks - Low. Selection-time check only; the guards match the channel case arm literally and fail closed (with the same actionable message) if that line is ever reformatted ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.2 |
||
|
|
35aaaa0bd0 |
feat(server): preserve task timestamps and hierarchy through company import/export (#11193)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company export/import moves a whole company — agents, tasks, comments — between instances as a portable bundle > - The bundle never carried task timestamps or parent links: the export writes neither, the importer lets database defaults stamp "now", and sub-tasks arrive flattened > - Boards sort by recency, so every imported task showing "created just now" collapses the task list into import order, and the task hierarchy the user built is gone > - This pull request adds created/updated/started/completed/cancelled timestamps and a parent link to the bundle (schema v7), preserves them end to end on import, and keeps comment imports from clobbering a preserved updated time > - The benefit is that an imported company reads like the company the user left: same recency order, same task tree ## Linked Issues or Issue Description **What happened?** After a company import, every task showed as created at import time. Recency sorting collapsed to import order, and parent/child task nesting disappeared. The user called out losing "the meaningful task hierarchy and recency sorting". Cause: the export bundle has no fields for task timestamps or parent links, the importer lets `defaultNow()` win on insert, and the comment importer bumps every touched task's `updatedAt` to now. **Expected behavior** An imported company preserves each task's creation/update/start/completion times and its position in the task tree, so sorting and nesting on the destination match the source. **Steps to reproduce** 1. On a source instance, create tasks over several days, including sub-tasks nested under parents. 2. Export the company and import it into another instance. 3. Every task shows the import moment as its creation/update time and all tasks are top-level. ## What Changed - Export writes `createdAt`/`updatedAt`/`startedAt`/`completedAt`/`cancelledAt` (ISO, only when set) and `parent: <taskSlug>` into each task's bundle extension; a parent outside the export selection drops the edge with an aggregate warning, mirroring the existing blocker-edge warning (`server/src/services/company-portability.ts`). - Bundle schema version 6 → 7. All new fields are optional: v5/v6 bundles import unchanged with a version-aware downlevel warning; bundles newer than the board still fail closed. - Manifest parsing validates the new timestamps like comment timestamps (invalid → warn and ignore, never a hard failure); shared types and the zod validator carry the new optional fields. - Import resolves parent slugs to pre-generated destination ids, drops self-references and cycles from tampered bundles with warnings, and orders rows parents-first because the self-referencing FK is checked per insert chunk. - `importIssues` writes the preserved timestamps (falling back to insert time when absent; `startedAt` stays null unless bundle-carried, per #11191's semantics) and `parentId`. - `addImportedComments` no longer blanket-bumps `updatedAt = now()`; it takes `GREATEST(updated_at, newest imported comment createdAt)`, so a preserved update time never regresses while unpreserved rows keep the old behavior. ## Verification - `pnpm vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts server/src/__tests__/productivity-review-service.test.ts` — 102 passed, 1 pre-existing opt-in benchmark skip. Includes: full round-trip with exact timestamp equality and a 3-deep parent chain against embedded Postgres; v6 back-compat (defaults + warning); forward-compat rejection (v8); cycle/self-reference/invalid-timestamp tampered-bundle handling; comment-bump preserve-awareness in both directions. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/shared typecheck` — clean. ## Risks - **Rollout ordering**: a board on the previous build (max schema v6) refuses bundles exported by this build (stamped v7) — the existing newer-than-supported rejection, working as designed. Cross-instance moves need the importing board upgraded first. Called out here so operators aren't surprised during the transition window. - Parent edges from tampered bundles are dropped with warnings rather than failing the import; blocker relations already behave this way. - Timestamps are data-only; no destination schema migration. Stacked on #11191 (its commit is included here) — merge #11191 first; this PR then shows only the v7 changes. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4c062a0eb2 |
fix(ui): stop coercing imported agents to the destination CEO adapter (#11192)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import lets an operator bring a package of agents into an instance, and each agent declares which adapter runs it (Claude Code, Codex, and so on) > - The export and the server importer preserve each agent's adapter faithfully, but the Import page seeds an adapter override for every agent with the destination CEO's adapter before the user touches anything > - Every imported agent therefore arrives as the CEO's adapter (usually Claude Code) even when the source package holds a mix, and the picker shows the coerced value as if it were the source's, so nothing looks wrong > - This pull request makes the manifest adapter the default, sends overrides only for agents the user actually changed, and replaces the silent coercion with an explicit per-agent fallback warning when the destination truly lacks the source adapter > - The benefit is that a mixed Claude/Codex team imports as a mixed Claude/Codex team, and any real adapter gap is visible instead of silent ## Linked Issues or Issue Description **What happened?** A user imported a company package whose agents were a mix of Claude Code and Codex on the source instance. After the import, every agent was configured as Claude Code. The import preview showed no sign that anything had been changed. Cause: the Import page initializes its adapter-override map by assigning every agent the destination CEO's adapter type and sends that override for every agent, overriding the manifest's per-agent adapter server-side. For imports into a new company, the "CEO adapter" is read from whichever unrelated company is currently selected. **Expected behavior** Imported agents keep the adapter declared in the package. An override is sent only when the operator explicitly picks a different adapter, or when the source adapter is not installed on the destination — and in that case the page must say so per agent, not silently substitute. **Steps to reproduce** 1. On a source instance, create a company with one Claude Code agent and one Codex agent, and export it. 2. Import the package on another instance whose CEO uses Claude Code, changing nothing in the import dialog. 3. Both agents arrive configured as Claude Code; the Codex identity is gone. ## What Changed - The preview no longer seeds adapter overrides; the override map starts empty, and the picker displays each agent's manifest adapter (`ui/src/pages/CompanyImport.tsx`). - `buildFinalAdapterOverrides` sends an entry only when the effective adapter differs from the manifest or the agent's adapter config was edited — untouched agents flow through with no override. - The page fetches the destination's installed adapters (existing `adaptersApi.list()` client). When a manifest adapter is missing or disabled on the destination, only that agent defaults to the CEO's adapter, with a visible amber warning naming both adapters. If the adapters request fails, the page fails open: manifest adapters are kept and no coercion happens. - Tests: untouched mixed-adapter import sends no overrides; a user-changed agent sends exactly one; a missing destination adapter produces the fallback plus rendered warning for that agent only; an adapters-endpoint failure produces no coercion. ## Verification - `npx vitest run ui/src/pages/CompanyImport.test.tsx` — 19 passed (15 pre-existing + 4 new). - `pnpm --filter ./ui typecheck` (`tsc -b`) — clean. ## Risks - Behavior change: users who previously relied on the silent conversion (importing packages that reference adapters they don't have) now get an explicit per-agent fallback with a warning — same outcome, visible. The server's hard rejection of unknown adapter types remains the backstop for API callers. - UI-only change; no server or schema impact. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d816eb8095 |
fix(server): keep imported tasks quiescent under the productivity review sweep (#11191)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import brings a full company package — agents, tasks, routines — into an instance, with `pauseAutomations` promising a quiet landing > - The pause covers the imported entities, but the destination's own productivity-review sweep does not know the difference between imported rows and live work > - The importer stamps every imported in-progress task with `startedAt = now()`, so six hours later the sweep's long-active check fires on every one of them and floods the board with review tasks and agent wakeups > - This pull request stops fabricating `startedAt` on import and makes the sweep skip tasks whose assignee agent is paused > - The benefit is that an import lands quietly: no surprise review-task storm, and paused teams stay paused until the operator activates them ## Linked Issues or Issue Description **What happened?** After importing a company package with automations paused, a batch of "productivity review" tasks appeared roughly six hours later — one for every imported in-progress task — each with an owner-agent wakeup. The user described it as jarring and wasteful. Cause: `importIssues` fabricates `startedAt = now()` for imported in-progress rows, and `reconcileProductivityReviews` considers any assigned in-progress task without checking whether the assignee agent is paused, so its long-active-duration evidence (6 h threshold) trips on the fabricated timestamp. **Expected behavior** An import with paused automations must be quiescent: no destination sweep should generate work from imported rows until the operator unpauses the imported team. A paused agent must not accumulate review tasks it cannot act on. **Steps to reproduce** 1. Import a company package containing tasks with status `in_progress` assigned to agents, with "pause automations" enabled. 2. Wait for the productivity-review reconcile (runs at startup and on the heartbeat scheduler tick) more than six hours after the import. 3. Observe one new review task plus an owner wakeup per imported in-progress task. ## What Changed - `importIssues` no longer fabricates `startedAt` for imported `in_progress` rows; it inserts null (`server/src/services/issues.ts`). Audited every consumer of `issues.startedAt` — all are null-tolerant, and normal checkout/status-transition paths set the value when work really starts. - `reconcileProductivityReviews` skips candidates whose assignee agent is `paused`, counting them as skipped (`server/src/services/productivity-review.ts`). This is a general rule, not import-specific: a paused agent cannot act on a review. - Tests: paused-assignee candidate with an old `startedAt` creates no review, and creates one after unpausing; imported in-progress issue lands with null `startedAt` (embedded-Postgres import test); the pre-existing long-active regression test still passes. ## Verification - `pnpm vitest run server/src/__tests__/productivity-review-service.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — 20 passed, 1 pre-existing opt-in benchmark skip. - `pnpm vitest run server/src/__tests__/company-portability.test.ts` — 78 passed. - `pnpm --filter @paperclipai/server typecheck` — clean. ## Risks - Behavior change beyond imports: tasks assigned to paused agents no longer receive productivity reviews anywhere. This is intended — the review would target an agent that cannot respond — and reviews resume on the first reconcile after unpausing. - Imported in-progress tasks now carry no `startedAt` until real work starts on the destination. The one sweep that read the fabricated value is the one this PR quiets; all other consumers fall back safely (audit in the commit body). - Low risk otherwise: no schema change, no API shape change. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (multi-agent implementation with independent verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f7b8b3fda |
feat(release): add human-gated beta channel with stable soak enforcement (#11008)
> Follow-up to #11006 (merged): rebased onto master and ready for review. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem now publishes canary (every master push), nightly (scheduled, smoke-gated, added in #11006), and stable (manual) > - There is still no human-approved release-candidate lane between nightly and stable, and nothing enforces that a stable actually soaked anywhere before shipping > - Betas need a real approval gate, and stables need a soak policy that is data, not prose > - This pull request adds the beta channel: a manual promotion of a chosen nightly behind the `npm-beta` environment gate, re-smoked after publish, plus a stable preflight that enforces a 3-day beta soak with a written-justification bypass > - The benefit is a complete canary → nightly → beta → stable train where every stable shipped as a beta first, and emergencies leave a written trace ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** After #11006 the project has canary and nightly prerelease lanes, but no release-candidate lane. Stable promotion has no enforced soak: any ref can ship as stable directly. There is no approval boundary for a broader-audience prerelease, and no structured way to record why an emergency release skipped validation. **Proposed solution** Add a `beta` channel: a manual dispatch that promotes a chosen nightly's source commit, publishes behind the `npm-beta` GitHub environment (required reviewers are the gate), re-smokes the published beta, and tags `beta/vX`. Enforce in the stable path that the source commit shipped as a beta at least 3 days earlier (measured from the beta's npm publish time), with a `skip_soak_justification` input as the recorded emergency bypass. **Alternatives considered** Codifying the soak policy in docs only. Rejected: an unenforced policy decays; the preflight makes the policy executable while the justification input keeps the emergency path usable and auditable. ## What Changed - `scripts/release.sh` + `scripts/release-lib.sh`: `beta` channel — requires HEAD to carry a `nightly/v*` tag, publishes the package set as `YYYY.MDD.P-beta.N` under dist-tag `beta`, tags `beta/vYYYY.MDD.P-beta.N` - `.github/workflows/release.yml`: - `channel: beta` dispatch path: `select_beta` resolves the newest (or an explicit `source_version`) nightly and fails loudly on selection problems; `publish_beta` runs behind the `npm-beta` environment, pushes the tag, and dispatches `docker.yml`; `smoke_beta` re-runs the release smoke suite against the exact published beta version - stable path: new `preflight_stable` job enforces the 3-day beta soak from the beta's npm publish time; `skip_soak_justification` bypasses with the reason echoed into the job summary; dry runs report without blocking - `.github/workflows/docker.yml`: `beta/v*` tags publish `:beta` on both images, with exact version stamping - `.github/workflows/release-smoke.yml`: `beta` added to the dispatch choice list - Docs: `CHANNELS.md` beta entries; `RELEASING.md` beta lane, soak gate, and failure playbook; `RELEASE-AUTOMATION-SETUP.md` `npm-beta` environment setup, including the warning to create the environment before the first beta dispatch (GitHub auto-creates unprotected environments on first reference) - Tests: beta version-counting coverage in `scripts/release-registry-versions.test.mjs`; beta identity and nightly-tag guard coverage in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the two touched suites: 17 pass, including the 3 new beta tests - `bash -n` on both shell scripts and YAML parse of all three workflows - After merge, in order: create the `npm-beta` environment, dispatch `channel: beta` with `dry_run: true` to preview, then a real promotion of a published nightly through the approval gate, then a stable dry-run against a young beta to see the soak gate report ## Risks - If the `npm-beta` environment does not exist when the first beta dispatch runs, GitHub creates it with no protection rules and the beta publishes without approval. Mitigated by documentation and by creating the environment before merge (operator step) - Until the first beta exists, every stable dispatch requires `skip_soak_justification`. This is deliberate — the first beta ships immediately after this merges — but it is a behavior change to the stable dispatch - The soak clock reads the beta's npm publish time from the registry; a registry outage makes the preflight fall back to requiring justification (fail-closed) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.811.0-canary.0 beta/v2026.811.0-beta.0 nightly/v2026.811.0-nightly.0 |
||
|
|
459469638e |
feat(adapter-codex-local): add secure device-login building blocks (#11097)
## Thinking Path > - Paperclip connects AI agents to local and remote runtimes. > - The Codex local adapter needs a safe device-login flow. > - A future sandbox integration needs strict prompt validation, secret protection, cleanup, and private credential storage. > - This pull request adds tested building blocks for that flow. > - The result gives a later Daytona integration a clear security boundary. ## Linked Issues or Issue Description No public issue covers this change. **Problem or motivation** The Codex local adapter has no safe, reusable flow to prove device login inside an isolated sandbox. **Proposed solution** Add parser, runner, credential export, and proof helpers. Validate the prompt, protect login data, store credentials in a private run-scoped home, and dispose all sandbox resources. **Alternatives considered** Do not connect a production Daytona driver in this change. Use an injected sandbox driver and focused tests first. This keeps the security controls testable before live provider integration. **Roadmap alignment** The change extends the Codex local adapter. It does not add a core Paperclip route or duplicate a planned core feature. **Additional context** The flow keeps the login URL, code, and token out of logs, results, and errors. The proof home uses a company-scoped root and a run-scoped private directory. ## What Changed - Add a pure parser for the exact Codex device-login URL and one-time code shape. - Add a sandbox runner with prompt handling, timeout, cancellation, and disposal. - Add a credential export step with company scoping, path checks, payload checks, private modes, locking, and cleanup. - Add redacted device-login fixtures and focused tests for parsing, secret redaction, runner outcomes, credential export, and cleanup. ## Verification - Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`. - Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc --noEmit`. - Review tests for strict URL and code validation, timeout, cancellation, disposal, secret redaction, path safety, payload safety, file modes, and cleanup. ## Risks - This change provides building blocks, not a live Daytona proof. - A later integration must connect the runner to a concrete sandbox driver. - Credential export depends on existing Codex authentication cache helpers. - Incorrect path or payload assumptions can reject valid credentials. ## Model Used Codex, GPT-5, tool use, code execution, and repository review. The Paperclip runtime controls the exact context window and reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.810.0-canary.6 |
||
|
|
5a0985f80a |
test(release-smoke): update onboarding spec for the mission-first wizard (#11190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly publish on the release smoke suite, which drives real onboarding in a browser against the published artifact > - With the harness fixed (#11187, #11189), the gate reached the Playwright suite for the first time in CI — and the spec still walks the old onboarding wizard, so it fails at "Create your first agent" on every current build > - The wizard was redesigned to a mission-first five-step flow, and the spec rotted silently because the suite never ran in CI before > - This pull request rewrites the spec to drive the current wizard end to end > - The benefit is a smoke gate that actually tests today's product, verified against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `tests/release-smoke/docker-auth-onboarding.spec.ts`. **Problem or motivation** Nightly run 31431273139 failed in the smoke Playwright suite: the spec expects the old wizard step "Create your first agent", but current builds show the redesigned mission-first flow (front door → company → mission → team lead → connect model → review). The page snapshot in the run artifact shows the "Define your mission" step where the spec expected the agent step. Both retries failed identically — this is deterministic spec drift, not flake. **Proposed solution** Rewrite the spec for the current flow: fill the company name, define the mission directly (confirming creates the company), name the team lead, hire it through the adapter step — the adapter environment probe reports unhealthy in the CLI-less smoke container by design and must not block the hire — then launch to the dashboard. Assert the company, the ceo-role agent, and the company goal through the API. The first-task and assignment-run assertions are removed together with the wizard flow that created them. ## What Changed - `tests/release-smoke/docker-auth-onboarding.spec.ts`: rewritten for the mission-first wizard; sign-in and wizard-opening helpers and the company-name step are unchanged ## Verification - Full local run against the real nightly candidate: launched the smoke container for `paperclipai@2026.810.0-canary.3` via `scripts/docker-onboard-smoke.sh` (with the #11189 bind fix), then ran `pnpm run test:release-smoke` against it — 1 passed (4.5s) - After merge: dispatch `release.yml` with `channel: nightly` to run the full gate in CI ## Risks - Low. Test-only change. The spec now asserts less about first-task creation because the wizard no longer creates a first task; if a first-run trigger returns to onboarding, the spec should grow that assertion back ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI artifact forensics, UI source tracing, local Docker + Playwright reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.5 |
||
|
|
f94f6003c6 |
fix(release-smoke): pin the smoke container to the lan bind preset (#11189)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly on the release smoke suite, which boots the published artifact in a Docker container and drives real onboarding > - The gate kept failing even after the readiness budget fix (#11187), and the new container-log dump revealed the server was healthy but listening on 127.0.0.1 inside the container, unreachable through Docker's port mapping > - `onboard --yes` without an explicit `--bind` prefers trusted-local quickstart defaults: it writes a loopback bind into the instance config and ignores the deployment env vars the harness passes, and that config outranks `HOST` at runtime > - This pull request pins the smoke container to the `lan` bind preset and adds a wiring test for it > - The benefit is a working nightly gate, verified end to end against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `docker/Dockerfile.onboard-smoke`, `scripts/__tests__/release-verify-workflow.test.mjs`. **Problem or motivation** Nightly run 31428558684 failed in smoke with the server unreachable at the mapped port for the full 420 second budget. The container logs (captured thanks to #11187) show a fully booted server with `Bind loopback (127.0.0.1)`. The harness sets `HOST=0.0.0.0` and the deployment env vars, but `onboard --yes` without `--bind` deliberately prefers trusted-local defaults, writes `bind: loopback` into the instance config, and the config outranks `HOST` at runtime. A loopback listener inside a container is invisible to the port mapping, so the health check can never pass. This behavior predates the current stable, so the harness was silently broken against every recent version — it only surfaced now because the nightly lane is the suite's first CI consumer. **Proposed solution** Pass `--bind lan` in the smoke container command (the flag is supported by `latest` and canary alike; it selects the all-interfaces preset and keeps the env-driven authenticated deployment), and pin the flag with a wiring test so it cannot regress silently. ## What Changed - `docker/Dockerfile.onboard-smoke`: the onboard command is now `onboard --yes --bind lan --data-dir ...`, with a comment explaining why the flag is load-bearing - `scripts/__tests__/release-verify-workflow.test.mjs`: wiring test asserting the smoke Dockerfile pins a non-loopback bind preset ## Verification - Full local harness run against the real nightly candidate `2026.810.0-canary.1`: container healthy, bind banner shows `lan (0.0.0.0)`, authenticated bootstrap completed (admin created, bootstrap invite accepted, board session verified), `/api/health` returns `bootstrapStatus: ready` - `node --test scripts/__tests__/release-verify-workflow.test.mjs`: 4 pass - After merge: dispatch `release.yml` with `channel: nightly` to run the gate end to end in CI ## Risks - Low. The change only affects the smoke container. `--bind lan` inside a container exposes the port to the container network only; reachability from outside still goes through Docker's explicit port mapping ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI log forensics, upstream source tracing, local Docker reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.4 nightly/v2026.810.0-nightly.0 |
||
|
|
30f6999cbe |
fix(release-smoke): configurable readiness timeout and diagnostics for slow containers (#11187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane (#11006) gates every nightly publish on the release smoke suite, which boots the published artifact in a Docker container > - The suite's first CI execution failed at the health readiness check: the harness hard-codes a 90 second budget, but a CI container cold-installs paperclipai from npm and initializes embedded postgres with no warm caches > - When the timeout expired with the container still running, the harness printed no container logs, so the failure gave no diagnostics > - This pull request makes the readiness budget configurable, raises it for CI, and dumps container logs on timeout > - The benefit is that the nightly gate measures the artifact, not the runner's cold caches, and a red smoke run is diagnosable from its logs ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `scripts/docker-onboard-smoke.sh`, `.github/workflows/release-smoke.yml`. **Problem or motivation** Run 31426044332 (first forced nightly after #11006) failed in `smoke_nightly` with `server did not become ready at http://localhost:3232/api/health` after exactly 90 seconds. The harness's readiness window is hard-coded to 90 attempts at 1 second. Locally that works because the npm cache is warm; in CI the container downloads the full package set and embedded postgres first. The timeout path also printed no container logs when the container was still running, so there was no way to see how far boot had progressed. **Proposed solution** Make the readiness budget an environment variable (`SMOKE_READY_TIMEOUT_SECONDS`, default unchanged at 90 for local use), set it to 420 in the CI workflow, and dump the last 150 container log lines when the readiness check times out on a still-running container. ## What Changed - `scripts/docker-onboard-smoke.sh`: `SMOKE_READY_TIMEOUT_SECONDS` env var (default 90) replaces the hard-coded readiness budget; timeout with a still-running container now prints the tail of `docker logs` - `.github/workflows/release-smoke.yml`: sets `SMOKE_READY_TIMEOUT_SECONDS=420` for CI runs ## Verification - `bash -n` on the harness and YAML parse of the workflow - The real proof is the next `channel: nightly` dispatch of `release.yml`, which re-runs this suite in CI with the new budget ## Risks - Low. The local default is unchanged; CI runs simply wait longer before declaring failure, and a genuinely broken artifact still fails (with logs now) ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use. Diagnosis from CI run logs; patch model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.3 |
||
|
|
5ca752dc81 |
fix(server): raise company import zip upload limit to 1 GB and make it operator-configurable (#11184)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Company import/export lets an operator move a full company package between instances, with the Import page uploading the package as one compressed `.zip` > - The server caps that upload at 128 MB, and real company packages with attachments now exceed it — imports fail at the preview step > - The failure message tells the user to use the CLI folder import, but that path posts inline JSON capped at 64 MB, so the advice is a dead end for exactly these packages > - This pull request raises the zip upload cap to a 1 GB default, makes it operator-configurable through an environment variable, scales the decompression-bomb guards from the cap in effect, and replaces the misleading hint > - The benefit is that large real-world company packages import successfully, and operators with unusual needs can tune the cap without a code change ## Linked Issues or Issue Description **What happened?** A company import fails at the preview step with `Preview failed: Import package exceeds 134217728 bytes`. The package is a valid Paperclip export. Its compressed size is larger than the 128 MB server cap (one reported package is 257 MB). The error panel suggests the CLI folder import, but that path sends the package as one inline JSON body capped at 64 MB, so it also fails. **Expected behavior** A valid company package of realistic size imports successfully through the Import page. If a package is too large, the error must state the limit clearly and suggest a step that can work. **Steps to reproduce** 1. Export a company with enough attachments to make the compressed package larger than 128 MB. 2. Open the Import page and upload the `.zip`. 3. Click "Preview import". 4. The preview fails with `Import package exceeds 134217728 bytes`. **Deployment mode** Reported from a managed deployment; the limit applies to all deployment modes. ## What Changed - Raise `PORTABLE_ZIP_UPLOAD_LIMIT_BYTES` from 128 MB to a 1 GB default (`server/src/http/body-limits.ts`). - Add the `PAPERCLIP_IMPORT_ZIP_MAX_BYTES` environment override. Invalid or non-positive values fall back to the default. - Scale the zip decompression-bomb guard from the configured cap at the import route: the aggregate inflated ceiling is 4x the cap. The per-entry ceiling stays at 512 MB because V8's string length limit applies to an entry regardless (`server/src/routes/companies.ts`, `packages/shared/src/portability-zip.ts`). - Report the 422 limit error in MB instead of raw bytes. - Replace the "use the CLI folder import for very large packages" hint on preview failure with advice that works: re-export the package without large attachments (`ui/src/pages/CompanyImport.tsx`). - Update the stale comment in `ui/src/lib/import-preflight.ts` that made the same CLI claim. - Add tests for the new default, the env override, and the invalid-override fallback. ## Verification - `pnpm vitest run server/src/__tests__/body-limits.test.ts packages/shared/src/portability-zip.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/__tests__/company-portability-import-batching.test.ts` — all pass. - `pnpm vitest run ui/src/pages/CompanyImport.test.tsx` — passes, including the updated failure-panel copy assertion. - `pnpm typecheck` — clean across the workspace. - Manual: upload a `.zip` larger than the configured cap; the preview fails with `Import package exceeds the 1024 MB upload limit` and the new hint. A package between 128 MB and 1 GB now previews and imports. ## Risks - Peak per-import memory rises with the cap: the upload is buffered in memory and unzipped in one pass. A 1 GB compressed package can use several GB transiently. Imports are instance-admin actions, so the exposure is a deliberate operator action, not anonymous traffic. Operators on small hosts can lower the cap with `PAPERCLIP_IMPORT_ZIP_MAX_BYTES`. - The aggregate bomb guard moves from a fixed 512 MB to 4x the configured cap. It still bounds expansion far below what a decompression bomb needs. - No migration and no API shape change. The 422 message text changes; no code matches on the old text. ## Model Used - Claude Fable 5 (`claude-fable-5`) via Claude Code CLI, with extended thinking and tool use (file edits, local test runs, live-instance inspection). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.2 |
||
|
|
f9173782cd |
feat(release): add smoke-gated nightly channel and lane-separated Docker tags (#11006)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem publishes the `paperclipai` npm package set and the Docker images on two lanes: canary on every master push, and stable on manual promotion > - There is no middle ground between those lanes. Users must track every merge or wait weeks for a stable. Docker `:latest` also tracks master, so Docker users have no stable image at all > - A calm prerelease lane needs to exist, and it must never ship a build that failed its checks > - This pull request adds the nightly channel: a scheduled job that selects the newest master commit with a green canary publish, runs the full release smoke suite against that exact published canary, and only then republishes it as the nightly. It also separates Docker tags by lane, so `:latest` finally means stable > - The benefit is that users can follow prereleases at a nightly cadence with a smoke-tested guarantee, and Docker users get real `:canary`, `:nightly`, and stable image tags ## Linked Issues or Issue Description **Subsystem affected** Release automation: `scripts/release.sh`, `scripts/release-lib.sh`, `.github/workflows/release.yml`, `.github/workflows/docker.yml`, `.github/workflows/release-smoke.yml`. **Problem or motivation** The project publishes only `canary` (every master push) and `latest` (manual stable). Users who want prereleases without per-merge churn have no option. Docker has a second problem: master builds overwrite `:latest`, and CI-published stables never produced Docker images, because tags pushed with `GITHUB_TOKEN` do not fire the `v*` tag trigger in `docker.yml`. No stable-versioned image exists in ghcr today. **Proposed solution** Add a `nightly` channel. A scheduled job selects the newest canary-tagged master commit, smoke-tests that exact published canary, and republishes the same commit as `YYYY.MDD.P-nightly.N` under the `nightly` dist-tag. Separate Docker tags by lane (`:canary` for master, `:nightly` for nightly tags, `:latest` plus version tags for stable tags only), and have the release jobs dispatch `docker.yml` at the new tag so lane images actually build. **Alternatives considered** Moving the `nightly` dist-tag to the existing canary version without a republish. Rejected: the version string would say `canary` while the user is on nightly, which breaks at-a-glance lane identification in bug reports and `--version` output. ## What Changed - `scripts/release-lib.sh`: channel-parameterized `next_prerelease_version` and `prerelease_tag_name` helpers (canary helpers delegate to them), a `require_channel_tag_at_head` guard, and the no-provenance retry for Sigstore transparency-log duplicates now covers the `nightly` dist-tag as well as `canary` - `scripts/release.sh`: new `nightly` channel. It requires HEAD to carry a `canary/v*` tag, publishes the full public package set as `YYYY.MDD.P-nightly.N` under dist-tag `nightly`, and tags the source commit `nightly/vYYYY.MDD.P-nightly.N` - `.github/workflows/release.yml`: scheduled nightly chain (09:00 UTC) — select candidate, smoke it via `release-smoke.yml`, publish on green under the existing `npm-canary` environment, push the tag, dispatch `docker.yml`. New `channel` dispatch input (default `stable`, so existing stable dispatches are unchanged) with `nightly_source_version` and `dry_run` support for forced runs. The stable path now also dispatches `docker.yml` at the new `v*` tag - `.github/workflows/docker.yml`: lane tag mapping for both image jobs — master pushes publish `:canary` and no longer move `:latest`; `nightly/v*` tags publish `:nightly`; only stable `v*` tags publish `:latest` and the versioned tags. New `workflow_dispatch` trigger for the release-job dispatches. Build-version stamping uses the exact nightly version on nightly tag builds - `.github/workflows/release-smoke.yml`: `nightly` added to the dispatch choice list - `doc/CHANNELS.md` (new): user-facing guide to the channels - `doc/RELEASING.md`: nightly lane documentation, Docker tag mapping table, and a nightly failure playbook - `doc/RELEASE-AUTOMATION-SETUP.md`: note that nightly reuses `npm-canary` and needs no npm trusted-publisher changes - Tests: channel-parameterized version helper coverage in `scripts/release-registry-versions.test.mjs`, and nightly flow coverage (publish identity, notes not required, canary-tag guard) in `scripts/__tests__/release-dry-run-notes.test.mjs` ## Verification - `node --test` on the release script suites: 68 pass, including 6 new tests. The only failure, `acpx-patch-packaging.test.mjs`, needs installed `node_modules` and fails identically on a pristine checkout of master in the same environment - `bash -n` on both shell scripts and YAML parse of all three workflows - Live fail-path check: `./scripts/release.sh nightly --print-version` from a master tip with no canary tag fails with `HEAD has no canary/v* tag` - Live success-path check: the same command from the `canary/v2026.806.0-canary.7` commit prints `2026.806.0-nightly.0` - Live selection check: the candidate-selection shell logic run against the real repository selects the commit of `canary/v2026.806.0-canary.7`, which matches the current npm `canary` dist-tag exactly - After merge: dispatch `release.yml` with `channel: nightly` and `dry_run: true` to preview, then a real forced run to validate end to end before the first scheduled run ## Risks - Docker `:latest` changes meaning from "latest master build" to "latest stable release". This is deliberate and will be announced. Users who want the old behavior pull `:canary`. Until the first stable release after this change, `:latest` stays at its current (master-built) image - The nightly is a rebuild of the same source commit, not the byte-identical canary artifact that was smoked. The lockfile pins dependencies, and the publish path's registry-visibility and clean-prefix install gates still run on the nightly artifacts - All npm publishing must stay inside `release.yml` because npm trusted publishing pins that workflow file per package. The nightly jobs were added to `release.yml` for exactly that reason; this constraint is now documented in `RELEASING.md` - The stable-lane Docker dispatch fails gracefully (a warning with manual instructions) when the source ref predates `docker.yml`'s `workflow_dispatch` trigger ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (repository exploration, local test execution, live registry and git verification). All code, tests, and docs in this PR were model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.810.0-canary.1 |
||
|
|
6a4e2e1b8c |
fix(routes): return 409 for routine checkout conflicts (#3790)
## Thinking Path > - Paperclip orchestrates AI agents and relies on issue checkout as the core task-claiming primitive > - The issue checkout route is the HTTP boundary that translates service and database outcomes into agent-usable API responses > - Routine-linked issues are protected by the partial unique index `issues_open_routine_execution_uq`, which covers only rows whose `execution_run_id` is set > - `svc.checkout` sets `execution_run_id`, so a concurrent claim moves the row into that index and can raise a 23505 mid-request > - Unhandled, that surfaces as a 500 and crashes the agent run instead of being a recoverable conflict > - Drizzle wraps driver failures in its own `Failed query: ...` error, so the Postgres error carrying `code` and the constraint name is reachable only through `cause` > - This pull request translates that violation into a 409 at the checkout route, detecting it through the cause chain the way `isReviewPathRecoveryIdempotencyConflict` already does > - The benefit is that agents handle routine execution contention through the normal heartbeat conflict path instead of failing on an internal server error ## Linked Issues or Issue Description Fixes #3660 Related pull requests found while searching for duplicates: - #3699 — an earlier attempt at this same route-level fix, closed unmerged. Same shape, and its check has the flat-error bug described under Verification. - #3633 — related work on postgres.js `constraint_name` handling in conflict detection. - #5662 — covers the adoption path (`assertCheckoutOwner`) that this pull request does not. ## What Changed - Added `server/src/db-errors.ts` with `isUniqueViolation(error, constraintName?)`, which walks the `cause` chain (depth-capped) and accepts the postgres.js `constraint_name`, the node-postgres `constraint`, or the driver message as evidence of SQLSTATE 23505. - Wrapped `svc.checkout()` in `POST /issues/:id/checkout` with a narrow try/catch that uses that helper to return **409 Conflict** for `issues_open_routine_execution_uq`, and rethrows every other error unchanged. - Added `server/src/__tests__/db-errors.test.ts` covering the wrapped and bare error shapes, both constraint field names, the message fallback, non-matching constraints, non-unique-violation codes, and a self-referential cause chain. ## Verification - The new unit test includes the wrapped case `{ cause: { code: "23505", constraint_name: ... } }` that a flat `error.code` check fails, so it is a real regression guard rather than a restatement of the implementation. - The wrapped shape is what this codebase observes in practice: `server/src/__tests__/plugin-tenant-isolation.test.ts` asserts `cause?.code === "23505"` against embedded Postgres, `packages/db/src/pipelines-schema.test.ts` asserts that constraint failures throw `Failed query`, and `server/src/services/recovery/review-path-recovery.ts` walks the same chain. - CI (verify, e2e, policy) exercises this change against current master through the pull request merge ref. - Not verified locally: no monorepo install or typecheck was run in this environment. ## Risks - Low. One route gains a catch that matches a single constraint and rethrows all other errors, so no unrelated failure can be swallowed. - The 409 body `{ error: ... }` matches the other 409 responses this route already returns. - Scope limit: this covers the checkout route only. The adoption path reached through `assertCheckoutOwner` (heartbeat, plugins, and pipelines routes) can still surface the same violation as a 500; #5662 targets that path. - `isUniqueViolation` is new and intentionally generic. Existing flat 23505 checks elsewhere in the server are left untouched by this pull request. ## Model Used - Original change: OpenAI Codex, GPT-5-class tool-using coding agent in the Codex CLI environment; exact backend model revision is not exposed in that runtime. - Follow-up revision (cause-chain detection plus tests): Anthropic Claude Opus 5 (`claude-opus-5`), tool-using coding agent with extended thinking and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>canary/v2026.810.0-canary.0 |
||
|
|
ebf2b8ff79 |
fix(server): persist worktree runtime port when ambient PORT does not match (#1849) (#1930)
## What was done Replaced the strict `!nonEmpty(process.env.PORT)` guard in `maybePersistWorktreeRuntimePorts` with a new `isPortPinnedByRuntimeEnv` helper function. This function checks if `process.env.PORT` is set, but only suppresses persisting the port to configuration if the ambient `PORT` matches the newly allocated `selectedPort`. ## Why it matters Fixes issue #1849. Previously, if an ambient `PORT` environment variable was exported globally (like inheriting from the shell running the parent workspace), worktrees would silently fail to write their collision-avoiding ports (e.g. 3103 instead of 3100) back to their respective local `config.json` files. This resulted in orphaned sub-worktrees and lost port tracking on reboot. With this fix, worktrees correctly persist their assigned ports even while nested under an inherited environment variables stack, while continuing to respect manual, explicit pinning. ## How to verify 1. Export a port in the shell explicitly: `export PORT=3100`. 2. Launch a sub-worktree instance which receives an auto-assigned free port (e.g., `3103`). 3. View the underlying `config.json` for that worktree inside `.paperclip/worktrees/`. 4. The config file should correctly contain `{"server": {"port": 3103}}` rather than dropping the write operation. ## Risks None expected. The `Number()` and `Number.isInteger()` checks handle parsing edge cases cleanly, defaulting robustly to preventing writes if `process.env.PORT` is somehow malformed (e.g., set to a non-integer), ensuring absolute safety during misconfigurations. Co-authored-by: manavshrivastavagit <manavshrivastava@users.noreply.github.com>canary/v2026.809.0-canary.0 |
||
|
|
19be4cf927 |
refactor(a11y): add scope=col to the agent costs table headers (#1789)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page has a Costs section. That section renders a data table of per-run spend. > - The table has a header row, but its `<th>` elements carry no `scope` attribute. > - A screen reader uses `scope="col"` to bind each data cell to its column header. Without it, the reader announces a number without telling the user which column it belongs to. > - A table of costs is exactly the case where that hurts. Every cell is a bare figure. > - This pull request adds `scope="col"` to the five header cells in that table. > - The benefit is that assistive technology announces the cost table correctly. The change is markup only, so sighted users see no difference. ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR, following the enhancement template. **What existing behavior does this improve?** The Costs table rendered by `CostsSection` in `ui/src/pages/AgentDetail.tsx`. **Subsystem affected** ui/ — React + Vite board UI **Current behavior** The table renders five header cells: Date, Run, Input, Output, and Cost. None of them set `scope`. A screen reader must guess the header-to-cell relationship, so a user hears a value with no column name attached to it. **Proposed behavior** Each header cell sets `scope="col"`. A screen reader then announces the column name together with each cell, so a cost figure is read as part of the Cost column. **Reason and benefit** `scope` is the standard way to associate header cells with data cells in an HTML table. The attribute has no visual effect, so the fix carries no design cost and makes the table usable with a screen reader. **Breaking changes** None. `scope` is a presentational-neutral HTML attribute. No component API, no styling, and no test changes. **Related pull requests** - #2215 proposed the same attribute for the Routines table. It is closed, because that table no longer exists on master. - #1524 and #1522 applied `scope="col"` to other tables. Both are closed. ## What Changed - Added `scope="col"` to the five `<th>` elements in the `CostsSection` table in `ui/src/pages/AgentDetail.tsx`. - Rebased the branch onto current master. - Dropped the original `ui/src/pages/Routines.tsx` hunks. Master rebuilt the Routines page around folder-grouped rows, so the table those hunks targeted no longer exists. - Dropped the original `HintIcon` opacity change. It altered a visible colour, which is out of scope for a markup-only accessibility fix. ## Verification - Run `pnpm --filter @paperclipai/ui typecheck` and `pnpm --filter @paperclipai/ui build`. This is a markup-only change, so a clean type-check and build is the relevant automated signal. - Open an agent detail page and go to the Costs section. Inspect the header row. Each `<th>` now carries `scope="col"`. - Navigate the same table with a screen reader, cell by cell. Each cell is announced with its column name. - Compare the rendered page before and after. It is unchanged, because `scope` has no styling effect. ## Risks Low risk. The change adds one standard HTML attribute to five header cells in a single table. It introduces no code path, changes no component API, and has no visual effect. The worst case is that the attribute is redundant for a reader that already infers the column, which is harmless. ## Model Used Anthropic Claude Opus 5, exact model ID `claude-opus-5`. It ran with extended thinking and repository read/write tools, inside a maintainer-operated triage agent. The model rebased the branch, dropped the two out-of-scope hunks, and wrote this description. The original change was authored by @bluzername, and the model used for that work is not recorded here. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` references) - [x] I have considered and documented any risks above - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass — not run. This is a markup-only change and the package has no test covering this table. - [ ] I have added or updated tests where applicable — no test added, which is why this PR is titled `refactor:`. - [ ] I have updated relevant documentation to reflect my changes — no documentation describes this markup. - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work — not checked by the maintainer who rebased this. - [ ] All Paperclip CI gates are green — CI re-runs on this push. - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — Greptile re-reviews on this push. --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>canary/v2026.808.0-canary.4 |
||
|
|
cc35c3c395 |
feat: structure and humanize recovery notices (#11075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip posts system comments when automatic run recovery cannot continue > - These comments currently mix the main event with recovery identifiers and routing details > - The task chat shell also renders these comments as large raw text blocks > - Operators need a short explanation first and inspectable evidence on demand > - This pull request emits structured recovery notices and renders them as compact humanized rows > - The benefit is a quieter task thread that keeps the full recovery evidence available ## Linked Issues or Issue Description Related prior extraction source: #11070. This pull request replaces only its structured recovery notice slice with a focused branch based on current master. **What existing behavior does this improve?** Paperclip recovery escalations and the experimental task chat system-comment renderer. **Current behavior** Recovery escalation comments put action identifiers, owner details, run details, and failure codes into the visible markdown body. The task chat shell renders the complete system comment as a large text block. **Proposed behavior** The server emits a short system notice with typed metadata sections. The task chat shell classifies known recovery families and renders one compact row. An operator can expand the row to inspect the full body and metadata. **Reason and benefit** The main thread stays readable during repeated recovery activity. Typed links and evidence remain available without exposing raw failure text in the default view. **Breaking changes** The visible recovery comment body is shorter. Recovery action deduplication now reads the structured metadata and still recognizes legacy body markers. No API schema or database migration changes. ## What Changed - Emit stranded recovery escalations with `system_notice` presentation and typed recovery, owner, run, and failure-code metadata. - Share bounded metadata row builders across recovery notice producers and preserve legacy deduplication compatibility. - Humanize known recovery notice families and render compact expandable task-chat rows. - Route system-authored comments ahead of derived agent authorship so recovery notices do not appear as agent bubbles. - Add focused server and UI regression coverage. ## Verification - `pnpm check:token-gates` — 3/3 clean. - `pnpm --filter @paperclipai/server typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/shared exec vitest run src/validators/issue.test.ts` — 32 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/stranded-notice.test.ts src/__tests__/issue-recovery-actions.test.ts` — 57 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/services/recovery/successful-run-handoff.test.ts src/services/recovery/stranded-notice.test.ts` — 39 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'escalates an exhausted failed successful-run handoff without using generic continuation recovery first|escalates an exhausted successful handoff run that still leaves no disposition'` — 2 tests passed. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts -t 'blocks assigned todo work after the one automatic dispatch recovery was already used'` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/lib/system-notice-humanizer.test.ts src/components/task-chat/TaskChatSystemNotice.test.tsx src/components/task-chat/task-chat-adapter.test.ts` — 15 tests passed. - Storybook visual baselines were not updated because this chat-shell path has no affected snapshot baseline. Focused rendering tests and token gates cover this change. ## Risks - Consumers that parse recovery action identifiers from comment markdown must move to structured metadata. Server deduplication remains backward compatible with legacy comments. - The humanizer uses stable recovery-family phrases. Unknown notices use a generic truncated first-sentence fallback. - The UI changes only the experimental task chat presentation. The stored comment body and expanded metadata remain available. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5, reasoning mode, repository tools, shell execution, and GitHub integration. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4e9a78db58 |
feat(ui): persist task chat composer drafts (#11076)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People send task instructions through the board task chat. > - A page refresh or task switch can discard an unfinished message in the redesigned composer. > - The existing task chat already supplies a task-specific draft key. > - The redesigned composer must use that key without changing attachment or send behavior. > - This pull request restores, saves, and clears text drafts in the redesigned composer. > - The benefit is that users can return to unfinished task messages without losing their text. ## Linked Issues or Issue Description Related prior work: #11070. This pull request extracts only the final composer draft behavior from that larger draft. **Subsystem affected** ui/ — React + Vite board UI. **Problem or motivation** The redesigned task chat composer does not use the draft key that the task thread already provides. A refresh, navigation, or unmount can lose an unfinished message. **Proposed solution** Persist text drafts by task key in local storage. Restore a draft when the composer mounts. Save changes after a short delay and flush pending text during unload or unmount. Clear the draft only after a successful send. **Alternatives considered** The composer could save on every keystroke. A short delay avoids unnecessary synchronous storage writes. The feature could also stay in the larger predecessor PR, but a focused PR is easier to review and verify. **Roadmap alignment** This is a focused usability improvement for the task conversation surface. It does not add or duplicate a roadmap capability. ## What Changed - Added safe draft storage helpers for load, save, and clear operations. - Connected the task-specific draft key to the redesigned task chat composer. - Preserved drafts across debounce windows, unmounts, page unloads, failed sends, and React Strict Mode probes. - Cleared drafts after successful sends without changing current attachment safeguards. - Added focused composer and thread integration tests. ## Verification - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatComposer.test.tsx src/components/TaskChatThread.test.tsx` ## Risks - Local storage can be unavailable or full. The helpers catch storage errors and keep the composer usable. - Only text is persisted. Attachments, work mode, and assignee selections remain session state. - Draft keys remain task-scoped, so text does not cross task boundaries. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with model `gpt-5`. The context-window size is not exposed in this environment. The model used agentic reasoning, tool use, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.808.0-canary.3 |
||
|
|
34fe57a024 |
fix(server): ignore sibling worktrees in dev watch (#11074)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Developers can run Paperclip from linked Git worktrees. > - The server development watcher scans paths near the active checkout. > - A main checkout can contain many complete sibling worktrees under `.paperclip/worktrees`. > - Scanning those sibling checkouts can stall the watcher before it starts the server. > - This pull request excludes the shared worktree directory from the development watcher. > - The benefit is that development startup stays responsive as the number of worktrees grows. ## Linked Issues or Issue Description **What happened?** The server development watcher traversed sibling checkouts under `.paperclip/worktrees`. Large worktree collections could make `pnpm dev` stall before the watcher started the server process. **Expected behavior** The watcher must observe only source paths that can reload the active checkout. It must ignore sibling worktrees in both a main checkout and a linked worktree. **Steps to reproduce** 1. Create several linked worktrees under `.paperclip/worktrees`. 2. Add normal dependency and build output trees to those worktrees. 3. Run `pnpm dev` from the main checkout or one linked worktree. 4. Observe the watcher scan sibling worktrees before it starts the server. **Paperclip version or commit** Reproduced on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Detect whether the active server root is inside the managed linked-worktree directory. - Ignore the shared `.paperclip/worktrees` root from both main and linked checkouts. - Add regression coverage for the resolved ignore path and its globstar form. ## Verification - `./node_modules/.bin/vitest run server/src/__tests__/dev-watch-ignore.test.ts --reporter=verbose` - `pnpm --filter @paperclipai/server typecheck` ## Risks - Low risk. The change affects only local development watch exclusions. - A non-standard checkout that copies the same `.paperclip/worktrees` directory layout will receive the same exclusion. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5, context window not disclosed, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1791979512 |
refactor(a11y): add ARIA progressbar attributes to BudgetPolicyCard (#1805)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Humans oversee those agents in teams, so each team needs its own spend controls > - The BudgetPolicyCard component shows how much of a budget is consumed > - The utilization bar in that card is a styled div with no ARIA role, value, or label > - A screen reader user therefore cannot hear how much budget is used > - This pull request adds role="progressbar" and the matching ARIA value attributes > - The benefit is that assistive technology announces budget utilization the same way sighted users see it ## Linked Issues or Issue Description No existing issue. Related pull requests: #1869 and #1878 add the same progressbar semantics to ProviderQuotaCard and QuotaBar. They touch different files. The problem follows the bug report template: **What happened?** The budget utilization bar in `BudgetPolicyCard` renders as a plain `div`. It has no `role`, no `aria-valuenow`, and no accessible name. A screen reader announces nothing for it. The user can read the "Remaining" amount, but not the utilization percentage. **Expected behavior** The bar is announced as a progress bar. It reports the current utilization percentage, with its minimum and maximum. **Steps to reproduce** 1. Open a project or agent page that shows the budget card. 2. Start VoiceOver (Cmd+F5 on macOS). 3. Move the cursor to the budget utilization bar. 4. VoiceOver announces nothing. **Paperclip version or commit** master, `ui/src/components/BudgetPolicyCard.tsx`. **Deployment mode** Local dev (pnpm dev). ## What Changed - Added `role="progressbar"` to the inner bar element. - Added `aria-valuenow` with the rounded utilization percentage. - Added `aria-valuemin={0}` and `aria-valuemax={100}`. - Added `aria-label` in the form `Budget utilization: 73% used`. `aria-valuenow` and `aria-label` use the same `progress` value. That value is already capped at 100 by `Math.min(100, summary.utilizationPercent)`, so the reported value stays inside the min/max range when a scope is over budget. An earlier revision also changed the budget amount `Input` to `type="number"`. That change is removed. It changed input behavior and was not an accessibility fix. See "Risks". ## Verification 1. Open a page that shows the budget card. 2. Start VoiceOver (Cmd+F5 on macOS) and move to the utilization bar. 3. VoiceOver announces "Budget utilization: X% used, progress indicator". 4. Inspect the element. `aria-valuenow` equals the displayed percentage, and it stays at 100 when utilization is above 100%. The change adds attributes only. There is no visual change. ## Risks Low risk. The change adds ARIA attributes to one element in one component. It changes no logic, no styling, and no layout. The `type="number"` change is removed on purpose. With `type="number"`, a browser reports an empty string for text it cannot parse. `parseDollarInput("")` then returns `0` instead of `null`, so the existing "Enter a valid non-negative dollar amount." error never appears and unparseable input silently becomes a $0.00 budget. The `inputMode="decimal"` input keeps that validation path working. ## Model Used Original implementation by @bluzername. The author did not state a model. Scope reduction, branch update, and this description: Claude Opus 5 (1M context, extended thinking, tool use), prepared under maintainer review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change - [x] I have considered and documented any risks above - [ ] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>canary/v2026.808.0-canary.2 |
||
|
|
d5208d30c1 |
feat(adapter-utils): add host-side pack span to managed-runtime tarball build (#11072)
## Thinking Path > - Paperclip runs AI agents through adapter execution services. > - The adapter runtime records spans for each stage of agent startup. > - The managed runtime packs workspace tarballs before it uploads them. > - These pack operations had no host span, so `stage.sync` omitted pack time. > - This pull request adds one `pack` span around both tarball builds. > - The span nests under the active `stage.sync` step and improves trace detail. ## Linked Issues or Issue Description **What existing behavior does this improve?** The managed runtime workspace sync builds a git-history tarball and a workspace-overlay tarball before upload. **Subsystem affected** The change affects `packages/adapter-utils`, which provides adapter execution and managed runtime support. **Current behavior** The host builds both tarballs without an OpenTelemetry span. The `stage.sync` trace therefore omits the host pack duration. **Proposed behavior** The host wraps both tarball builds in one `pack` span. The executor parents this span under the active startup step. **Reason and benefit** The trace shows the time that the host spends packing workspace data. Operators can use the existing runtime span tree to find sync delays. **Breaking changes** None. The default span runner remains a no-op runner, and the existing control flow remains unchanged. ## What Changed - Add an optional `runtimeSpan` runner to the managed runtime preparation path. - Create one host `pack` span around the two workspace tarball builds. - Parent the `pack` span under the active startup step. - Add unit coverage for span emission and span nesting. ## Verification - Run `tsc --noEmit` for `@paperclipai/adapter-utils`. - Run the `@paperclipai/adapter-utils` Vitest suite. - Confirm the suite reports 448 passed tests and 4 skipped tests. - Confirm the trace test records `pack` under `stage.sync`. ## Risks The change adds optional tracing only. The default no-op runner preserves behavior when tracing is not configured. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The model reviewed the handoff and opened this pull request. The implementation author supplied the commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.808.0-canary.1 |
||
|
|
677242344c |
chore(lockfile): refresh pnpm-lock.yaml (#11068)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.808.0-canary.0 |
||
|
|
384e5f6178 |
revert(ui): back out onboarding port (#10786) (#11067)
This reverts commitcanary/v2026.807.0-canary.16 |
||
|
|
6b7e0814a0 |
feat(acp): stream Daytona sandbox agent output and remove the host output poll (#11049)
## Thinking Path > - Paperclip is the open source app that manages AI agents for work > - Sandbox providers let agents run in remote and isolated environments > - Daytona session commands need a path that sends agent output to the host without host polling > - Host polling adds delay and repeats provider output work > - This pull request adds typed execute.log notifications and a log sink for incremental output > - This pull request adds an optional ACP session stream with final-result replay protection > - The benefit is lower output delay while the default flags keep current behavior unchanged ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting. The change spans the plugin SDK, Daytona provider, adapter utilities, and server execution services. **Problem or motivation** The Daytona ACP bridge polls a host output file while an agent command runs. This adds delay and can repeat work. The host also needs a safe route for provider output chunks. **Proposed solution** Add a typed `execute.log` notification with host-issued invocation correlation. Add an ordered log sink to the environment execute path. Add an optional ACP session-log path that parses newline-delimited JSON frames and removes the host output poll for that path. **Alternatives considered** Keep the output-file poll as the only path. This keeps the current behavior but does not provide timely output. The new path stays behind flags, so the existing path remains the default fallback. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`, including Daytona support. ## What Changed - Add the typed `execute.log` worker-to-host notification and company-scoped host route. - Add ordered `stdout` and `stderr` chunk delivery before the final execute result. - Add the Daytona session log sink and the optional ACP streamed session path. - Add monotonic frame handling so live and final output reach the host once. - Keep `useLogStream` and `streamAgentSessionOutput` off by default. - Add unit and integration coverage for the notification, execution target, runtime, and Daytona paths. ## Verification - Run adapter-utils tests: 445 tests pass locally. - Run server environment tests: 73 tests pass locally. - Run Daytona plugin tests: 131 tests pass locally. - Run TypeScript checks for shared, adapter-utils, and server. - Review the pull request checks after GitHub completes them. - All required GitHub checks pass on the current head. ## Risks The new paths change output delivery only when a feature flag enables them. The final execute result remains available for parsing and fallback. The main risk is a provider stream or frame-order error; the final-result parser limits that risk. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The runtime did not supply a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes No operator documentation change applies because both new flags remain disabled by default. - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.15 |
||
|
|
0a511ed1b0 |
feat(apps): support multiple provider connections (#11060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem connects company tools through governed provider connections. > - A company can need more than one account for the same provider. > - The current database constraint and Apps flow assume one named connection per company. > - New quarantined actions also need an explicit review decision before activation. > - This pull request supports multiple provider connections and complete action review decisions. > - The benefit is safer access control and a clear multi-account Apps workflow. ## Linked Issues or Issue Description Refs: #11040 **Subsystem affected** Cross-cutting. This change affects the Apps UI, the tool access API, the shared request contract, and the database schema. **Problem or motivation** The connection name constraint prevents a company from keeping more than one connection for a provider. The Apps UI also reuses an existing OAuth connection when a user asks to connect another account. Action review can enable selected entries without recording a decision for every quarantined action. **Proposed solution** Remove the company and connection name uniqueness constraint. Let users open, count, edit, and create multiple provider connections. Require the finish request to cover every quarantined action exactly once before the server activates reviewed entries. **Alternatives considered** The UI could generate unique internal names and keep the database constraint. This would preserve a one-connection assumption in the data model and would make display names part of identity. The server could also infer review decisions from enabled actions. This would not distinguish a reviewed disabled action from an action that the user did not review. **Roadmap alignment** This change extends the completed MCP Tool Gateway and Apps milestone. It also supports the Connected Apps roadmap item. It follows the navigation and connection management work in #11040. ## What Changed - Remove the company-scoped connection name uniqueness index with an ordered and idempotent migration. - Add a reviewed action list to the finish-app contract and reject incomplete or duplicate review decisions. - Activate reviewed entries and keep unreviewed quarantined entries blocked. - Enable a completed connection and preserve the company and connection scope in all updates. - Show provider connection counts and open the provider setup page from Browse. - Let users edit existing connections or connect another account without reusing an active OAuth connection. - Update focused server and UI coverage for multiple connections and action review. ## Verification - Ran the focused Apps UI suite. All 116 tests passed in 11 files. - Ran the focused server and CLI suite. All 276 tests passed in 3 files. - Ran `pnpm --filter @paperclipai/db check:migrations`. The migration safety check passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One worktree-safety assertion failed because the execution workspace reloads its worktree marker. The same test passed with an isolated non-worktree marker. - Ran `pnpm check:token-gates`. It reports 12 existing violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. - Started the six affected Playwright specifications. Chromium could not start because the host does not provide `libatk-1.0.so.0`. The GitHub e2e jobs will verify these specifications. - GitHub Actions passed every final-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed final commit `9af9200426` at 5/5 with zero review threads. ## Risks - Removing the name uniqueness index permits duplicate display names. Stable connection IDs and UIDs remain unique within a company. - The finish-app endpoint accepts the new review field as optional for backward compatibility. When clients send it, the server requires a complete decision for all quarantined actions. - Multiple OAuth connections depend on the explicit new-connection route flag. Focused tests cover active and draft connection reuse. - The migration is ordered after migration 0210. Its `DROP INDEX IF EXISTS` statement is safe to repeat. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The agent used repository tools, code execution, test execution, and agentic reasoning. The Codex runtime manages the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.14 |
||
|
|
a71b9cf628 |
feat(skills): add MCP integration preparation skill (#11063)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Skills Store lets a company find and install reusable agent procedures > - The MCP integration preparation procedure existed outside the app catalog > - Paperclip users could not find or install that procedure from the product > - This pull request adds the procedure as an optional software development skill > - The skill keeps research, human approval, and connector delivery as separate gates > - The benefit is a repeatable and governed path from vendor research to one connector pull request ## Linked Issues or Issue Description **What existing behavior does this improve?** The app-shipped Skills Store can install optional skills, but it does not include the MCP integration preparation workflow from `paperclip-content`. **Subsystem affected** `packages/skills-catalog`. **Current behavior** An agent must know where the external workflow lives. The agent cannot find or install it from the Paperclip skills catalog. **Proposed behavior** The optional catalog includes `prepare-mcp-integration`. The installed skill directs agents through cited research, a research-only content pull request, an exact-revision human gate, and one Paperclip connector pull request per approved connection. **Reason and benefit** This change makes the existing integration and connector playbooks available as one installable Paperclip workflow. It also prevents agents from starting connector code before the research gate is approved. **Breaking changes** None. The skill is optional and markdown-only. Related source work: paperclipai/paperclip-content#13. ## What Changed - Add the optional `prepare-mcp-integration` catalog skill under software development - Add Paperclip catalog metadata for roles, requirements, tags, and trust classification - Add a worked Notion MCP research-gate example - Regenerate the checked-in catalog manifest - Add the new key to the shipped optional skill test ## Verification - `pnpm --filter @paperclipai/skills-catalog build:manifest` - `pnpm --filter @paperclipai/skills-catalog validate` - `pnpm --filter @paperclipai/skills-catalog test` - Confirm the generated catalog contains `paperclipai/optional/software-development/prepare-mcp-integration` - Confirm the trust level is `markdown_only` and compatibility is `compatible` ## Risks - Low risk. This change adds one optional markdown-only catalog entry. - The workflow can become stale if the two upstream playbooks change. The skill requires agents to read the current playbooks before each phase and to update upstream rules when reusable specifications change. ## Model Used OpenAI Codex with GPT-5.4, reasoning mode, shell tool use, and code editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.13 |
||
|
|
3435920ae1 |
docs: minimize plan task graphs (#11057)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip skills guide agents through repeatable work. > - The plan-to-task skill guides agents when they create issue graphs from plans. > - The prior guidance could encourage an issue for every concrete deliverable. > - That guidance can create unnecessary subtasks for work that one owner can complete end to end. > - This pull request defines the boundaries that justify a separate task. > - It also adds a merge-back pass that removes task splits without a qualifying boundary. > - The benefit is a smaller issue graph with clear ownership, dependencies, and review gates. ## Linked Issues or Issue Description **Issue type** Unclear or confusing documentation. **Where is the issue?** `skills/paperclip-converting-plans-to-tasks/SKILL.md` **What's wrong?** The skill says that each concrete deliverable must become an issue. This can make agents split one end-to-end job into tasks for each step, file, component, or phase. The result is more coordination work without a real execution boundary. **Suggested fix** Tell agents to start with one end-to-end task. Permit separate tasks only for ownership, parallel work, dependencies, independent review or approval, or substantial follow-up work. Require a merge-back pass before agents create the issue graph. ## What Changed - Added a rule to use the fewest tasks that can complete and verify the work. - Defined the boundaries that qualify work for a separate issue. - Added a merge-back pass for proposed subtasks without a qualifying reason. - Updated dependency, parallel work, verification, and checklist guidance to enforce the smaller graph. ## Verification - Ran `pnpm exec vitest run packages/shared/src/frontmatter.test.ts`. - Result: 1 test file passed and 21 tests passed. - Ran `git diff --check public-gh/master...HEAD`. - Confirmed that the PR changes one skill file and has no lockfile, workflow, image, or migration changes. ## Risks - Low risk. This change updates skill guidance only. - Agents might combine work too aggressively. The qualifying boundaries preserve separate ownership, parallel execution, dependencies, review, approval, and substantial follow-up work. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 through Codex. The exact deployment ID and context window are not exposed. The agent used reasoning, repository tools, command execution, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.12 |
||
|
|
b18b0fc39b |
feat: refine app connections and legacy worktree startup (#11040)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps UI manages app discovery and app connections. > - The managed worktree runtime starts agent work in repository worktrees. > - The Apps routes do not match the main discovery flow, and the connections view lacks a delete action. > - Legacy managed worktrees can also start before their pending seed operation runs. > - This pull request makes app discovery the main Apps route and makes connection management explicit. > - It also seeds legacy managed worktrees before runtime startup and makes the CLI read the repository-local config. > - The benefit is a clearer Apps workflow and a safer managed-worktree startup path. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Apps navigation, app connection management, managed git-worktree startup, and CLI worktree selection. **Subsystem affected** Cross-cutting. The change affects `ui/`, `server/`, `cli/`, and development documentation. **Current behavior** The `/apps` route opens the connections list while discovery uses a nested route. The connections list has no delete action. Some legacy managed worktrees can start runtime work before their pending seed operation runs. The CLI can also read an ambient Paperclip config instead of the repository-local config. **Proposed behavior** The `/apps` route opens Browse, and `/apps/connections` opens the connection list. Users can delete a connection after confirmation. Runtime startup seeds legacy managed worktrees when required. The CLI resolves the current worktree from the repository-local `.paperclip/config.json` file. **Reason and benefit** Users can discover apps from the canonical Apps route and can manage existing connections from a dedicated route. Legacy worktrees receive their required repository content before agent runtime starts. CLI worktree selection stays scoped to the current repository. **Breaking changes** The `/apps` and `/apps/browse` route behavior changes. Old Browse links redirect to `/apps`. The change does not modify an API schema or database schema. ## What Changed - Make Browse the canonical `/apps` page and move the connection list to `/apps/connections`. - Align Apps navigation, redirects, attention links, empty states, and connection actions with the new routes. - Add connection deletion with confirmation and clear failure feedback. - Seed legacy managed git worktrees before runtime startup when their seed status is pending. - Read the CLI worktree selection from the repository-local Paperclip config. - Update focused UI, server, CLI, and development documentation coverage. ## Verification - Ran 202 focused UI, server, and CLI tests. All tests passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. The server and UI stages passed 7,168 tests. The CLI stage found one environment-sensitive secrets test because this workspace injects static AWS credentials. The isolated CLI file passed all 8 tests after those injected variables were unset. - Ran `pnpm check:token-gates`. It reports 12 existing color-token violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. This pull request does not modify that file. - Ran focused regression coverage for repository-root CLI config resolution and connection deletion state. All tests and affected package typechecks passed. - Collected all 27 tests in the six changed Playwright specifications successfully. - GitHub Actions passed every latest-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed the final commit at 5/5 with zero unresolved threads. ## Risks - Existing bookmarks for `/apps/browse` redirect to `/apps`. - Connection deletion changes visible connection state and requires user confirmation. - The legacy seed path runs only for managed git worktrees with pending seed state. Tests cover the startup condition. - The rebase preserves the target branch's direct OAuth policy for the Notion connection flow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the GPT-5 model family assisted this change. The agent used reasoning, repository tools, code execution, and test execution. The runtime does not expose the exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.11 |
||
|
|
42c73562c5 |
fix(heartbeat): backfill projectWorkspaceId when restoring a reused execution workspace (#10171)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Each run executes inside a persisted **execution workspace** (a row
in `execution_workspaces`) that is either freshly created or
**restored/reused** across runs of the same issue
> - Before adapter launch, a guard rejects a restored workspace whose
`projectWorkspaceId` is null while the issue resolves a concrete project
workspace (`persisted_workspace_missing_project_workspace_id`) — a
safety check against binding a run to a workspace with no
project-workspace link
> - The reuse/**restore** path updated the existing row (cwd, branch,
status, metadata…) but **never set `projectWorkspaceId`**, so a row
persisted with a null value stayed null on every restore
> - Result: for an issue that resolves a project workspace, the guard
fires, `reuse_existing` re-selects and re-binds the *same* stale null
row on the next attempt, and the run crash-loops forever with no
self-heal
> - This pull request backfills `projectWorkspaceId` during restore
(prefer the existing binding, fall back to the resolved one) so the row
heals on first reuse and the guard stops firing
> - The benefit is that reused workspaces created before their project
had a primary project workspace self-repair on next use instead of
crash-looping, while genuine mismatches are still surfaced by the guard
## Linked Issues or Issue Description
No public GitHub issue exists — describing the bug inline per the bug
report template (`.github/ISSUE_TEMPLATE/bug_report.yml`):
### What happened?
In `heartbeatService`, the execution-workspace reuse/restore branch
calls
`executionWorkspacesSvc.update(reusableExistingExecutionWorkspace.id, {
… })` without a `projectWorkspaceId` field. Only the sibling CREATE
branch sets `projectWorkspaceId`. So an execution workspace that was
persisted with a null `projectWorkspaceId` (e.g. created before its
project had a primary project workspace) is never backfilled on restore.
When such a workspace is later reused for a run whose issue resolves a
concrete project workspace, the pre-launch guard throws
`persisted_workspace_missing_project_workspace_id`, the run fails, and
`reuse_existing` re-binds the identical stale row on the next attempt —
an unbounded crash-loop with no self-heal.
### Expected behavior
On restore, the reused workspace's `projectWorkspaceId` is backfilled
from the resolved project workspace when it is currently null, so the
guard passes and the run launches. An existing non-null binding is never
overwritten (a genuine mismatch is still surfaced by the separate
`project_workspace_mismatch` guard).
### Steps to reproduce
1. Have an `execution_workspaces` row with `project_workspace_id = NULL`
that is eligible for reuse.
2. Give its project a primary project workspace (so the issue now
resolves a concrete `projectWorkspaceId`).
3. Dispatch a run for an issue in that project that reuses the
workspace. The restore `update()` leaves `project_workspace_id` null,
the launch guard throws
`persisted_workspace_missing_project_workspace_id`, and every subsequent
reuse re-binds the same null row and fails identically.
### Paperclip version or commit
`master` (branched from `14f20be92`); reproduced on a live self-hosted
instance.
### Deployment mode
Self-hosted, embedded Postgres, local adapters.
## What Changed
- New exported pure helper
`reconcileReusedExecutionWorkspaceProjectWorkspaceId(existing,
resolved)` in `server/src/services/heartbeat.ts`, returning `existing ??
resolved ?? null`. It prefers an existing binding (never nulls out a
good value or silently rebinds a genuine mismatch — the guard still
surfaces those), backfills a null binding from the resolved value, and
stays null when neither is present.
- Wire the helper into the reuse/restore
`executionWorkspacesSvc.update(...)` call so the restored row's
`projectWorkspaceId` is set to
`reconcileReusedExecutionWorkspaceProjectWorkspaceId(reusableExistingExecutionWorkspace.projectWorkspaceId,
resolvedProjectWorkspaceId)`. The CREATE branch already set
`projectWorkspaceId`; this brings the restore branch to parity.
## Verification
- Added 3-case unit coverage in
`server/src/__tests__/heartbeat-workspace-session.test.ts` for the
helper: (a) backfills a null existing binding from the resolved value,
(b) never overwrites an existing binding even when a resolved value is
present, (c) returns null when both existing and resolved are absent
(null and undefined inputs).
- Confirmed the `update()` patch type accepts the field:
`executionWorkspacesSvc.update` takes `Partial<typeof
executionWorkspaces.$inferInsert>`, and `projectWorkspaceId` is a column
on that table; both
`reusableExistingExecutionWorkspace.projectWorkspaceId` and
`resolvedProjectWorkspaceId` are `string | null`, matching the helper's
`string | null | undefined` params / `string | null` return.
- Live-instance exposure check (embedded Postgres): 354
`execution_workspaces` rows carry a null `project_workspace_id`; all of
them belong to projects with **no** project workspace, so
`expectedProjectWorkspaceId` currently resolves null and the guard does
not fire today. The fix is durable heal-on-reuse protection for the
moment any such project gains a primary project workspace (or a null row
is reused for an issue that resolves one).
- CI (full pnpm workspace install) runs the authoritative test +
typecheck for this change on this PR.
## Risks
- Low risk; scoped to the execution-workspace restore path, no schema or
API change.
- The helper only ever *adds* a `projectWorkspaceId` where the row had
none; it never overwrites an existing binding, so it cannot mask a real
`project_workspace_mismatch` (that guard still runs after).
- Complementary to (not overlapping with) #10130, which escalates a
terminal `workspace_validation_failed` run to `blocked` from the
recovery side; this PR prevents the guard from firing on reuse in the
first place. Neither depends on the other.
## Model Used
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use /
code execution (repo edit, embedded-Postgres exposure query, unit-logic
verification).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (searched open PRs touching heartbeat / execution-workspace /
reuse; only #10130 is related, and it is complementary)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes (no
user-facing docs affected)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (in progress)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.807.0-canary.10
|
||
|
|
9abb600e72 |
docs(connections): MCP-direct/DCR playbook section + Notion dry-run appendix (#11030)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools through catalog connectors. > - The connector playbook is the repeatable template for adding such connectors. > - Notion just shipped as the first MCP-direct connector with RFC 7591 dynamic client registration (#11009). > - The playbook had no guidance for MCP-direct connections, OAuth discovery, or DCR. > - This pull request documents that path and encodes three mandatory standards into the connector template. > - It also adds a Notion dry-run appendix recorded from the live probe and the shipped implementation. > - The benefit is that the next MCP-direct connector follows a recorded, verified path. > [!IMPORTANT] > **Depends on #11009.** This PR documents the Notion MCP connector that ships in #11009 (RFC 7591 DCR, discovery-first endpoints, `redirectConstraints` enforcement, refresh-rotation hardening). Reviewing this doc against master before #11009 merges will show the documented behavior as "not implemented" — that is expected. Draft until #11009 lands, then re-review. ## What Changed `doc/connections/CONNECTOR-PLAYBOOK.md` only (+364 lines, no code): - New **"MCP-Direct Connections"** section: RFC 9728/8414 endpoint discovery chain, an **RFC 7591 dynamic client registration** subsection (public client, PKCE S256, env-client precedence, persist-and-reuse), and a **redirect-URI constraints** subsection (`https-or-loopback-http`, fail-fast wizard error). - Three mandatory documentation standards encoded into the connector template itself: a service-involvement statement (DCR providers need neither Paperclip ID nor Paperclip Connect — instance-local per the PAP-14828 spec §10 item 8.4; cloud and self-hosted use the same path), a required **Connection Flow** section (sequence diagram + exact authorize/token/registration/callback endpoints), and a required **Administrator Setup** section. - A **Notion dry-run appendix** mirroring the Linear appendix, recorded from the live PAP-16649 probe: verified request sequence, redirect-URI probe results table, sequence diagram (mermaid), shipped manifest sketch, representative tool risk classes, admin setup (nothing to register), governance defaults, and validation hooks. ## Verification - Docs-only change; no code paths affected. `git diff --stat` shows exactly one file. - Every endpoint, error code, and constraint in the appendix was checked against the shipped implementation on the #11009 head: `server/src/services/tool-access.ts` (`assertOAuthRedirectConstraints`, DCR registration metadata, refresh serialization), `server/src/routes/tool-access.ts` (`POST /api/tools/oauth/:connectionId/start`, `GET /api/tools/oauth/callback`), and `packages/shared/src/app-definitions/notion.json` (`redirectConstraints: "https-or-loopback-http"`). - The request log mirrors the live probe record from PAP-16649 (2026-08-06/07), not vendor docs alone. - Mermaid source renders cleanly (rendered PNG attached to PAP-16653). ## Risks - Low: documentation only. Main risk is doc/implementation drift if #11009 changes before merging — mitigated by the dependency note above and re-review after #11009 lands. - The template changes add mandatory sections for future connector proposals; existing proposals are not retroactively invalidated. ## Model Used Claude Fable 5 (claude-fable-5). ## Linked Issues or Issue Description PAP-16653 (parent PAP-16637 P5). Companion catalog package landed in paperclip-content (`integrations/catalog/platforms/notion/areas/mcp/`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.807.0-canary.9 |
||
|
|
01b51dc0e5 |
fix(runtime): stop Live badge and Working shimmer after task teardown (#10985)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task list and the chat views show a Live badge and a Working shimmer for an issue that has an active run. > - A finished task kept the Live badge and the Working shimmer after the run ended and the sandbox stopped. > - The user interface reads run liveness from the `heartbeat_runs.status` row. The run finalizer writes the terminal status in a step that is separate from the agent `status=done` update. When the sandbox or the run process stops between the two steps, `heartbeat_runs.status` stays `running` forever. > - A run row that stays `running` makes a finished task look perpetually Live, and the user interface has no guard for an issue that already reached a terminal status. > - This pull request closes the invariant "environment lease released implies the run is terminal" on the server, and adds a user interface guard that suppresses live state for a terminal issue. > - The benefit is that a finished task stops showing Live and Working, both at the source (the run row) and at the surface (the badge and the shimmer). ## Linked Issues or Issue Description **Bug description** - A completed task kept the Live badge and the Working shimmer after its run ended and the sandbox was torn down. **Steps to reproduce** - Run an agent task to completion. Let the sandbox tear down while the run finalizer is between the `status=done` update and the terminal run-status write. - Open the task list or the chat view for the finished task. **Expected behavior** - A finished task shows no Live badge and no Working shimmer. **Actual behavior (before this change)** - The finished task showed the Live badge and the Working shimmer because its `heartbeat_runs.status` row stayed `running`. This pull request supersedes the two separate pull requests #10954 (frontend) and #10955 (backend). It carries all of their changes for the same race. ## What Changed Server: - Run teardown terminalizes a still-running or still-queued run before it releases the environment lease. It writes `succeeded` when the issue already reached `done`, `cancelled` when the issue is `cancelled`, and `interrupted` otherwise. It never overwrites a status that another path already made terminal. - The recovery stale-lock sweep terminalizes an orphaned running run to `interrupted` after it confirms the process and the sandbox are both gone. It requires recorded process metadata, so it never terminalizes a live run, a queued run, or a scheduled retry. - Each terminal transition writes a run event. - The stale-lock sweep continues and clears the lock when the audit write fails. It logs the failure loudly. - New server tests cover both invariants. User interface: - A shared guard suppresses the Live badge and the Working shimmer when the issue status is terminal. - The guard keeps non-terminal `queued` and `running` issues live. - The guard prefers the newest issue live-status snapshot. - New user interface tests cover the guard and the snapshot preference. ## Verification Server: - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps && tsc --noEmit` in `server/` — 0 errors. - `pnpm --filter server test heartbeat-run-lease-release-terminalization.test.ts recovery-stale-issue-lock-sweep.test.ts` — 12 tests pass. User interface: - `pnpm --filter @paperclipai/ui typecheck` — 0 errors. - `pnpm exec vitest run ui/src/lib/liveIssueIds.test.ts ui/src/lib/issue-chat-messages.test.ts` — 40 tests pass. ## Risks - Low risk. The server change only forces a still-live run row to a terminal status when the lease releases or when the recovery sweep confirms the process is dead. It never overwrites an existing terminal status, and it guards the recovery path with process metadata to avoid terminalizing a live run. - The user interface change is additive. The guard only suppresses live state for a terminal issue and keeps queued and running issues live. - No database migration. No change to any external endpoint. ## Model Used - Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.8 |
||
|
|
4683f26c97 |
chore(lockfile): refresh pnpm-lock.yaml (#11036)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
11e56654f8 |
feat(ui): port onboarding flow from prototype; add cloud + local variants (#10786)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - First-run onboarding is the subsystem that turns a brand-new install
into a working company: it creates the company, its goal, a lead agent,
and that agent's first task
> - The existing `OnboardingWizard` carried all of that wiring
correctly, but its UI had drifted from the current design direction, and
a separate design prototype (`paperclip-onboard`) existed as a
standalone visual mock with no backend
> - Porting the prototype's *logic* would have thrown away working,
well-tested backend orchestration; leaving the two apart meant the
design never shipped
> - Separately, cloud and local (self-hosted) installs need meaningfully
different first runs — local has no sign-in and must let the user pick a
locally-installed CLI adapter — so a single linear wizard could not
serve both
> - This pull request rebuilds the presentational layer from the
prototype on top of the existing backend orchestration, and splits it
into two thin flow containers over a shared core
> - The benefit is that the shipped onboarding matches the intended
design, cloud and local can diverge without duplicating logic, and each
can later ship to a different app version while sharing one set of step
components
## Linked Issues or Issue Description
No existing issue — describing inline (feature request).
**What problem does this solve?**
Onboarding is the first thing a new user sees, and the shipped wizard
had drifted from the current design. In parallel, cloud and local
installs need different first-run paths: local has no hosted sign-in,
and its agent runs on a CLI adapter installed on the user's machine,
which the cloud path never has to ask about. There was no way to express
that difference without either forking the whole wizard or bolting
conditionals onto a single linear flow.
**Proposed solution**
Extract the onboarding step views and shell into a shared core, then
compose two thin flow containers (cloud and local) over it. Keep all
backend orchestration in the existing `useOnboardingFlow` hook so no
working logic is rewritten.
**Alternatives considered**
- *Single flow with a `variant` prop* — most DRY, but the two flows are
intended to ship on different app versions, and a shared file would have
to be split later anyway.
- *Two fully independent copies* — simplest per-flow, but every shared
refinement (spacing, motion, copy) would have to be made twice and would
drift.
## What Changed
- **Shared core** under `ui/src/components/onboarding/`:
`OnboardingScaffold` owns the full-screen shell and the single
`AnimatePresence` step crossfade, so both flows transition identically;
step views (Start / Company / Agent / Task), `FooterNav`, `AgentPreview`
and the motion constants are extracted for reuse.
- **`CloudOnboardingFlow`** — `start → company → agent → task`; mounted
in the real app via `OnboardingWizardVariant`. Behaviour matches the
retired wizard, including `previewMock` and the existing-company ("add
an agent") entry point.
- **`LocalOnboardingFlow`** — skips sign-in and adds an optional email
ask (with a privacy assurance), a local model/adapter step that hires
with `requireEnvProbe: true`, and a "star us on GitHub" interstitial
before completing. **Harness-only for now** — the real app still mounts
the cloud flow.
- **Deleted `OnboardingWizard.tsx`** (1,786 lines); updated its
Storybook stories and the `OnboardingWizardVariant` test to the new
components.
- **Orbiting 3D paperclip backdrop** behind the auth and welcome screens
(`three`), code-split so it only downloads on those screens; honours
`prefers-reduced-motion` and disposes its GL context on unmount.
- **`motion`** added for step transitions and the agent-capsule
choreography.
- Visual values routed through design tokens per `DESIGN.md`; `Stepper`
generalized to take a step total (backward compatible); `/design-guide`
page and the component index updated.
- **Standalone preview harness** (`ui/onboarding-preview.html`) with
`?flow=` and `?step=` for backend-free review, wired as a second Vite
rollup input.
- **Adapter env probe bound to the adapter it ran against.**
`hireLeadAgent` reused `adapterEnvResult` for any adapter, so when a
hire failed and the user picked a *different* local adapter and retried,
the previous adapter's verdict satisfied the `requireEnvProbe` guard
while the hire posted the new adapter's config — hiring it unprobed. The
cache is now keyed on the adapter type plus the exact config posted to
the test endpoint, the config is built once and shared by probe and
hire, a failed probe clears the cache, and `clearAdapterEnvResult()`
(called on adapter change) stops the step displaying a stale verdict.
Cloud is unaffected — it hires with `requireEnvProbe: false`. Reported
by Greptile.
- **E2E specs re-pointed at the new flow.** Four specs still drove the
deleted wizard (`onboarding`, `conference-room-typing-intro`,
`planning-mode-visual-verification`, `nux-phase4-screenshots`) and
failed with `element(s) not found` on `"Name your company"` /
`input[placeholder="Acme Corp"]`. Rather than repeat the new drive
sequence four times, `tests/e2e/onboarding-flow.ts` adds one driver per
step (`startCloudOnboarding`, `completeCompanyStep`,
`completeAgentStep`, `completeTaskStep`, `completeCloudOnboarding`) and
the specs import it, so the next flow change touches a single file. Two
now-dead `**/test-environment` route stubs went with it — the cloud flow
hires with `requireEnvProbe: false`, so that probe never fires.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `npx vitest run` over the onboarding suites
(`OnboardingWizardVariant`, `AgentCapsule`, `onboarding-launch`,
`onboarding-goal`, `onboarding-route`, `onboarding-adapter-config`) — 33
tests pass.
- `pnpm --filter @paperclipai/ui build` — succeeds; the three.js chunk
splits out separately (522 kB raw / 133 kB gzip) rather than entering
the main bundle.
- Both flows driven end-to-end in the preview harness in `previewMock`
(no database writes), plus the cloud flow rendered in the real
authenticated app at `/onboarding` to confirm the mount swap.
- The four re-pointed e2e specs pass locally against the new flow.
- New `ui/src/hooks/useOnboardingFlow.test.tsx` — 4 cases pinning the
adapter-probe cache (switch-adapter retry, cold path, explicit clear,
and the cloud flow's `requireEnvProbe: false`). Verified non-vacuous:
the switch-adapter case fails against the pre-fix code.
- Rebased onto current `master`; `pnpm-lock.yaml` is deliberately
**not** committed — `.github/workflows/pr.yml` regenerates it when a
manifest changes and shares it with downstream jobs as the `pr-lockfile`
artifact.
## Risks
- **Deleting `OnboardingWizard.tsx` is the one change that alters
existing app behaviour.** The cloud flow is intended to be
behaviour-equivalent, and its entry points are covered by the updated
`OnboardingWizardVariant` test, but this is the area to review most
closely.
- **Conflict risk with open PRs that touch the old wizard**: #9900,
#9501, #8982 and #6636 all modify
`ui/src/components/OnboardingWizard.tsx`, which this PR removes.
Whichever lands second will need its change re-applied to the new step
components. Flagging so ordering can be decided deliberately.
- **New dependencies**: `motion` and `three` (+ `@types/three`). `three`
is large, so it is lazily imported and code-split — it does not affect
the main bundle. Both are MIT.
- The **local flow is not reachable in the app** yet (harness/canary
only), so it carries no runtime risk today; wiring it up is a follow-up.
- The auth screens remain **presentational only** — they are not wired
to real auth, unchanged from before this PR.
- **Pre-existing, not introduced here:** `OnboardingWizardVariant`
renders outside `<Routes>` in `App.tsx`, so its `useParams()` never
resolves `:companyPrefix` and `/{prefix}/onboarding` opens the welcome
screen instead of jumping to the agent step. `master` has the identical
structure, so this PR faithfully ports existing behaviour; the working
"add an agent" entry is the launcher card behind the overlay, which is
what the screenshot spec drives. Worth a separate fix.
## Model Used
Claude Opus 5 (`claude-opus-5`) via Claude Code, with extended thinking
and tool use (repo search/edit, local test + build execution, and
browser-driven visual verification of the rendered flows). Portions of
the session also ran on `claude-opus-4-8` and `claude-fable-5`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
canary/v2026.807.0-canary.7
|
||
|
|
b67c512f82 |
feat(ui): live run label → run detail, running row → task detail (#11034)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The agent detail page has a Dashboard tab. The Dashboard tab shows a "Live Run" section for the agent's current heartbeat. > - The "Live Run" section has two clickable pieces: the section heading and the running row. Both pieces linked to the same run detail page. > - Two controls that go to the same place waste a navigation affordance and hide the task the agent runs. > - This pull request splits the two destinations. The heading goes to the run. The running row goes to the task. > - The benefit is that a user reaches the run internals from the label and the work item from the row, in one click each. ## Linked Issues or Issue Description No public GitHub issue exists for this change. The description follows the enhancement issue template. **What existing behavior does this improve?** The "Live Run" section on the agent detail page, Dashboard tab (the `LatestRunCard` component in `ui/src/pages/AgentDetail.tsx`). **Current behavior** The "Live Run" heading and the running row both link to the run detail page (`/agents/:agentId/runs/:runId`). The row shows the run code and an invocation-source chip. There is a separate "View details →" link that also goes to the run detail page. A user cannot reach the task the run works on from this section. **Proposed behavior** The heading becomes a link to the run detail page and appends the short run code, shown as `Live Run · <run code>`. The redundant "View details →" link is removed. The running row links to the task detail page when the run's context snapshot resolves to a known issue, and the row then shows the task status glyph, the task slug, and the task title. A pure timer heartbeat with no resolvable task keeps the previous behavior: run code plus source chip, linking to the run detail page. **Reason and benefit** The heading and the row now go to distinct, intuitive destinations. A user reaches the run internals from the label and the work item from the row, each in a single click. The fallback keeps heartbeats with no task readable and avoids a blank or broken row. ## What Changed - Made the "Live Run" / "Latest Run" heading a `Link` to the run detail page and appended the short run code (`run.id.slice(0, 8)`) in a mono span, formatted `Live Run · <run code>`. Kept the pulsing live dot. - Removed the redundant "View details →" link. - Changed the running row `Link` target to the task detail page (`/issues/:identifier`) when a task resolves, falling back to the run detail page otherwise. - Resolved the task from the run context snapshot (`contextSnapshot.issueId`, falling back to `contextSnapshot.taskId`) against a `Map` of the agent's assigned issues threaded in from `AgentOverview`. - When a task resolves, replaced the run code and source chip in the row with the task status glyph (`StatusGlyph`), the task slug, and the task title. Kept the running spinner, the run status badge, and the timestamp. ## Verification - `pnpm check:token-gates` → 3/3 gates clean. - `pnpm --filter ui typecheck` → passes. - `pnpm --filter ui exec vitest run src/pages/AgentDetail.progress.test.ts src/pages/AgentDetail.instructions.test.tsx` → 10/10 pass. - Manual (needs a reviewer with a browser): open an agent detail page → Dashboard tab. - For a live issue-execution run: the heading reads `Live Run · <run code>` and opens the run detail page; the row shows the task status icon, slug, and title and opens the task detail page. - For a pure timer heartbeat with no task: the row falls back to run code + source chip and opens the run detail page. No blank row. - Confirm both states in light and dark mode. ## Risks Low risk. The change is presentational and scoped to one component. The task lookup is defensive: it reads the context snapshot with a fallback key and only renders the task row when the issue is present in the already-loaded assigned-issue set, so an unknown or missing issue degrades to the previous run-detail behavior rather than breaking. ## Model Used - Provider: Anthropic (Claude). - Model: claude-opus-4-8 (Opus 4.8). - Context window: 200K. - Reasoning mode: extended thinking, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.807.0-canary.6 |
||
|
|
4e76227f12 |
feat(plugin-daytona): stream session command logs behind useLogStream (#11021)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox provider plugins let agents run commands in remote environments. > - The Daytona provider polls the exit code while it waits for command logs. > - Polling delays log delivery and does not support long-lived streamed commands. > - This pull request adds an opt-in Daytona log stream with one reconnect and a poll fallback. > - The benefit is faster log delivery while the existing default path stays unchanged. ## Linked Issues or Issue Description Refs: #10941 **Subsystem affected** packages/plugins — sandbox provider plugins. **Problem or motivation** The Daytona provider polls the command exit code every 50 milliseconds while it waits for logs. This delays output and does not support a long-lived streamed command. **Proposed solution** Add the `useLogStream` provider option. Stream stdout and stderr from the Daytona callback log form, read the exit code after the stream ends, retry the read with bounded backoff, and fall back to the existing poll path after a disconnect. **Alternatives considered** Keep polling for all commands. This keeps the current behavior but does not provide timely logs or a path for long-lived commands. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents, including Daytona. ## What Changed - Add the opt-in `useLogStream` option with a default of `false`. - Stream stdout and stderr from the Daytona callback log form. - Drop replayed log prefixes by delivered byte offset after reconnect. - Retry once after disconnect, then use the existing poll path. - Read the exit code once after a successful stream and retry when the code is not ready. - Add tests for ordered output, exit-code reads, disconnect fallback, and reconnect replay handling. ## Verification - Run `node_modules/.bin/vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`. - Confirm that the full Daytona plugin test project passes with 127 tests. - Confirm that the changed files type-check against Daytona SDK 0.203.0 types. - Review the PR against parent PR #10941 before it reaches `master`. ## Risks - The stream path changes behavior only when `useLogStream` is `true`. - A stream failure can add one reconnect attempt before the existing poll fallback. - The stream path has no command deadline because it supports long-lived commands. - No endpoint, stored data, telemetry shape, authentication rule, or result shape changes. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The context window and reasoning mode are not exposed by this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.5 |
||
|
|
9485ffea70 |
fix(config): preserve env files during managed updates (#10980)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The CLI and server both update Paperclip values in `.env` files > - The server preserved operator content, but the CLI rebuilt the complete file > - A CLI rerun could remove comments, custom values, ordering, and newline style > - Both paths need one editor with one value encoding and duplicate key policy > - The final integration also needs one regression test across the related setup and sync safety mechanisms > - This pull request moves the editor to the shared package and adds cross-cutting rerun-survival coverage > - The benefit is safe setup and worktree repair reruns that preserve operator edits ## Linked Issues or Issue Description **What happened?** The CLI rebuilt the complete `.env` file when it wrote a managed Paperclip value. This action removed comments, blank lines, custom keys, original quoting, and the original newline style. **Expected behavior** Paperclip must update only the managed assignments. It must preserve all unrelated bytes. It must skip the file replacement when all managed values are current. **Steps to reproduce** 1. Add comments, custom keys, quoted values, and CRLF newlines to the Paperclip `.env` file. 2. Run a CLI path that calls the agent JWT secret setup. 3. Observe that the old writer replaces the complete file. **Paperclip version or commit** The problem exists on `master` before this pull request. Related public context: Refs #437. ## What Changed - Add one shared line-preserving `.env` editor for the CLI and server. - Define minimal and JSON value encodings in the shared helper. - Update every stale duplicate of a managed key and preserve current duplicate encodings. - Preserve comments, ordering, blank lines, unknown keys, export prefixes, trailing comments, and newline style. - Write changed files through a same-directory temporary file and atomic rename. - Limit CLI updates to non-empty `PAPERCLIP_*` entries. - Skip the write when all managed values are current. - Add shared, CLI, and server regression coverage. - Refresh the branch after the related config, sandbox, and skill safety changes landed. - Add a cross-cutting integration test for config, env-file, managed-sandbox, and managed-instructions rerun survival. ## Verification - `pnpm exec vitest run packages/shared/src/env-file.test.ts packages/shared/src/config-schema.test.ts cli/src/__tests__/agent-jwt-env.test.ts cli/src/__tests__/config-store.test.ts server/src/__tests__/config-file.test.ts server/src/__tests__/worktree-config.test.ts` passes 39 tests. - `pnpm exec vitest run server/src/__tests__/rerun-survival.integration.test.ts` passes 4 tests. - `pnpm -r typecheck` passes on the previous head. GitHub CI reruns it on the refreshed head. - The previous head passed the complete general, serialized, workspace, and E2E matrix. GitHub CI reruns that matrix on the refreshed head. - `pnpm build` passes on the previous head. GitHub CI reruns it on the refreshed head. ## Risks - Low risk. The production change only affects managed `.env` assignments. - Existing managed assignments can keep their original quoting when their decoded values are current. - Changed CLI values keep the prior minimal encoding policy. Changed server values keep the prior JSON encoding policy. - Duplicate managed assignments now follow one explicit rule: Paperclip updates each stale occurrence. - The master refresh had one import-block conflict. The resolution keeps both the config merge imports and the env-file imports. - The added integration file is test-only. It has no database, API, or UI contract effect. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex from the GPT-5 family produced this change with reasoning, tool use, and code execution. The runtime did not expose the exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5da382fd59 |
feat(skills): require explicit merge modes (#10978)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can select company skills and synchronize them to adapter runtimes > - The skill sync API replaced the complete selection without an explicit destructive choice > - Company package import also replaced conflicting skills by default > - These defaults could remove operator edits during setup and import reruns > - This pull request adds explicit assignment merge modes and safe package conflict handling > - The benefit is that reruns preserve operator work unless the caller explicitly requests replacement ## Linked Issues or Issue Description **What existing behavior does this improve?** This change improves agent skill synchronization and company package import. **Subsystem affected** This is a cross-cutting change across the shared contracts, server, CLI, and UI. **Current behavior** Agent skill synchronization replaces the full desired skill set from a modeless request. Package import replaces a conflicting skill when the caller does not select a conflict mode. **Proposed behavior** Agent skill synchronization requires `add`, `remove`, or `replace`. Package import skips conflicts by default. Each imported skill reports whether it was created, renamed, replaced, or skipped. **Reason and benefit** Setup and import reruns must preserve operator edits by default. Explicit destructive modes make data loss less likely and make each outcome inspectable. **Breaking changes** Callers of the agent skill sync API must now send `mode`. Callers that need the former behavior must send `replace`. Package import now uses `skip` when `onConflict` is absent. ## What Changed - Added required `add`, `remove`, and `replace` modes to the shared agent skill sync contract. - Added actionable `422` validation for missing or invalid modes. - Updated first-party UI and CLI callers with explicit modes. - Changed package skill conflict handling to use `skip` by default. - Kept plugin-owned and built-in stock skill imports on explicit `replace`. - Added created, renamed, replaced, and skipped results to company imports. - Added regression coverage for merge modes and package conflict outcomes. ## Verification - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm build` - `pnpm test:run:serialized` (128 suites passed) - `pnpm --filter @paperclipai/skills-catalog test` (20 tests passed) - Focused agent skill route, company skill service, portability, CLI, and UI tests passed. - GitHub CI passed build, typecheck, canary, all general and serialized test shards, all browser shards, policy, security, and final verification on commit `2cfbb3e4c5`. - Greptile reviewed the latest commit at 5/5 with zero unresolved threads. ## Risks - This change intentionally rejects modeless agent skill sync requests. - The safe package default can leave an existing skill unchanged where the old default overwrote it. - All first-party callers now select a mode. Regression tests cover each outcome. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI `gpt-5.6-sol` through Codex. The runtime used agentic reasoning, tool use, code execution, and repository editing. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.4 |
||
|
|
f6c6452b25 |
fix(server): preserve managed environment drift on boot (#10979)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip creates a managed sandbox environment for each company during boot > - Operators can change environment fields after Paperclip creates the environment > - The boot reconciler replaced those changes without checking for drift > - This pull request adds stock hashes and transactional drift reconciliation > - The benefit is that Paperclip can update untouched stock fields without losing operator work ## Linked Issues or Issue Description **What happened?** The managed sandbox boot reconciler rewrote the stock description, configuration, metadata, and status on every start. It did not detect operator changes first. A restart could therefore remove an operator's changes. **Expected behavior** Paperclip must preserve operator changes by default. It must update an untouched stock environment when Paperclip ships new stock values. It must perform each row update and stock-hash update atomically. **Steps to reproduce** 1. Start Paperclip and let it create the managed sandbox environment. 2. Change one Paperclip-owned stock field on that environment. 3. Restart Paperclip. 4. Observe that the previous reconciler replaced the change. **Paperclip version or commit** This bug reproduces on `master` before this change. **Deployment mode** Local development and self-hosted server boot are affected. ## What Changed - Add a shared deterministic stock-hash and drift classifier for built-in resources. - Track the managed sandbox stock hash with the company-scoped built-in resource binding. - Reconcile the environment and its stock metadata in one transaction with a row lock. - Preserve operator-modified and unmanaged rows and report their skipped update state. - Use archive ownership tokens so provider recovery reactivates only Paperclip-archived rows and preserves later operator archive decisions. - Keep operator-owned environment variables and unrelated metadata out of the stock fingerprint. - Add activity records for managed environment creation, updates, skipped drift, tracking initialization, and archive changes. - Add regression tests for current stock, available stock updates, operator drift, unmanaged rows, archive and reactivation, user-owned fields, and concurrent reconciliation. ## Verification - `pnpm -r typecheck` - `pnpm build` - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm test:run -- --mode serialized` (128 suites passed) - Repository general server, UI, CLI, shared, skills catalog, database, adapter, plugin SDK, and plugin creator projects passed. Two embedded-Postgres tests exceeded the host's five-second default under the aggregate run and passed in the complete database project with `--testTimeout=20000`. One timing-sensitive sandbox stream test passed on its focused retry. - Focused managed-environment unit and integration coverage passed: 49 tests across the drift classifier, boot report, and environment service suites. ## Risks - The main risk is an incorrect ownership boundary in the stock fingerprint. The fingerprint includes only Paperclip-owned stock fields. Tests confirm that environment variables and unrelated metadata survive reconciliation. - Concurrent reconciliation could otherwise overwrite a late operator edit. The implementation locks the environment row and updates the row and hash binding in one transaction. A concurrency test covers this path. - There is no schema migration. Existing managed rows initialize tracking without replacing their current values. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex. The exact serving snapshot and context-window size were not exposed. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
35132af161 |
fix(config): preserve extensions and guard invalid repairs (#11005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The CLI and server share a JSON configuration contract for local installations and worktrees. > - Existing config writes removed extension keys because Zod stripped unknown object properties. > - Invalid config files could also be replaced with defaults before an operator preserved the original bytes. > - Configuration updates must preserve operator edits and must not rewrite files when the effective value is unchanged. > - This pull request adds extension-preserving merges, guarded invalid-config repair, atomic writes, and focused regression tests. > - The benefit is safe setup and configuration reruns without data loss or unnecessary mtime changes. ## Linked Issues or Issue Description **What happened?** Known-field updates through the CLI or server removed unknown top-level and nested config keys. Non-interactive configure and onboard paths could replace a present but invalid config with defaults. **Expected behavior** Writers preserve extension keys, skip semantic no-op writes, and require explicit interactive confirmation before an invalid config is replaced. Repair preserves an exact collision-safe backup first. **Steps to reproduce** 1. Add an unknown top-level key and an unknown nested provider key to `config.json`. 2. Update a known field through the CLI or worktree config writer. 3. Observe that the extension keys are removed on the base branch. 4. Write invalid JSON and run configure or onboard without an interactive terminal. 5. Observe that the original file can be replaced without a durable invalid-file backup on the base branch. **Paperclip version or commit** `master` at the pull request base commit. ## What Changed - Accept unknown properties at each extensible config object boundary while keeping every known field validated. - Merge known-field updates into the parsed source config and preserve only unknown extension data. - Warn about near-match key names without deleting or changing them. - Skip writes when the effective config is unchanged, which keeps file mtimes stable. - Write config changes through a temporary file, file sync, rename, and directory sync. - Distinguish a missing config from an invalid config in configure and onboard. - Back up invalid bytes as `config.json.invalid-N` and verify the source still matches that backup before repair. - Require interactive repair confirmation and reject non-interactive replacement with an actionable message. - Document the config preservation and repair behavior. ## Verification - `pnpm exec vitest run packages/shared/src/config-schema.test.ts cli/src/__tests__/config-store.test.ts cli/src/__tests__/configure-repair.test.ts cli/src/__tests__/configure.test.ts cli/src/__tests__/onboard.test.ts server/src/__tests__/config-file.test.ts server/src/__tests__/worktree-config.test.ts` - `pnpm -r typecheck` - `AWS_ACCESS_KEY_ID= AWS_SECRET_ACCESS_KEY= VITEST_MAX_WORKERS=1 pnpm test:run` - `pnpm build` - Confirm all pull request checks are green on the latest commit. - Confirm Greptile reports 5/5 with no unresolved comments. ## Risks - Passthrough keeps misspelled keys. Near-match warnings make this visible without destructive cleanup. - Merge behavior must distinguish unknown extension keys from optional known keys. Schema-aware regression tests cover preservation and known-key deletion. - Repair must not overwrite bytes that changed after backup. The writer compares the current source with the selected backup before atomic replacement. - The change does not alter database schema, company scoping, or activity logging. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 model family. The exact deployment model ID and context window are not exposed. Agentic reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ace548fd2 |
feat(observability): rename sandbox provider spans and add run-time wrapper spans (#10999)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapter and sandbox code to start agents and run sandbox work > - The current sandbox spans use mixed names and do not group related run-time work > - Mixed names make traces harder to read and compare across providers > - This pull request renames provider spans, adds run-time wrapper spans, and keeps the host allowlist closed > - The benefit is clearer traces with the same sandbox behavior and trust boundary ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves OpenTelemetry span names and grouping for sandbox startup, execution, callback relay, and agent session work. **Subsystem affected** Cross-cutting (multiple of the above): adapter utilities, sandbox providers, shared telemetry documentation, and server instrumentation. **Current behavior** Sandbox provider spans use mixed names. Related run-time operations expose inner `sandbox.exec` spans without a named wrapper span. The host mapper uses a closed allowlist for provider span names. **Proposed behavior** Use descriptive provider-scoped span names. Add wrapper spans for agent session input, agent session output polling, and callback relay. Keep the host mapper allowlist closed and map unknown names to `other`. **Reason and benefit** Clear names make traces easier to read and reduce ambiguity during sandbox operation analysis. Wrapper spans show the full operation while preserving the inner execution spans. **Breaking changes** None. This change updates telemetry span names and grouping only. It does not change sandbox behavior, endpoint behavior, or the host trust boundary. **Additional context** Related prior work: [#10758](https://github.com/paperclipai/paperclip/pull/10758). ## What Changed - Rename Daytona provider sync and session spans with descriptive provider-scoped names. - Add three run-time wrapper spans for agent session input, output polling, and callback relay. - Add a shared span runner that preserves no-op behavior without a real tracer. - Keep the host mapper allowlist closed and map unknown names to `other`. - Update telemetry documentation and span-name tests. ## Verification - Focused adapter-utils span tests pass for startup timing, callback relay, and sandbox execution. - Focused Daytona plugin span tests pass for renamed leaf spans and session open or close spans. - Focused server tests pass for host mapping and instrumentation. - The stacked diff contains one commit on top of `feat/daytona-persistent-session-model`. ## Risks - Span names change for existing telemetry consumers. - The wrapper spans add trace structure but do not change sandbox execution. - The host mapper keeps the existing closed allowlist and `other` bucket. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 (Codex agent); exact deployment revision and context window are not exposed in this run; tool use and code execution enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.3 |
||
|
|
03cfad7ceb |
feat(apps): connect Notion through MCP OAuth (#11009)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - The Apps gallery lists Notion, but the server required manually configured OAuth credentials. > - Notion's hosted MCP server supports OAuth discovery and dynamic client registration. > - Notion also requires HTTPS or a loopback HTTP redirect URI. > - This pull request adds a direct Notion MCP OAuth path with PKCE and reusable dynamic clients. > - It also adds the current Apps UI states for connect and reauthorization. > - The benefit is a secure Notion connection with no manual client credential setup. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps gallery, Apps connect route, OAuth token lifecycle, and managed MCP gateway. **Subsystem affected** `server/`, `packages/shared/`, `scripts/`, and `ui/`. **Current behavior** The Notion gallery cards are disabled. The server uses the classic Notion OAuth endpoints and requires operator-supplied client credentials. It does not register an OAuth client from provider metadata. Concurrent refreshes can also replay a rotating refresh token. **Proposed behavior** Enable the Notion Apps flow. Discover OAuth metadata from `https://mcp.notion.com/mcp`. Register and reuse a public RFC 7591 client with PKCE. Require HTTPS or loopback HTTP callbacks. Serialize refreshes, store each rotated refresh token before the new access token can be used, and show a reconnect state for `invalid_grant`. **Reason and benefit** Operators can connect the built-in Notion MCP app without creating or copying OAuth credentials. Paperclip keeps dynamic clients and rotating tokens in the company secret store. **Breaking changes** None. Explicit environment client credentials still take priority. Existing Slack and Linear OAuth endpoint hints remain unchanged. Other OAuth apps remain disabled unless they are allowlisted. **Additional context** PR #10910 is a related, broader Connections v3 wizard replacement. This PR is the focused current Apps flow. The MCP Tool Gateway and Connected Apps items in `ROADMAP.md` cover this planned capability. ## What Changed - Classify all 20 reviewed Notion MCP tools with provider-scoped read and write defaults. - Require approval for selected Notion mutations, including move, duplicate, and convert actions that generic verb matching missed. - Preserve company-scoped connection and catalog resolution for Notion profiles and policies. - Add RFC 7591 dynamic client registration with `token_endpoint_auth_method=none` and mandatory PKCE. - Store the dynamic client ID on the connection and store any returned client secret in the company secret store. - Reuse the registered client for later connects and keep explicit environment credentials as the first choice. - Discover protected-resource and authorization-server metadata from the Notion MCP endpoint. - Add `redirectConstraints: "https-or-loopback-http"` to the generated Notion app definition and shared contract. - Reject non-loopback plain HTTP callbacks before network access with a TLS setup error. - Serialize client registration and token refresh operations within the server process. - Store a rotated refresh token before publishing the refreshed access token. - Treat `invalid_grant` as terminal and move the connection to a clear reauthorization state. - Add focused coverage for registration reuse, callback constraints, refresh rotation, and terminal grants. - Enable the Notion Apps route and add connect, redirect, success, error, and reconnect UI states. - Keep non-allowlisted OAuth apps blocked and cover the UI policy with regression tests. ## Verification - The focused Notion policy integration test passed with embedded PostgreSQL. - The focused 20-tool classification test passed. - The server typecheck passed on the governance head. - `pnpm -r typecheck` passed on the rebased head. - `pnpm --filter @paperclipai/server typecheck` passed after the security follow-up. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t \u0027DCR|refresh tokens|invalid_grant|abandoned lease\u0027` passed 10 focused security tests. - `pnpm build` passed on the rebased head. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts -t 'OAuth|oauth'` passed 14 tests. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts` passed 5 tests. - The complete server group passed 3,686 tests with 4 skipped. - The complete UI group passed 3,656 tests. - The full local runner found one environment-only CLI failure because this agent runtime injects static AWS credentials into a test that expects `AWS_PROFILE` only. `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` passed all 8 tests. - The prior UI verification passed 55 focused tests, `pnpm check:token-gates`, the Storybook build, and review of six 1440 x 1000 screenshots. - OAuth request sequence: protected-resource metadata `GET https://mcp.notion.com/.well-known/oauth-protected-resource/mcp`; authorization metadata `GET https://mcp.notion.com/.well-known/oauth-authorization-server`; dynamic registration `POST https://mcp.notion.com/register`; authorization `GET https://mcp.notion.com/authorize`; token exchange and refresh `POST https://mcp.notion.com/token`; MCP traffic `POST https://mcp.notion.com/mcp`. - The live metadata and registration probe confirmed that Notion accepts HTTPS and loopback HTTP redirects. It rejects a plain HTTP private hostname. - A later QA task owns the full browser consent and managed gateway tool-list dry run against a configured HTTPS deployment. ## Risks - Notion can add tools. Unrecognized names use the generic classifier, and new or changed risky tools stay quarantined after connection activation. - A deployment that uses a private non-loopback hostname must configure HTTPS before it can connect Notion. - Dynamic registration creates a provider-side client. Paperclip reuses it because registration does not provide a standard delete operation. - Refresh coordination uses a database CAS lease across service instances. An unclean crash leaves an uncertain lease and requires reconnect instead of risking refresh-token replay. - The current Apps surface overlaps with PR #10910. Merge order can require a small conflict resolution if that PR lands first. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex on a GPT-5 runtime. The exact deployment ID and context window are not exposed. The runtime used reasoning, repository tools, code execution, and network tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>canary/v2026.807.0-canary.2 |
||
|
|
ea83c5c822 |
feat(task-chat): bring back copy/👍/👎 actions on the agent bubble footer (#11025)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task view shows a threaded conversation between the user and the agent, with agent replies rendered as bubbles that carry a "✓ Worked · N tools" summary line. > - The conference-room chat already offers per-message copy and thumbs-up / thumbs-down feedback, but the redesigned task thread dropped these controls from the agent bubble footer. > - Users lose a quick way to copy an agent reply or send feedback on it, and the redesign silently ignored the feedback-vote props it was already given. > - This pull request prepends a copy · thumbs-up · thumbs-down cluster to the bubble summary line and wires the existing feedback-vote props through. > - The benefit is a consistent feedback surface across both chat views, with no new API. ## Linked Issues or Issue Description No public GitHub issue exists for this change; the underlying issue is described inline below following the feature template (`.github/ISSUE_TEMPLATE/feature_request.yml`). #### Problem or motivation The redesigned task thread renders each agent reply with a "✓ Worked · N tools · <timestamp>" summary line, but it dropped the copy and thumbs-up / thumbs-down controls that the conference-room chat still shows. Users can no longer copy an agent reply or vote feedback from the task thread. The redesign component already received `feedbackVotes` and `onVote` props but ignored them. #### Proposed solution Prepend a copy · 👍 · 👎 cluster to the summary line, leading the always-visible timestamp, reusing the shared `IssueChatFeedbackButtons` so both chat views speak the same feedback language. Anchor the cluster to the turn's summary row (a sibling of the expandable tool-history fold) so it stays on the summary line whether the tool history is collapsed or expanded. #### Alternatives considered Placing the cluster inside the expandable fold — rejected because expanding the tool history then re-centered the actions to the middle of the tall fold. #### Roadmap alignment UI polish to the task thread; no core-roadmap overlap. ## What Changed - Add `TaskChatBubbleActions`: a copy · thumbs-up · thumbs-down cluster built on the shared `IssueChatFeedbackButtons`. - Render the cluster on the agent bubble's "✓ Worked · …" summary line, leading the timestamp; runless agent replies get the same cluster with the timestamp trailing. Human and system bubbles are unchanged. - Add a `leading` slot to `TaskChatTurn` so the actions sit on the summary row, a sibling of the tool-history fold, and stay anchored when the fold expands. - Wire the redesign to the `feedbackVotes` / `onVote` props it already received. - Add a demo binding in the `TaskChatLab` dev harness. ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN. - `pnpm typecheck` — clean across all packages. - `cd ui && pnpm vitest run src/components/task-chat/TaskChatBubble.test.tsx src/components/task-chat/TaskChatTurn.test.tsx` — 31/31 pass. - Manual: open a task thread, confirm the copy / 👍 / 👎 cluster shows on the agent bubble summary line before the timestamp, copy works, votes toggle, and the cluster stays on the summary line when the tool history is expanded. Visual change: snapshot baselines are intentionally not updated, per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. UI-only change scoped to the redesigned task-chat bubble footer. It reuses an existing shared feedback component and existing vote props; no API, schema, or server change. Human and system bubbles are untouched. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.807.0-canary.1 |
||
|
|
cfed36ea6b |
feat(plugin-daytona): persistent session model with plain command dispatch (#10941)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies > - One core subsystem runs agent work inside sandboxes > - The Daytona provider uses that path to run user commands > - The current one-shot model does not keep a shell alive across commands > - This pull request adds an opt-in persistent session model for Daytona > - The benefit is faster command dispatch with the same sandbox boundaries ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting. This change touches `packages/adapters`, provider tests, span names, and sandbox command behavior. ### Problem or motivation The Daytona provider needs a persistent shell for repeated command dispatch. The old advisory wrapper path does not reach that goal. It also adds cost and removes the session speed gain. ### Proposed solution Add a `useSessions` driver flag. Keep it off by default. Open one Daytona session per lease when the flag is on. Send each user command into that session. Read stdout and stderr from the session logs endpoint. Run each command in a subshell so `exit` does not stop the shell. Remove the advisory `bwrap` wrapper path and its lease metadata. Add session setup and teardown spans. Keep a hard delete on teardown. ### Alternatives considered Keep the advisory `bwrap` wrapper. That path does not give a real persistent session. It also keeps extra command overhead. Keep a one-shot fallback for user commands. That would weaken the session model and hide a missing session case. ### Roadmap alignment This work fits the `Cloud / Sandbox agents` milestone in `ROADMAP.md`. It also supports the control plane goal of safe remote sandbox execution. ### Additional context The handoff verification reported `tsc --noEmit` clean and 119 Daytona unit tests passing. The handoff also reported a clean host span allowlist test and five expected commits on the branch. The security review gate remains required before merge. ## What Changed - Added an opt-in persistent session model for the Daytona sandbox provider. - Routed user commands through `executeSessionCommand` when sessions are enabled. - Removed the advisory `bwrap` command wrapper path and the lease metadata it used. - Added session lifecycle spans and span allowlist coverage. - Documented the leak bound in `DIRECTORY-CONSTRAINT-FINDINGS.md`. ## Verification - `tsc --noEmit` clean for the Daytona plugin, per handoff verification. - Daytona unit suite passes, with 119 tests, per handoff verification. - Host span allowlist test passes, per handoff verification. ## Risks - Persistent sessions can leak if teardown fails. - Session logs must keep stdout and stderr separate. - The flag stays off by default to limit rollout risk. ## Model Used OpenAI GPT-5, Codex, tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.807.0-canary.0 |
||
|
|
d84c5eae7a |
feat(ui): hide task priority from the UI (keep data model) (#11024)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks and issues carry a `priority` field that renders across many product surfaces: the detail header, the Triage properties panel, Kanban and thread cards, the New Task composer, list Sort/Group/Filter menus, search filters, and the dashboard chart. > - Product feedback found the priority level adds visual noise and decision cost without clear value in day-to-day task flow. > - We want to remove priority from the interface, but keep the data model, API, validation, and search DSL fully intact so the choice is reversible with no migration. > - This pull request hides every priority indicator and control behind one compile-time flag, `SHOW_TASK_PRIORITY_UI`, set to `false`. > - The benefit is a calmer, simpler UI now, with a single-boolean revive path and zero data loss. ## Linked Issues or Issue Description <!-- No public GitHub issue exists. Describing in-PR per the feature template. --> **Subsystem affected** The web UI (`ui/src`): issue detail, properties panel, Kanban/thread cards, New Task dialog, issues list Sort/Group/Filter menus, search filter bar/sheet, dashboard charts, and the design-guide showcase. **Problem or motivation** The task/issue priority level appears across many surfaces and adds visual clutter and decision overhead without pulling its weight in normal task flow. We want it gone from the interface without discarding the underlying data or breaking anything that depends on it. **Proposed solution** Add a single compile-time UI flag, `SHOW_TASK_PRIORITY_UI` (default `false`), and gate every priority indicator and control behind it. Leave the data model, API params, Zod validation (including the `"medium"` default), and the search filter DSL untouched. Reviving priority is a one-line flip of the flag back to `true`. **Alternatives considered** Deleting the priority code and schema outright. Rejected: it is irreversible, needs a data migration, and throws away a field the API and search still support. A gated flag keeps the change reversible and low risk. **Roadmap alignment** UI simplification. This is a presentation-only change; it does not alter core agent or data behavior. ## What Changed - Added `ui/src/lib/ui-flags.ts` exporting `SHOW_TASK_PRIORITY_UI: boolean = false` (typed `boolean` so gated branches are not flagged as dead code). - Gated the priority row in the Triage properties panel and the editable priority control in the issue detail header (plus its skeleton seed). - Gated the per-card priority icon in `KanbanBoard` and in `IssueThreadInteractionCard`. - Hid the priority chip and the mobile "more" menu priority section in the New Task dialog. The submit path still sends the `"medium"` default. - Removed the Priority options from the issues list Sort and Group-by menus; the comparator and grouping logic stay dormant. - Hid the Priority sections in the issue filters popover and in the search filter bar and sheet. The `priority:` search DSL and filter state stay functional at the data layer. - Suppressed active-filter priority pills for consistency. - Gated the "Tasks by Priority" dashboard chart and the design-guide priority showcase subsection. - Left activity-feed "changed priority" history text intact as a historical record. - Updated call-site tests to assert priority UI is absent while the flag is off, added focused hidden-surface tests, and added a test that proves creating a task still persists `priority: "medium"`. ## Verification - `pnpm check:token-gates` — all 3 gates clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `pnpm --filter @paperclipai/ui exec vitest run` on the touched surfaces (IssueProperties, IssueFiltersPopover, IssuesList, NewIssueDialog, IssueDetail, PriorityIcon and its interaction test) — all green under `TZ=UTC`. - Manual: with the flag off, priority does not appear in the detail header, Triage panel, New Task composer, Sort/Group/Filter menus, or the dashboard chart. Creating a task still persists `priority: "medium"`, and the `priority:` search token still filters at the data layer. ## Risks Low risk. The change is presentation-only and additive: no data model, API, validation, or search-DSL changes. The priority code paths remain compiled and tested; flipping `SHOW_TASK_PRIORITY_UI` to `true` restores the full UI. Visual snapshot baselines are intentionally not updated per the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, extended thinking enabled, with tool use / code execution in an agentic coding harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.806.0-canary.10 |
||
|
|
75acc4650f |
feat(sandbox-providers): pre-fill environment form with default sizing and image values (#11004)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runs execute inside sandbox environments. Sandbox provider plugins (Daytona, Modal, exe.dev, and others) declare a JSON-Schema `configSchema`. The Environment configuration form renders from that schema. > - The sizing and image fields open empty. Users must guess working values. The Modal form cannot submit at all until the user types an app name and an image by hand. > - The form renderer already pre-fills every field that declares a JSON-Schema `default`. The manifests do not use this mechanism for sizing or image fields. > - This pull request adds optional `default` values to the Daytona, Modal, and exe.dev manifest schemas. > - The benefit is a form that opens with known-good values. Users can create a working environment without provider research. ## Linked Issues or Issue Description **Current behavior** The Environment configuration form opens with empty sizing and image fields for the Daytona, Modal, and exe.dev sandbox providers. Users must find working values in provider documentation. Modal declares `appName` and `image` as required with no default, so the form blocks submission until the user invents both values. **Proposed behavior** The provider manifests declare JSON-Schema `default` values. The existing form renderer pre-fills them: - Daytona: CPU `4`, memory `4` GiB, disk `10` GiB, image `daytonaio/sandbox:0.8.0` - exe.dev: CPU `4`, memory `4GB`, disk `20GB` - Modal: app name `paperclip`, image `node:22` Secret-ref fields (API keys, tokens) get no defaults on purpose. The form persists a raw string in a secret-ref field as a company secret on save. A placeholder default would become a stored secret with a bogus value. Each plugin test suite now guards this invariant. **Reason and benefit** New users can create a working sandbox environment without guessing. The defaults stay optional: users can clear or change every value, and the schema marks no new field as required. The Modal image default `node:22` satisfies the sandbox runtime contract in `SANDBOX-REQUIREMENTS.md` (`node`, `sh`, and `tar` on PATH). **Subsystem affected** Sandbox provider plugins (`packages/plugins/sandbox-providers/*`): environment driver `configSchema` manifests. **Breaking changes** None. Defaults only seed the create-mode form. Saved environments keep their stored config. E2B, Novita, Cloudflare, and Kubernetes manifests do not change: E2B and Novita already default to their base templates, and the Cloudflare bridge and Kubernetes cluster fields have no sensible universal value. ## What Changed - Add `default` values for `cpu`, `memory`, `disk`, and `image` in the Daytona manifest. Trim the memory description to match. - Add `default` values for `cpu`, `memory`, and `disk` in the exe.dev manifest. - Add `default` values for `appName` and `image` in the Modal manifest. Extend the image description with the runtime-contract rationale. - Add manifest tests in all three plugins: defaults match expected values, defaults satisfy their own schema constraints, and no secret-ref field declares a default. - Bump plugin versions: daytona and modal `0.1.0` → `0.1.1`, exe-dev `0.1.1` → `0.1.2`. ## Verification - Run `pnpm test` in `packages/plugins/sandbox-providers/daytona`, `.../modal`, and `.../exe-dev`. The new `* manifest form defaults` suites pass. - Run `./node_modules/.bin/tsc --noEmit` in each of the three packages. Typecheck passes. - Manual: rebuild the plugins (`pnpm build` in each package), let the plugin dev-watcher refresh the manifest, then open Environments → New environment. The Daytona form shows CPU 4, Memory 4, Disk 10, and image `daytonaio/sandbox:0.8.0`. The Modal form shows `paperclip` and `node:22`. The API key fields stay empty. - Verified live on a local instance: the served `configSchema` in the plugin registry carries the new defaults, and existing environments are unchanged. ## Risks - Low risk. The change touches only manifest schema metadata and tests. No runtime code path changes. - New environments created with untouched forms now request 4 CPU / 4 GiB / 10 GiB from Daytona instead of provider minimums. This can raise cost per sandbox for users who previously saved empty fields. - The Daytona image default pins `daytonaio/sandbox:0.8.0`. The default needs a manual bump when Daytona ships new sandbox images. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Fable 5 (Anthropic, model ID `claude-fable-5`), via the Claude Code CLI, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.806.0-canary.8 |
||
|
|
f258b34bbd |
fix(ui): unify cloud-managed sign-out (#10994)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip can run as a self-hosted app or as a Cloud-managed tenant. > - These modes need different sign-out sequences because Cloud owns three sessions. > - Several visible controls implemented sign-out separately and could choose different paths. > - This pull request adds one Cloud-aware sign-out action and moves every visible control to it. > - The benefit is one safe sign-out path in Cloud and unchanged local sign-out in self-hosted deployments. ## Linked Issues or Issue Description Refs #2073. That older PR adds a separate company-settings sign-out surface. This change centralizes the existing account, company, and instance-settings surfaces and preserves the self-hosted behavior described there. **What happened?** Visible sign-out controls used separate implementations. A Cloud-managed control could call the app-local sign-out endpoint and open the local auth page. That path did not enter the Cloud-owned logout sequence for the tenant, Cloud, and identity sessions. **Expected behavior** Every visible sign-out control must use one action. Cloud-managed instances must navigate the top-level window to the same-origin `/cloud/logout` route without a local sign-out call first. Authenticated self-hosted instances must keep the local sign-out API and cache invalidation behavior. **Steps to reproduce** 1. Open a Cloud-managed tenant. 2. Use the account menu, company menu, or instance-settings sign-out control. 3. Observe that independently implemented controls can enter different sign-out paths. **Paperclip version or commit** The problem reproduces at `656ecfa585b31938e2685ffab3db22e794474803`. **Deployment mode** Cloud-managed tenant built from source. The regression tests also cover authenticated self-hosted mode. ## What Changed - Added `useSignOut` as the shared Cloud-aware sign-out action. - Navigated Cloud-managed sessions to `/cloud/logout` exactly once without calling local auth first. - Preserved local API sign-out and cache invalidation for authenticated self-hosted sessions. - Migrated the account menu, company menu, and instance general settings to the shared action. - Added focused tests for mode selection, menu closure, pending state, failure state, and settings behavior. ## Verification - `pnpm exec vitest run ui/src/hooks/useSignOut.test.tsx ui/src/components/SidebarAccountMenu.test.tsx ui/src/components/SidebarCompanyMenu.test.tsx ui/src/pages/InstanceGeneralSettings.test.tsx` — 26 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — passed. - `git diff --check origin/master..HEAD` — passed. ## Risks - Low risk. The change centralizes existing behavior and adds no schema, API, telemetry, or style-token changes. - Cloud mode depends on the existing health decision. Tests pin both mode branches. - The change does not alter Fetch Metadata, CSRF, cookie, prefetch, or return-URL protections owned by the Cloud logout route. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The runtime does not expose a more specific model ID or context-window size. The agent used high-reasoning mode, shell tools, and API tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.806.0-canary.9 |
||
|
|
52b8741b8e |
perf(server): cut steady-state DB hot paths in dashboard, attention, and productivity sweeps (#10992)
<!-- ASD-STE100 --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server keeps fleet health with periodic sweeps and shows a dashboard with run activity > - The Paperclip instance became slow again after the first round of recovery-sweep indexes landed > - Live profiling found four steady-state hot paths that read much more data than they use > - This pull request bounds the dashboard recursion, adds the missing taskKey index, and narrows two wide reads > - The benefit is a large drop in constant database load and a responsive server ## Linked Issues or Issue Description **Describe the bug** The server becomes slow while agents work. Live query sampling shows four hot paths: 1. The dashboard run-activity recursive CTE reads every run a company ever had on each call. One call takes 2.85 seconds. The UI calls it after almost every fleet event through the dashboard and sidebar-badges routes. 2. The productivity-review sweep runs each 30 seconds. Its run-scope filter is `issueId OR taskId OR taskKey` on the run context JSONB. No index exists for `taskKey`. The planner must detoast every run snapshot for the agent. One query takes 444 ms and the sweep makes one for each of ~152 candidate issues. 3. The attention failed-run section selects the full `context_snapshot` for every run newer than the oldest exhausted run. That fetch moves 29 MB for each feed build. 4. The retention sweep pages the attention feed with a cursor. Each page makes a full feed rebuild. **Expected behavior** Periodic sweeps and dashboard queries read only the data they use, and use indexes. **Actual behavior** The database stays saturated. Users see a slow server. ## What Changed - `server/src/services/dashboard.ts`: bound both arms of the `recovered_runs` recursive CTE to the chart window. A retry is always newer than the run it retries, so the bound cannot change visible chart data. Live time went from 2,852 ms to 54 ms. - `packages/db/src/migrations/0210_heartbeat_context_taskkey_index.sql`: add the `taskKey` expression index that completes the issueId/taskId/taskKey trio. With all three, the planner uses a BitmapOr. Live time for the productivity run-scope query went from 444 ms to 1.9 ms. - `packages/db/src/schema/heartbeat_runs.ts`: mirror the new index in the Drizzle schema. - `server/src/services/productivity-review.ts`: select only the seven run fields the evidence code reads. Before, the query pulled full rows with `result_json` (up to 43 kB per row, 100 rows per issue). - `server/src/services/attention.ts`: project `issueId`/`taskId` text fields instead of the full `context_snapshot` in the failed-run newer-runs query (29 MB per feed build before). - `server/src/index.ts`: the retention sweep now builds the attention feed once per company with `all: true` instead of one full rebuild per cursor page. - `packages/db/src/heartbeat-context-snapshot-index-migration.test.ts`: cover the new index and re-run migration 0210 statements to prove idempotency. ## Verification - `pnpm --filter @paperclipai/db typecheck` (includes migration numbering and safety checks) — pass. - `npx tsc --noEmit` in `server/` — pass. - `npx vitest run packages/db/src/heartbeat-context-snapshot-index-migration.test.ts` — pass (embedded Postgres, full migration chain, planner assertions, idempotent re-run of 0209 and 0210). - `npx vitest run` on attention, dashboard, productivity-review, decision-retention, issue-blocker-attention, and issue-review-attention test files — 72/72 pass. - Live EXPLAIN ANALYZE before/after numbers are in the What Changed list. ## Risks - Migration 0210 builds one btree index without CONCURRENTLY inside the transactional migration runner. The table is not in the large-table bucket. The 0209 twin built in seconds on a 100k-row live table. - The CTE bound excludes retry ancestors that are older than the chart window. Those rows are not visible to the chart query, so chart output does not change. - The attention projection changes JSONB scalar handling in one edge case: a non-string `issueId`/`taskId` value now casts to text instead of reading as absent. These keys are always strings in practice. - The retention sweep now holds one full feed in memory per company. The cursor loop already accumulated all pages into one array, so peak memory is unchanged. ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic, Mythos-class tier, extended thinking + tool use) via Paperclip agent runtime. - [x] I searched existing PRs and issues and this change is not a duplicate. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
656ecfa585 |
fix(server): keep Date fields intact through secret redaction; harden chat notice timestamps (#10984)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The task chat thread renders issue comments, system notices, and run
transcripts
> - The server routes comment payloads through the run-secret redaction
walker before it sends them
> - The walker rebuilds each object with `Object.entries`, and this
collapses `Date` instances to `{}`
> - The chat renderer then calls `.toISOString()` on an invalid date and
throws, and the thread falls back to the error banner
> - This pull request keeps `Date` instances intact in redacted
responses and makes the renderer safe against bad timestamps
> - The benefit is that task threads with system notices render
correctly again
## Linked Issues or Issue Description
**What happened**
Task threads that contain a system notice showed the banner "Chat
renderer hit an internal state error." in place of the conversation.
This occurred on many tasks.
**Expected behavior**
The thread renders all comments and system notices with correct
timestamps.
**Steps to reproduce**
1. Open a task that has at least one system notice comment (for example
a "Workspace ready" notice).
2. `GET /api/issues/{id}/comments` returns `createdAt: {}` for every
comment because the secret-redaction walker collapses `Date` objects.
3. The system-notice row calls `new Date({}).toISOString()`. This throws
`RangeError: Invalid time value` and trips the thread error boundary.
**Version / deployment**
Regression from #9934 (`e43f187ca`). It applies to all deployments that
include that commit.
## What Changed
- `server/src/services/run-secret-redaction.ts`:
`redactRegisteredSecretValues` now returns `Date` instances as-is. Dates
hold no redactable text, and the `Object.entries` rebuild turned them
into `{}`.
- `ui/src/components/IssueChatThread.tsx`: the system-notice row formats
its timestamp with a new `toValidIsoString` helper. A value that does
not parse as a date now degrades to "no timestamp" instead of a render
crash.
- Regression tests at three layers:
- Walker unit tests: `Date` values survive with and without registered
secret values.
- Route test: `GET /issues/:id/comments` serializes `createdAt` /
`updatedAt` as ISO strings.
- Render test: a system notice with a malformed `createdAt` renders
without the error boundary.
## Verification
- `npx vitest run --root server
src/__tests__/run-secret-redaction.test.ts` — 5 passed.
- `npx vitest run --root server
src/__tests__/issue-comment-redaction.test.ts` — 4 passed (embedded
Postgres route test).
- `cd ui && npx vitest run src/components/IssueChatThread.test.tsx
src/components/IssueChatThreadSystemNotice.test.tsx
src/lib/issue-chat-messages.test.ts` — 121 passed.
- Each new test was run against the unfixed code and failed there, which
confirms it guards the regression.
- A local sweep rendered 47 real issue threads through
`IssueChatThread`: 7 tripped the boundary before the fix, 0 after.
## Risks
- Low risk. The server change only preserves `Date` objects that the
walker destroyed before. String redaction behavior does not change, and
the registry-key stripping does not change.
- The UI change only affects the timestamp of system-notice rows and
omits it when the value is invalid.
## Model Used
- Claude Fable 5 (`claude-fable-5`), Anthropic. Agentic coding session
with tool use (file edits, shell, Vitest). No extended-context or
special reasoning mode.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
e591e75b8c |
perf(db): index context_snapshot/payload issueId lookups used by recovery sweeps (#10969)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server runs a recovery sweep every 30 seconds. The sweep makes
sure assigned issues do not stall.
> - The sweep reads the latest heartbeat run for each candidate issue.
It filters `heartbeat_runs` on `context_snapshot ->> 'issueId'`. No
index covers this expression.
> - Each lookup scans ~26k rows and detoasts each row's JSONB context
(~645 ms per query, measured with EXPLAIN ANALYZE). With ~460 candidate
issues, one sweep schedules ~500 seconds of database work every 30
seconds.
> - The database saturates. Users see the whole server as slow.
> - This pull request adds expression indexes so these lookups become
index scans.
> - The benefit is that the recovery sweep drops from ~500 seconds of
database work per tick to milliseconds, and the server becomes
responsive again.
## Linked Issues or Issue Description
**What happened?**
The server became slow for all users. `pg_stat_activity` sampling showed
the same two recovery-sweep queries active continuously (15/15 samples).
Each `getLatestIssueRun` call filtered `heartbeat_runs` on the unindexed
expression `context_snapshot ->> 'issueId'`, scanned ~26k rows, and
detoasted a 2.1 GB TOAST region. `heartbeat_runs` accumulated 10.8
billion sequential tuples read. `hasActiveExecutionPath` also scanned
`agent_wakeup_requests` (1.8M rows) on the unindexed expression `payload
->> 'issueId'`.
**Expected behavior**
Per-issue run lookups in the recovery sweep complete in milliseconds.
The sweep finishes well inside its 30-second interval. Background
maintenance does not degrade interactive latency.
**Steps to reproduce**
1. Run a board with several hundred issues in `todo` / `in_progress` /
`in_review` and a large `heartbeat_runs` table with big
`context_snapshot` payloads.
2. Let the heartbeat scheduler run its 30-second recovery sweep.
3. Observe `pg_stat_activity`: the per-issue `heartbeat_runs` lookups
run continuously; `EXPLAIN ANALYZE` shows a filter on `context_snapshot
->> 'issueId'` removing tens of thousands of rows per call.
**Paperclip version or commit**
master (
|
||
|
|
2ea22d6cea |
fix(ui): white text on light-mode user chat bubbles (#10952)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail view shows the conversation as chat bubbles. The requester's own messages sit in a solid accent-colored bubble. > - The bubble container sets `text-white`, but the message body renders through `MarkdownBody`. Tailwind prose tokens (`--tw-prose-body`) win over the inherited container color. > - `prose-invert` only lightens the prose text in dark mode. In light mode the prose body stayed its default dark color, so the text read as near-black on the blue bubble and was hard to read. > - This pull request maps the human bubble's prose tokens to the inherited text color in both themes. > - The benefit is that the requester's chat text is readable white-on-blue in light mode, and dark mode stays exactly as it was. ## Linked Issues or Issue Description <!-- No public GitHub issue exists; described in-PR per the bug report template. --> **What happened?** In light mode, the text inside the user's own chat bubbles in the task detail view rendered as dark (near-black) on the solid blue accent background. This made the requester's messages hard to read. **Expected behavior** The text inside the user's accent-colored chat bubbles should be white in light mode, matching the bubble's `text-white` intent. Dark mode already rendered correctly and should not change. **Steps to reproduce** 1. Open a chat-style task detail view in light mode. 2. Post a message as the requester (human) so it renders in the solid blue accent bubble. 3. Observe the body text renders dark on blue instead of white. **Paperclip version or commit** Reproduces on `master` (branched from `814cb3367`). **Agent adapter(s) involved** Not adapter-specific (core UI bug). ## What Changed - Add the existing `paperclip-markdown-on-accent` class to the human-branch `MarkdownBody` in `TaskChatBubble.tsx`. This class (already used by `IssueChatThread` for the same accent bubble) maps prose body/heading tokens to `currentColor`, so the text follows the bubble's `text-white` in both themes. - Apply the same class to the human-branch `MarkdownBody` in `TaskChatDescriptionBubble.tsx` (the description-as-first-bubble surface) for consistency. - Add unit tests covering that the human accent bubble carries the on-accent class and the agent/neutral bubbles do not. ## Verification - `pnpm check:token-gates` → 3/3 CLEAN. - `pnpm --filter ./ui vitest run src/components/task-chat/TaskChatBubble.test.tsx` → 9/9 passing. - Manual: in light mode, the requester's chat bubble text renders white on blue; agent/neutral bubbles unchanged; dark mode unchanged. This is a visual change. Snapshot baselines are intentionally not updated, per `doc/design/DECISION-SHEET.md` → "Per-change snapshot verification demoted to dormant (Jul 13 2026)". ## Risks Low risk. The change is scoped to the human-branch `MarkdownBody` className on two chat-bubble components and only remaps prose color tokens to the inherited text color. Agent and neutral bubbles are untouched, and dark mode behavior is unchanged. ## Model Used Claude Opus 4.8 (Anthropic), model id `claude-opus-4-8`, ~200K context window, extended thinking mode, with tool use / code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>canary/v2026.806.0-canary.7 |