mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 06:25:16 +02:00
7435b2ee9cc6c71e733fd77fee523d67b8b2badf
4369
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7435b2ee9c |
ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images that contain the native Rust Runner. > - The image already builds that Runner before copying ordinary app source. > - A Rust source change still invalidates its entire compiled dependency layer. > - Compiled dependencies can survive source changes when their recipe is unchanged. > - This PR adds a separate locked dependency build before compiling the real workspace. ## Linked Issues or Issue Description Refs #13195. A search of related Docker and Cargo cache PRs found no duplicate dependency-recipe change. **What existing behavior does this improve?** Docker image build time after Rust source or embedded protocol changes. **Current behavior** The `runner-build` stage compiles dependencies and workspace code in one layer. In Cloud readiness run 34698143548, that stage took about 3m48s when its cache was unavailable. **Proposed behavior** Generate a recipe with pinned cargo-chef 0.1.73. Build locked release dependencies in `runner-deps`, then copy and compile real Rust source and embedded protocol inputs in `runner-build`. Source edits can reuse the dependency layer from the existing registry cache. **Reason and benefit** Reduce dependency recompilation during source changes and merge bursts. Expected savings are roughly 2–4 minutes when the old native layer would miss but dependency layers are available. Full cold builds also pay for the recipe tool installation. Ordinary app-only cache hits gain little from this change. **Breaking changes** None to the shipped application or image tags. The recipe tool and compiled dependencies remain in build stages. ## What Changed - Install a pinned recipe generator with its locked dependencies and the existing package-owned compiler. - Add recipe planning and compiled dependency stages. Use the same release profile, package, binary, and lockfile enforcement as the real native build. - Remove generated source stubs before copying actual source. Preserve protocol inputs, timestamp normalization, binary staging, and application checks. - Add Docker cache wiring regressions and update the Docker cache documentation. - Run a two-build probe in Docker Runner check. It requires dependency reuse, changed real binary metadata after a source edit, and a changed recipe after a dependency declaration edit. It uses a disposable tracked-source context and exports only small metadata files. ## Verification - Passed all five Docker build-stamp and dependency-cache tests with `pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`. - Passed the local ARM64 `docker buildx build --target runner-build --progress plain`. Local Docker then hit storage errors during a runtime probe; cache invalidation verification continues on GitHub-hosted Linux. - Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`, and `git diff --check`. - Passed a [Linux AMD64 cache probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199) against the PR source: dependencies compiled in 3m49s for the baseline and were `CACHED` after a source edit; real source compilation took about 37 seconds. Binary metadata changed and dependency declaration changes altered the recipe. The permanent probe is also running in latest-head Docker Runner check. - Passed latest-head [Docker Runner check](https://github.com/paperclipai/paperclip/actions/runs/34711145160), including the permanent source/dependency invalidation probe. - Passed full [PR verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2): typecheck, all grouped tests, native verification, build, release dry run, and browser checks. One unrelated signoff-policy browser test failed waiting for a heartbeat run on attempt 1; only that failed shard and dependent checks were retried, and passed. - Latest-head Greptile is 5/5 with no unresolved findings. Full local tests/build were limited by local disk exhaustion; Linux CI completed those checks. ## Risks - The two-build CI probe has a 20-minute job limit to cover the cold build and source rebuild. It adds no AWS routing. - A fully cold build must install cargo-chef and populate the dependency layer. Both become reusable registry layers; no Actions cache is added. - The recipe and final build must keep the same compiler, build profile, package, binary, and directory layout. A source-change rebuild probe checks real cache reuse and binary invalidation. - Dependency or compiler changes still require rebuilding dependencies. Existing image verification and full-SHA publication gates remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.9 |
||
|
|
8df3ee2cf5 |
ci: rebalance serialized tests with current suite durations (#13328)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments wait for verified source commits. > - Verification splits serialized server tests across independent runners. > - The shard duration estimates came from August and no longer match current tests. > - Stale estimates put much more work on one runner than the others. > - This PR refreshes the estimates from a complete successful run to balance the existing runners. ## Linked Issues or Issue Description **What existing behavior does this improve?** The time spent waiting for the slowest serialized server-test shard in PR and release verification. **Current behavior** In [Cloud readiness run 34705914878](https://github.com/paperclipai/paperclip/actions/runs/34705914878), the five serialized shards spent 479, 371, 322, 390, and 365 seconds running tests. The recovery suite had a 55-second estimate but now takes about 156 seconds including process overhead. **Proposed behavior** Use fresh per-suite measurements with the existing deterministic duration balancer. Applying the same measured costs to the new assignment gives 385, 385, 386, 385, and 385 seconds. This predicts about 93 seconds less waiting for the slowest shard, before runner/setup overhead. Live CI will confirm the result. **Reason and benefit** Use the existing runners more evenly. No extra runner, test parallelism, cache, timeout, or routing change is needed. **Breaking changes** Suite-to-shard assignments change. The full suite set, assertions, and per-suite process isolation stay the same. **Additional context** Searched related CI and shard PRs. This updates the existing duration manifest, without duplicating a pending sharding implementation. ## What Changed - Refresh all 145 serialized suite weights from the same successful release verification run. - Record source job IDs and the measurement method in the manifest. Durations include process startup, imports, collection, tests, and shutdown. ## Verification - Passed all 30 shard and release-workflow tests: `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Confirmed every measured suite appears exactly once across all five source logs. - Compared old and new assignments using the same measured weights. The maximum fell from 478862ms to 385551ms. - In [PR CI run 34710696242](https://github.com/paperclipai/paperclip/actions/runs/34710696242), all five serialized jobs passed in 7m05s–7m27s including setup. The measured assignment is now balanced in a live run. - The same run passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Local full-suite verification on this base was limited by disk exhaustion; local typecheck and targeted shard tests passed. - Latest-head Greptile is 5/5 with no open findings. Every current-head CI check must be green or intentionally skipped before merge. ## Risks - Individual durations vary with load and future test changes. These estimates affect assignment only; missing or renamed suites receive the existing median weight. - Both PR and release verification read this manifest, so both receive the new assignments. Each suite still runs in its own serialized Vitest process. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed50a39c3f |
fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a59f5a8adc |
fix(server): stop reporting expected managed-cloud transients to Sentry (#13323)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The server reports crashes to Sentry so operators can find real
faults
> - Three expected conditions report as crashes: a client that closes
the connection mid-request, one stale pooled database socket after a
pooled endpoint recycles, and the short boot window where a supervised
cloud stack runs a new app image before its migration runner has caught
up
> - These events arrive in the hundreds and bury real errors
> - This pull request classifies each condition as expected and stops
the Sentry capture for exactly that condition, with behavior unchanged
everywhere else
> - The benefit is a Sentry feed where each event is a real fault
## Linked Issues or Issue Description
**What happened?**
Three noise classes fill the backend Sentry project on managed cloud
fleets:
1. `Error: aborted` (ECONNRESET) reports as a 500 crash when a client
closes the tab or loses its network mid-request. Observed 18 times in
one week from routine client disconnects.
2. `Error: write CONNECTION_CLOSED ...` reports from many query paths
after a pooled Postgres endpoint suspends. The existing single retry in
cloud actor resolution still fails, because a suspended endpoint kills
every pooled socket at once and the one replay draws another dead
socket.
3. `Error: PostgreSQL has pending migrations (...). Refusing to start`
reports from every supervised stack during a fleet upgrade. The
supervisor delivers the new app image before it runs the migration
runner, so each stack crash-loops briefly by design. One fleet roll
produced 329 events (11 per container).
**Expected behavior**
A client disconnect ends the request quietly. A transient dead socket is
replayed until a live socket answers. A supervised mid-upgrade boot
refusal logs and exits nonzero without a Sentry capture, while the same
refusal on a self-hosted deployment keeps reporting.
**Steps to reproduce**
1. Abort an HTTP request mid-flight: the error handler reports a crash
to Sentry.
2. Suspend a pooled Postgres endpoint under an idle server, then issue
two quick requests: the first replay can draw a second dead socket and
surface `CONNECTION_CLOSED`.
3. On a deployment with `PAPERCLIP_CLOUD_API_ORIGIN` set, add a
migration file without running the migration runner and boot: the
refusal reports to Sentry.
**Paperclip version or commit**
master (
|
||
|
|
24007a9804 |
fix: promote review tasks when continuations start (#13318)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service starts and tracks agent work on issues. > - An issue can be in review before a user comment starts more agent work. > - The previous checkout rule did not claim an issue from review for this continuation. > - The issue could stay in review while an agent actively worked on it. > - This pull request lets a resolved interaction continuation claim an issue from review. > - The existing checkout update then sets the issue to in progress. > - The benefit is that issue status now shows active agent work correctly. ## Linked Issues or Issue Description **What happened?** An issue stayed in review when a new agent continuation started work on it. **Expected behavior** The issue must move to in progress when an agent starts work. An idle issue must stay in review. **Steps to reproduce** 1. Put an assigned issue in review. 2. Resolve an interaction that starts a continuation. 3. Start the heartbeat run. 4. Observe that the issue remains in review while the run works. **Paperclip version or commit** The problem reproduces on the master branch before this change. ## What Changed - Allow a resolved interaction continuation to claim an assigned issue from review. - Keep idle review issues unchanged. - Add regression tests for both behaviors. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/heartbeat-auto-checkout.test.ts`. - Confirm that all four tests pass. ## Risks - Low risk. The change only expands checkout eligibility for resolved interaction continuations. - The existing guarded checkout update still controls ownership and the status update. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.5. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9021c6721 |
fix: require explicit native completion reviews (#13314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their outcome through paperclip_finish. > - The server previously turned incomplete reports into human approval requests. > - Those requests could block a later successful run, even when no person had requested review. > - This pull request creates review cards only for explicit attention requests and withdraws proven old fallback cards. > - Agents receive useful completion feedback, while explicit approval gates and task state protections remain in force. ## Linked Issues or Issue Description Related: #13266 removed reviews caused by policy upgrades. This change removes the separate completion fallback. **What happened?** An agent reported needs_review while waiting for checks without requesting a human decision. Paperclip created a generic Native completion review. A later successful report could not complete the task because that old card remained pending. **Expected behavior** Ordinary low-risk work completes after a valid done report, a successful run, and workspace finalization. Incomplete work stays with the agent. Explicit approval requests remain visible and must be resolved. **Steps to reproduce** 1. Complete a native run with needs_review and no attention requests. 2. Continue the task and submit a successful done report. 3. Observe that the old implementation leaves the task in review behind a generic confirmation card. ## What Changed - Require explicit attention requests to create native review cards. Route each request independently and preserve pending or declined decisions. - Withdraw only pending system cards with matching old decision, assessment, effect, contract, and prompt provenance. Preserve history and explicit or answered requests. - Reassess an affected current result without overwriting later task edits, runs, contracts, or workspace failures. - Return pending approval links and required actions through the completion tool. Reject contradictory done reports and empty review requests before accepting a result. - Allow one corrective continuation for incomplete results, then expose a recovery action. - Update status fixtures, database regressions, runner tests, and the completion contract documentation. ## Verification - `pnpm -r typecheck` passed after merging current master. - `pnpm build` passed after merging current master. - The combined branch passed 89 completion and Agent Chat tests. Other targeted tests passed: 170 external-chat and reconciliation tests; 50 runner-resume and control-plane tests; 13 arbiter tests; 7 chat delivery tests; 21 runner completion and runtime-context tests. - The full local test attempt exposed old review fixtures and a missing fake-provider binary. The fixtures are fixed and the helper is built. All affected suites pass in fresh reruns. The timing-sensitive Discord test also passed on rerun. - All latest-head CI checks passed, including build, typecheck, general and serialized tests, runner verification, browser tests, and canary dry run. Greptile is 5/5 with zero unresolved comments. ## Risks - Cleanup changes existing pending cards. It requires exact system provenance and only applies to low-risk agent-claim contracts. It does not delete history or dismiss explicit requests. - Status still commits after the turn and workspace finalization. Completion feedback reports current constraints and does not claim an early status commit. - Incomplete reports now request a bounded corrective run instead of an automatic approval. Repeated failures expose recovery. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and test tools. The exact deployment identifier and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.8 |
||
|
|
0e14c61da7 |
fix: fence native startup against cancellation (#13316)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can pause a task while its runner prepares to start. > - Cancellation must prevent preparation from creating new execution authority. > - Native runtime selection could run after cancellation and leave an unclaimed recovery coordinator. > - Saved user messages then waited for recovery that had no eligible worker. > - This pull request fences startup and lets explicit user continuation settle verified, unclaimed startup state. > - Tasks can continue after cleanup while keeping provider ownership and execution safeguards. ## Linked Issues or Issue Description Refs #13285 for startup controller leases and #13270 for explicit continuation and saved-message recovery. Related #13293 covers retained processes that actually started; this change covers cancellation before the native provider claim. Related #13315 covers legacy queued-message delivery. **What happened?** Pausing a task during startup could cancel its heartbeat before native runtime selection. Stale preparation then created an observed native coordinator on the cancelled run. The coordinator had no provider result or eligible recovery worker. A later Continue message stayed queued indefinitely. **Expected behavior** Cancellation fences native startup. After verified cleanup, a newer user message starts one fresh conversation turn. An unverified execution keeps its hold and a clear explanation. **Steps to reproduce** 1. Start a task with the native runner and delay startup preparation before runtime selection. 2. Pause the task, then release preparation. 3. Resume the task and send Continue. 4. Before this fix, native selection can persist after cancellation and block the saved message. 5. Repeat from persisted cancelled startup state after a server restart. **Paperclip version or commit** Reproduced against master at `586b5ec82` with isolated PostgreSQL regression fixtures. **Deployment mode** Self-hosted server built from source, with Paperclip Runner. ## What Changed - Serialize the cancellation fence and native runtime selection on the run row. Revalidate the startup controller lease. - Refresh the runtime before dispatching cancellation, and reject terminal or cancelled runs at the native provider claim. - Recognize never-claimed coordinators only after startup and environment cleanup are verified. Reject process, provider, owner, and conflicting launch evidence. - Settle that coordinator atomically with a new authenticated user turn. Preserve history, unknown outcomes, and attempt counts. - Reuse the saved-message worker after restart and retain pause, budget, approval, and ownership gates. - Add startup, restart, duplicate-admission, and negative-proof regressions. Document the rule. ## Verification - All 687 tests pass across the complete heartbeat recovery, explicit continuation, and native session executor suites on the rebased branch. - Six focused race regressions also pass: cancellation before and after native selection, Stop racing adapter registration, process termination during a database failure, and a run finishing during cancellation. - `pnpm -r typecheck` and `pnpm build` pass after rebasing on master at `ab15aff39`. - Greptile is 5/5 on `f3dea2ab27facdf0360b56172ebd3e5219e25538`, with no open review threads. Policy and security checks pass. - All 32 CI checks pass on the final head, including the full server/workspace test matrix, all three browser shards, runner verification, typechecks, build, and canary dry run. The two optional Storybook jobs are skipped. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34697945514). - Local full-suite limitation: the broad `pnpm test:run` attempt reported two failures outside the changed area after unusually long test durations (about 65 seconds for supporting skill-file saves and 933 seconds for setup-token login). Both cases passed isolated reruns, with no code changes. The broad local run was stopped after CI completed successfully; no clean full local-suite pass is claimed. ## Risks Cancellation and startup overlap. The run and coordinator locks provide the authority fence; cleanup and process evidence provide the containment proof. Historical runs without sufficient evidence remain blocked. A saved user message authorizes a fresh turn, not automatic replay. No schema migration or dependency change. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and tests. This session does not expose the exact backend revision or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.7 |
||
|
|
9132e8279f |
chore(skills): allow verified PR merges (#13313)
## Thinking Path > - Paperclip uses repository skills to guide agent work. > - The PR preparation skill controls the pull request workflow. > - One hard rule prevents an agent from merging a verified pull request. > - The requested workflow must permit that action when other authority allows it. > - This pull request removes only that one rule. > - All other PR safety and review rules stay unchanged. ## Linked Issues or Issue Description **What existing behavior does this improve?** The `prepare-paperclip-pr` agent workflow. **Current behavior** The skill always prohibits the agent from merging a pull request itself. **Proposed behavior** The skill no longer adds that universal prohibition. Other permissions and workflow rules still apply. **Reason and benefit** An authorized workflow can merge a verified pull request without conflict with this skill. **Breaking changes** The skill no longer blocks every agent-initiated merge. ## What Changed - Remove the single universal no-merge rule from the PR preparation skill. ## Verification - Confirm the pull request changes one file and deletes one line. - Run `git diff --check` against the pull request commit. ## Risks - An agent can merge when its other instructions and permissions allow it. - Existing review, CI, and work-preservation rules remain in place. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.6 |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.5 |
||
|
|
e830180139 |
chore(lockfile): refresh pnpm-lock.yaml (#13279)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
586b5ec828 |
fix(ui): improve mobile task spacing (#13304)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task detail view is the main place for task conversation. > - The mobile view used too much horizontal space for gutters and controls. > - The status control and identifier also reduced the title width. > - This pull request gives the title its own mobile row and reduces mobile padding. > - The benefit is a clearer task view with more space for useful content. ## Linked Issues or Issue Description **What existing behavior does this improve?** The mobile task detail header, conversation list, and message composer. **Current behavior** The title shares one row with status and identifier controls. The conversation and composer also use larger mobile gutters than needed. **Proposed behavior** The title uses the full mobile width. Status and metadata use the next row. The conversation and composer use smaller mobile gutters. **Reason and benefit** Long titles have more readable line lengths. The reduced padding gives task content more room on small screens. **Breaking changes** None. ## What Changed - Put the task title on a full-width mobile row. - Move the mobile status and metadata below the title. - Reduce the mobile conversation and composer padding. - Update the composer dock class test. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/task-chat/TaskChatComposer.test.tsx ui/src/pages/IssueDetail.test.tsx --reporter=dot` - `pnpm --filter @paperclipai/ui build` - `pnpm check:token-gates` - `git diff --check` ## Risks - Low risk. The layout changes apply only at the mobile breakpoint. - Desktop spacing stays unchanged. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f40b4ad4b |
fix(connections): repair and simplify Google Workspace setup (#13289)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - App connections give agents controlled access to external services.
> - Managed Google Workspace methods depend on profiles enabled for each
Paperclip instance.
> - The catalog used those profiles, but connection creation used static
availability and rejected enabled connections.
> - Switching capabilities also reset personal ownership to company-wide
ownership.
> - This PR uses the same availability rules for catalog and setup, and
preserves supported ownership choices.
> - Paperclip is the default Google authentication method. A small link
opens custom OAuth settings when needed.
> - Local and cloud instances can start the selected Workspace
connection without broadening its audience.
## Linked Issues or Issue Description
**What happened?**
A fresh enrolled instance showed managed Google Workspace apps as
available. The issue was first found with Gmail. Selecting Continue to
sign in returned HTTP 422 with “This app does not have an available
connection method.” Changing from Read & create drafts to Read only also
reset Just me to company-wide ownership.
**Expected behavior**
Start Google authorization for an enabled profile. Keep personal
ownership when the selected method supports it. Reject profiles that the
instance cannot use.
**Steps to reproduce**
1. Start a fresh source test-drive and enroll the instance with
Paperclip Cloud.
2. Open Apps and select Gmail.
3. Select Just me and choose the agents that can use the connection.
4. Continue to setup and select Read only.
5. Go back to inspect ownership, then continue to sign in.
**Paperclip version or commit**
Reproduced on master commit
canary/v2026.912.0-canary.4
|
||
|
|
eb9f954bae |
fix(ui): stabilize steered chat activity presentation (#13246)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task chat shows live agent work and user steering. > - Steered messages showed a second timestamp that did not match normal chat messages. > - Folded live work also changed height as new commands arrived. > - Status carets and dots did not use the same vertical alignment. > - This pull request gives these states one stable presentation. > - The benefit is a task chat that is easier to read while an agent works. ## Linked Issues or Issue Description **What happened?** Steered chat messages showed a separate queued or steered timestamp. Folded live activity could grow and shrink. Some carets and status dots did not align vertically. **Expected behavior** Steered messages use the normal timestamp. Folded live activity shows one latest line until the user opens it. Carets and dots use the same vertical center. **Steps to reproduce** 1. Open a task chat with a running agent. 2. Steer the chat and inspect the message timestamp. 3. Keep live activity folded while new commands arrive. 4. Compare the caret with the status dot in the Reconnecting state. **Paperclip version or commit** Current `master` before this change. **Deployment mode** Local development mode. ## What Changed - Use the normal timestamp for steered chat messages. - Keep folded live activity to one line and replace it with the latest activity. - Keep the full activity timeline available after explicit expansion. - Vertically align carets and status dots, including Reconnecting. - Add focused component tests and Storybook review states. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/task-chat/TaskChatRunnerTurn.test.tsx src/components/task-chat/TaskChatStatusPill.test.tsx src/components/task-chat/task-chat-adapter.test.ts` (63 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build-storybook` ## Risks - Low risk. The change is limited to task chat presentation and its tests. - A live activity item can replace the folded text. The expanded timeline keeps the full history. > This fix does not overlap with planned work in `ROADMAP.md`. ## Model Used - OpenAI Codex. The runtime did not expose the exact model ID or context window. The agent used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.912.0-nightly.0 canary/v2026.912.0-canary.3 |
||
|
|
30aa3740ba |
fix(ui): make feed cards fully clickable (#13294)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The activity feed shows cards for recent work. > - A feed card showed a linked entity, but only some card content behaved as one clear click target. > - Users expect the image, title, and blank card space to open the entity. > - This pull request makes the full visible card use the entity link. > - It also adds a regression test for the full-card click target. > - The benefit is a larger and consistent navigation target. ## Linked Issues or Issue Description **What happened?** Clicks on some feed card content or blank card space did not reliably open the linked entity. **Expected behavior** A click anywhere in a linked feed card opens its entity quick view. **Steps to reproduce** 1. Open the activity feed. 2. Find a card that links to an entity. 3. Click its image, title, or blank space. 4. Observe that only part of the card acts as the link before this change. **Paperclip version or commit** Current `master` before this pull request. ## What Changed - Made each linked feed card anchor use the full available width. - Made the visible card width account for its existing page margins. - Added a component test that clicks the card and confirms link navigation. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/FeedCard.test.tsx` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui typecheck` ## Risks - Low risk. The change only expands the existing link click target. - The width calculation can affect feed card layout. The regression test checks the expected width classes. > This bug fix does not add a roadmap feature. ## Model Used - OpenAI GPT-5.6, Codex agent, with reasoning and code execution tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.2 |
||
|
|
09e208c54f |
ci: activate shared PR dependency cache restores (#13302)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - PR checks share a repository cache budget with post-merge cloud verification. > - Per-PR dependency store copies consume about 700 MB each and displace useful build caches. > - The reviewed workflow now restores shared dependency caches without saving PR copies. > - The public entry point must pin an approved immutable workflow before it can use AWS runners. > - This PR activates the merged workflow after the runner group accepted its exact SHA. > - Native verification and the app build also run in parallel, as defined in the merged workflow. ## Linked Issues or Issue Description Refs #13300. Refs #13301. The cache definition is merged, and its exact SHA is authorized in the restricted runner group. This PR activates it. Its base also contains the two synthetic fixture fixes from #13301. No duplicate activation PR was found. ## What Changed - Advance the `pr-trusted.yml` pin from `03609aa6ecc9a047ed53d6b6469d8be554fbc46d` to merged commit `44dde2dec42a22746a2f36b595acacc9ccfa1df6`. - Update the adjacent pin comment to identify #13300 and the activated behavior. - Activate restore-only pnpm caching and removal of redundant Node setup steps from #13300. - Activate the already-merged separate native verification job, required by the aggregate `verify` check. The app build no longer waits for native verification inside the same job. - Activate the merged module-boundary and source-verification checks and explicit Runner Evalbook viewer build. ## Verification - Passed all 505 workflow, routing, cache, sharding, and source-verification tests. - Passed `actionlint` for both workflow files and `git diff --check`. - Passed the runner infrastructure's workflow routing regression tests. - Confirmed that the gate is unchanged between the old and new workflow definitions. The restricted runner group permits the exact merged SHA and retains its existing workflow pins. - Full workspace typecheck and build passed on the timeout-fix branch. Its 107 targeted runner tests and all Linux CI groups passed. The broad local Mac test command had 15 unrelated permission and missing-skill-path failures, documented in #13301; it stopped before later groups. - [Current-head CI run 34662664879](https://github.com/paperclipai/paperclip/actions/runs/34662664879) passed. All 33 checks are green or intentionally skipped at `ded4f1d093644b9279cdf56457f69a196fcf8cc2`. Greptile is 5/5, with no unresolved findings. - The live build restored the master pnpm cache, downloaded the resolved lockfile artifact, and completed a frozen install. It had no dependency-cache upload step. GitHub reports zero cache entries for this PR. Native verification and the app build started together and both passed. - [Post-merge cloud readiness for #13301](https://github.com/paperclipai/paperclip/actions/runs/34662389233) passed all 30 jobs in 12m42s from merge. Native verification restored the master Rust cache and passed the formerly failing fixtures. ## Risks A new PR-only dependency can require a download until a master cache contains it. The new pin also activates the merged workflow changes listed above, so the live PR must pass all checks. The repository storage cap is still 10 GB; this PR does not increase it. No allowlist or runner permission logic changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.1 |
||
|
|
5cc7784986 |
test: remove cold executable reads from runner integrity deadlines (#13301)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runner protocol tests verify that invalid authenticated input fails closed. > - Cloud readiness requires those tests to pass before a source commit can deploy. > - Two test fixtures hash the host Node executable twice even though its process launcher is synthetic. > - Cold reads of a large Linux executable consume time unrelated to protocol failure handling. > - This PR gives both synthetic runner fixtures tiny real artifacts and keeps their existing deadlines. > - Real authentication, encrypted frames, and request and notification failures remain covered. ## Linked Issues or Issue Description **What happened?** Cloud readiness repeatedly failed in `DurablePrpControlPlane > promptly fails the real transport request and notification paths on authenticated bad semantic input (throwing observer: false)` with `Test timed out in 5000ms`. The image built, but the failed source check prevented deployment. See runs [34656885170](https://github.com/paperclipai/paperclip/actions/runs/34656885170) and [34657111560](https://github.com/paperclipai/paperclip/actions/runs/34657111560). The first case took 5.6–11.4 seconds; the following case took about 0.4 seconds. The original test reads and hashes `process.execPath` once itself and once through the real transport. The Node executable is 122,678,944 bytes in the local Linux Node 24 container, versus 68,672 bytes on the development Mac. Instrumented Mac runs pass; cold executable reads are a likely cause of the CI-only timeout. **Expected behavior** The deadline should measure real protocol rejection and consumer failure, without reading a large unrelated executable as test fixture data. **Steps to reproduce** 1. Run the named test with the real transport and synthetic process launcher. 2. For a deterministic probe, inject a six-second delay into the first `readFileSync(process.execPath)` call. 3. The original test exceeds its existing five-second deadline. Both cases pass after this change because neither reads the host executable. The probe is temporary instrumentation, not part of this commit. **Paperclip version or commit** f12b647ae; the same failure also occurred on |
||
|
|
44dde2dec4 |
ci: reuse dependency caches without per-PR uploads (#13300)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud releases wait for source verification before deployment.
> - That verification reuses compiled Rust dependencies to finish
sooner.
> - PR jobs save large pnpm stores under separate merge refs and
different lockfile keys.
> - Those copies compete with master build caches for the repository's
10 GB cache limit.
> - This PR makes PR dependency caches restore-only and reuses
master-compatible keys.
> - A separate pin update will activate the reviewed workflow.
## Linked Issues or Issue Description
**What happened?**
PR merge refs accumulated roughly 700 MB copies of the same pnpm store.
Master Rust caches disappeared, and Cloud readiness run
[34656098157](https://github.com/paperclipai/paperclip/actions/runs/34656098157)
rebuilt dependencies after cache misses. The repository currently has a
10 GB limit. GitHub rejected a request for 50 GB; that setting needs
separate organization/billing access.
**Expected behavior**
PR jobs should reuse downloaded packages without evicting post-merge
compilation caches through duplicate uploads.
**Steps to reproduce**
1. Run several PRs while the checked-in lockfile needs policy
regeneration.
2. Compare the setup-node keys in PR jobs and master jobs.
3. List Actions caches by ref, key, and archive size. The PR keys repeat
across merge refs.
**Paperclip version or commit**
|
||
|
|
fe38085174 |
fix(ui): hide profile feedback flag on Cloud (#13292)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The sidebar provides account controls and a feedback link. > - Cloud now has Plain chat through the snippet support in #13168. > - The feedback flag creates a second feedback route on Cloud. > - This PR uses the existing Cloud metadata to hide the flag. > - Self-hosted users keep the existing feedback link. ## Linked Issues or Issue Description Refs #13168 and #10850. Related: #12778 proposes a separate native feedback relay; this PR only hides the existing flag on Cloud. **What happened?** Cloud users see both Plain chat and the profile feedback flag, which opens the external feedback form. **Expected behavior** Cloud users use Plain chat. Self-hosted users keep the profile feedback flag. **Steps to reproduce** Open a Cloud instance with Plain enabled and expand the sidebar. Both feedback controls are visible. ## What Changed - Use `useCloudInstance()` in both account-menu variants. - Render the feedback flag only when the sidebar is expanded and the instance is not Cloud. - Extend existing tests for Cloud and authenticated self-hosted behavior. ## Verification - Account-menu tests: 6 passed (`cd ui && pnpm exec vitest run src/components/SidebarAccountMenu.test.tsx`). - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - Full typecheck: failed in unchanged `plugin-workspace-diff` code (`editStateKey` required by `UseFileDiffInstanceProps`). - UI build: blocked by four missing exports in the installed `assistant-cloud` dependency, including `createRunReport`. - Repository build: failed on the same unchanged workspace-diff plugin types. - Full test suite: started, then stopped after broader validation hit dependency failures. Only the six targeted account-menu tests are claimed as passed. - Fresh installation used `--no-frozen-lockfile --lockfile=false` because the base lockfile and server manifest differ. No dependency or lockfile changes are included. - After deployment, check that Cloud has no profile feedback flag and self-hosted still links to `https://paperclip.ing/feedback`. ## Risks Low risk. The hook reads the existing health query cache and adds no request. CloudAccessGate loads health before the main UI mounts. The flag is hidden on all Cloud instances, so operators must keep Plain configured. No backend, settings, or Plain integration changes. ## Model Used OpenAI GPT-6 Codex, with code inspection and command execution. The exact model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9b7bd41833 |
feat(ui): refine dashboard cards, charts, and recent lists (#13269)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use the dashboard and Live runs page to inspect current and recent runs. > - Large transcript cards take space and make it hard to compare run states. > - A compact card must show the agent, linked task, and time, with access to run details. > - This pull request applies the supplied card design to both views. > - Operators can scan more runs and open a run or task for details. ## Linked Issues or Issue Description **What existing behavior does this improve?** The agent run cards on the dashboard and Live runs page, plus dashboard charts and recent lists. **Subsystem affected** `ui/` — React board UI. **Current behavior** Both views show large cards with embedded transcripts. Queued and running cards share a live treatment. The dashboard chart grid leaves an empty column, recent lists use different row sizes, and some activity labels expose raw event names. **Proposed behavior** Both views use the same compact grid. Each card shows the existing initials avatar, the agent name, one linked task row, and a right-aligned timestamp. There is no status chip next to the agent name. The timestamp uses small Inter text and the same muted gray as dashboard metric descriptions. Recovery details remain in the task view. Run status remains in the header tooltip and accessible label, and run transcripts remain in the run detail view. The shared in-progress task icon is now an animated circular spinner across the app. **Reason and benefit** Operators can scan task states without scrolling through embedded output. The in-progress spinner keeps the shared task icon sizes, circular shape, and stroke width. It honors reduced-motion preferences. **Breaking changes** The dashboard and Live runs page no longer show transcripts inside cards. Open the agent header to inspect a run. There are no API or database changes. **Additional context** Related public work found during the duplicate search: - Refs #4317 — related run-state display concerns. This change does not change server state reconciliation. - Refs #11394 — related live-run cache updates. This change retains the existing data flow. - Refs #2118 — related navigation from agent cards. The card header here opens the specific run. ## What Changed - Apply compact cards to the dashboard and Live runs page, with status icons and theme tokens. - Use 16 × 16px status icons in the shared cards, including the missing-task clock. - Keep the existing circular initials avatar. Remove the status chip beside the agent name while preserving the run label for tooltips and accessibility. - Replace the shared in-progress task glyph with an animated open circle. Use the same 10-unit radius and 2-unit stroke as the other task glyphs, with a reduced-motion guard across all task status surfaces. This is the requested workflow-status indicator across the app, including between runs; live indicators separately report active execution. - Use the same Tailwind blue tokens as nav dots and Live labels for the progress icon: blue-400 in dark mode and blue-600 in light mode. Covered-blocked icons follow the same blue token. - Match the Tasks by Status chart's In Progress bar and legend to the spinner's theme-aware color token. - Distribute the three visible dashboard charts across equal-width desktop columns, preserving existing gaps, card padding, and page margins. Retain four columns when the optional priority chart is enabled. - Use the done task icon's `--status-task-icon-done` token for green bars in Run Activity, Tasks by Status, and Success Rate, including their legends. - Use the blocked task icon's `--status-task-icon-blocked` token for red bars and legends in Run Activity, Tasks by Status, Success Rate, and priority charts. - Align Recent Tasks live counts to the sidebar's right edge. Reveal the ellipsis over a fading row surface on hover, keyboard focus, and while its menu is open, without moving the label or count. - Keep the ellipsis backdrop synchronized with the row background during fade-out, using the actual sidebar surface in both themes to prevent a darker flash. - On coarse-pointer devices, reserve space for the always-visible ellipsis and disable its fade so the live count remains readable beside it. - Order desktop dashboard Recent Tasks as status icon, task title, agent avatar/name, 11px monospace task ID, then timestamp. In narrow cards, keep the ID right-aligned beside the truncated title and place the agent and timestamp on the second line. - Give both dashboard lists the same 48px desktop rows, 24px leading slots, 80px desktop task ID columns, and 64px right-aligned timestamp columns. Use an 8px icon-to-content gap matching the task detail heading. Narrow layouts use matching 76px rows with intrinsic-width IDs on the title line; activity timestamps stay below. - Order dashboard Recent Activity as actor avatar, actor name, verb, task title, 11px monospace task ID, then timestamp. Use direct verbs such as “Board read …” instead of “issue read marked”. Truncate titles to preserve IDs and timestamps; keep full names and titles in tooltips. - Use 12px task status icons in the top breadcrumb and Properties status row. Keep the main task title icon at 20px. - Reduce monospace task IDs to the 11px micro type token in breadcrumbs and shared run cards, aligning them to the adjacent task titles' text baseline while keeping status icons centered. - Show the task title and identifier in one bordered row. Keep missing-task links usable and show lookup failures. - Remove the recovery chip and place the timestamp below the task row. - Keep the link to all runs available when the dashboard has four or fewer runs. - Avoid transcript polling for compact cards. - Fix bundled Inter font URLs and update the design guide and Storybook fixtures. - Cover run navigation, task status, failed lookups, shared icon sizes and stroke, and reduced-motion-safe animation in tests. ## Verification - Latest review fixes: targeted SidebarRecentTasks, SidebarNavItem, StatusGlyph, StatusIcon, ActiveAgentsPanel, activity-format, ui-font-assets, and Dashboard suites — 65 tests pass, including read/unread verbs and reduced-motion-safe task animation. - `pnpm --filter @paperclipai/ui typecheck` — passes. - `pnpm check:token-gates` — all gates pass. - `git diff --check` — passes. - Related component suites passed during development: run cards, status glyphs, breadcrumbs, issue properties, charts, sidebar navigation, recent-task actions, and settings sidebar. - Full repository `pnpm -r typecheck` and `pnpm build` pass on the final commit using the temporary Rust toolchain. Storybook build passed earlier in preparation. - All CI gates pass on `c85e949b12b471d425d216caa609aee41dc58bde`: all general/workspace and serialized server test shards, all three browser shards, build/runner verification, typecheck/release registry, canary dry run, policy, and security checks. The duplicate local `pnpm test:run` was stopped after CI completed the same test suites; it is not claimed as a completed local pass. The two Storybook jobs are skipped by the configured draft-PR workflow. - Greptile reviewed the final commit at 5/5 with no actionable findings and zero unresolved review threads. - Browser measurements at 1654px, 860px, and 390px confirm matching 48px desktop and 76px narrow rows (plus 1px dividers), right-aligned IDs on the title line, 64px timestamp slots, and no horizontal overflow. Narrow Recent Tasks rows retain readable agent names below the title. Activity reads as actor, verb, title, ID, time. Additional reference details can expand an activity row. - Verified three equal-width desktop charts with 16px gaps and aligned outer edges. Chart blues, greens, and reds resolve to the matching task-icon tokens. - Verified real queued, running, succeeded, failed, cancelled, and timed-out runs in the dev dashboard. Checked compact cards, task/run links, timestamps, light/dark themes, and narrow layout. - To inspect: open Dashboard, then Live runs. Both pages use the same small cards. Open an agent header for run output, or the task row for task details. - Verified the updated spinner on the real dashboard cards, Recent tasks, and task list. Checked actual SVG circle radius, rendered size, and stroke width. Confirmed the spinner, nav dots, and Live labels resolve to the identical blue-400 color in the dark-mode preview; light mode uses their shared blue-600 token. - Verified 12px computed width and height in the task breadcrumb and Properties pane, the restored 20px main task icon, and 11px task IDs with baseline alignment in the breadcrumb and shared dashboard/Live runs cards. - Verified the chart bar and legend resolve to the spinner color; Recent Tasks counts align with Dashboard's count. Checked the fading ellipsis overlay with keyboard focus and its open-menu state, then dismissed the menu. ## Risks - Operators must open run details to read output that was previously embedded. - Run outcomes are available in the header tooltip and run details instead of a visible status chip. An in-progress task icon reflects task status, independently of an individual run's outcome. - Long names and task titles truncate within the compact layout. Full labels remain available through links and tooltips. - Font loading changes affect all UI text that uses the bundled Inter font. The files remain served from the existing public fonts directory. - No schema, API, adapter, or execution-policy changes. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with extended reasoning, repository tools, code execution, and browser verification. The session does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Scott Tong <scott@scotts-mbp-m5-max.local> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.912.0-canary.0 |
||
|
|
dbf5ea432d |
fix: protect starting runs during overlapping deployments (#13285)
## Thinking Path > - Paperclip controls agent work across service deployments. > - A run can provision a remote sandbox before a process or invocation event exists. > - Each container previously treated its own missing process handle as proof that the run was orphaned. > - Overlapping deployments could therefore fail a run owned by another container. > - This pull request records and renews a controller lease before provisioning. > - A recovery worker must revoke an expired owner before it finalizes the run. ## Linked Issues or Issue Description Merged PR #13272 records startup adapter identity and restores explicit user continuation. This PR adds controller ownership on top of current master. Refs #7997 and #10442 for related replica and ownership problems. Related #13138 addresses silence and detached local processes; this change does not infer death from silence. **What happened?** During an overlapping hosted service deployment, a new container reaped a legacy conversation run that another container was provisioning. The run had no PID or adapter invocation yet. **Expected behavior** A live controller keeps its run. After controller loss, one recovery worker takes cleanup authority and the old controller cannot dispatch further work. **Steps to reproduce** Claim a legacy run in controller A. Start controller B against the same database before A finishes provisioning. Run the startup reaper in B. **Paperclip version or commit** Observed on `663c44cb2b9c28336d38d0b4a6971f4f1964bce6` in a hosted Railway deployment with a Daytona environment. ## What Changed - Add nullable controller boot ID, lease deadline, and execution stage columns. Claim ownership in the queued-to-running update. - Renew ownership independently of run output. Abort and reject dispatch if renewal fails. - Serialize reaper revocation against renewal. Let unfinished recovery claims expire after a restart. - Restrict graceful shutdown to legacy runs owned by the current controller. - Hand ownership back to the existing native coordinator when runtime selection becomes native. - Add twelve database regressions and document the lease contract. Update the task-drain regression to require controller expiry before reaping. ## Verification - `pnpm exec vitest run server/src/services/legacy-controller-lease.test.ts server/src/__tests__/heartbeat-task-drain-admission-release.test.ts`: 14 passed after rebasing onto master (`f12b647ae`). - Queue-interruption regressions in `heartbeat-process-recovery.test.ts`: 2 passed after preserving the new cleanup promotion from #13275. - `pnpm --filter @paperclipai/server exec tsc --noEmit`: passed after rebuilding runner TypeScript outputs for the updated master. Broad local tests are omitted at the maintainer’s request; CI owns broad coverage. - Latest-head CI passed on `f255e8e4d5ab2b24b12638a02434e6aa8a2285c5`: [run 34658248569](https://github.com/paperclipai/paperclip/actions/runs/34658248569). All test shards, browser suites, typecheck, build, canary, and security checks passed. Greptile is 5/5 with no unresolved review threads. ## Risks - Additive, idempotent migration; historical rows retain the previous recovery behavior. - Database unavailability aborts new dispatch rather than permitting an unfenced controller to continue. - Lease expiry is permission to clean up, not evidence that remote inference stopped. Follow-up PRs add persistent cleanup and automatic continuation. - Mixed-version deployment still includes old binaries whose reapers do not understand controller leases. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f12b647ae8 |
fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can collect more messages while its agent works. > - Legacy runners must stop the active process before they can receive those messages. > - The old Interrupt action cancelled the run but could leave the queue idle and hidden. > - Codex could also classify a cancelled run as successful or start a fresh process after cancellation. > - This pull request joins cancellation, preserves the provider session, and dispatches the current queue after cleanup. > - The benefit is reliable interruption with the saved message order, edits, and deletions. ## Linked Issues or Issue Description **What happened?** Interrupt could strand a legacy message queue. The UI could hide pending messages after the run stopped. A Codex signal exit could race the cancellation write. A stale session warning could also trigger a fresh process after an interrupted resume. **Expected behavior** Interrupt stops the active turn and sends the remaining messages once, in their saved order. Deleted messages stay deleted. An interrupted Codex turn keeps its session and does not restart itself. **Steps to reproduce** 1. Assign a task to a legacy Codex agent that runs a long command. 2. Queue three messages. Edit one, discard another, and move the last message first. 3. Click Interrupt in the queue. 4. Repeat the interruption while the resumed session runs another command. Related work: Refs #13160, which moves native queue steering into the wake-queue module. This change fixes legacy interruption and keeps native steering unchanged. ## What Changed - Add a revision-checked, company-scoped endpoint for legacy queue interruption. - Promote only the requested queue after the provider stops and releases its lease. Retry its persisted interrupt intent from the scheduler after a promotion error or server restart. - Keep pending legacy queues visible after a run stops. Use server state for the interrupt result. - Serialize owned process cancellation before classifying the adapter result. Preserve late session and log metadata. Acknowledge cancellation only when an actual process or process group was owned; scheduler placeholders retain their normal release policy. - Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the session has started. - Add cancellation race, multi-actor queue order, durable retry, resume fallback, and stale request regression tests. Document the behavior. ## Verification - Real browser tests passed with legacy Codex CLI and ACP engines, using Codex 0.153.4 and gpt-5.6-sol. - All three automated ACP browser scenarios passed locally: immediate Interrupt delivery, no replay of an unfinished write, and pause requiring Resume. Updated the old test expectation that required a separate “go” after Interrupt. - Browser tests covered queued edits, deletion, reordering, deleting the final message, and repeated interruption. - Two consecutive CLI interrupts kept one provider session. Both stopped processes exited. The final message arrived once. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed. - All 346 post-review scheduling, recovery, queue-route, archived-company, worktree-suppression, and stale-queue regression tests passed. - All 318 process-recovery and durable-chat tests passed after the final cancellation guard. - Codex adapter, queue UI, issue-page, and OpenAPI contract tests passed. - `pnpm build` passed. - Full local suite coverage completed with `PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI shards: 618 general server suites, all 145 serialized server suites, and all workspace groups. Every failing suite passed a targeted rerun after the fixes, rebuilding the native test fixture, correcting macOS temporary-path setup, or retrying setup/timing failures. Existing skips remain. - The original monolithic run reported failures before the final fixes; its failed suites were rerun rather than rerunning all 618 suites again. The final process-recovery/durable-chat regression run passed all 318 tests. - All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: [run 34654820774, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2), including typecheck, build, all test shards, E2E, and canary. The signoff and Cursor sandbox tests each hit a timeout in the initial attempt; both suites passed locally, and both failed shards passed their single CI rerun. All three corrected ACP browser scenarios passed in CI. - Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no open review threads. ## Risks Cancellation order affects local adapters. The tests cover signal exits, graceful exits, adapter exceptions, termination errors, and cancellation write errors. Embedded adapters keep their cancellation controls. Ordinary run cancellation and task pause keep their distinct queue policies. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code execution. The exact serving model ID and context-window size are not exposed in this session. The live test runner used OpenAI gpt-5.6-sol. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
30c63af0e6 |
fix(cli): recover abandoned workspace build locks (#13288)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The test-drive command starts an isolated instance for local testing. > - Source startup builds shared packages before it starts the server. > - An interrupted build can leave an empty lock directory. > - Later starts wait without output and fail after 60 seconds. > - This pull request recovers abandoned locks and shows build progress. > - Local testing can start again without manual lock removal. ## Linked Issues or Issue Description **What happened?** `pnpm paperclipai test-drive` stopped at “Starting Paperclip server…” in a source checkout. A leftover plugin build lock caused a silent 60-second wait and then a timeout. **Expected behavior** Startup should recover an abandoned build lock. It should show when it waits for a live build. An interrupted or failed build should not leave partial output that the next start accepts as complete. **Steps to reproduce** 1. Leave an empty `node_modules/.cache/paperclip-plugin-build-deps.lock` directory after an interrupted build. 2. Make the shared or plugin SDK build output out of date. 3. Run `pnpm paperclipai test-drive --api-key placeholder --no-browser` with a fresh data directory. 4. Observe the silent wait at server startup. **Paperclip version or commit** Reproduced at `2083bf6f9`. **Deployment mode** Local source checkout with an isolated embedded PostgreSQL instance. Related work: #12894 added test-drive. #12898 restored its credential inputs. Neither change handles abandoned workspace build locks. No duplicate fix was found. ## What Changed - Publish a lock directory with an owner record in one rename. - Recover locks after their owner and compiler exit. Recover legacy empty locks after two minutes. - Keep the lock until the compiler stops on SIGINT or SIGTERM. - Print build and lock-wait progress. - Record source, dependency, compiler-config, and output content fingerprints only after a successful compile. Recover partial output even when modification times are unchanged. - Add 12 process-level regression tests and update the development guide. ## Verification - `node --test scripts/__tests__/ensure-plugin-build-deps.test.mjs`: 12 tests pass. - `pnpm exec vitest run --config cli/vitest.config.ts cli/src/__tests__/test-drive.test.ts`: 32 tests pass. - `pnpm --filter paperclipai typecheck`: passed. - `pnpm --filter paperclipai build`: passed. - Live smoke tests: fresh startup and startup with an abandoned lock both reach ready state. The API and UI respond. The command creates the company and CEO and enables worktree execution. Test instances stop cleanly. - Full repository `pnpm -r typecheck` and `pnpm build`: passed. - Full Vitest suite coverage completed using the repository-supported server, chat, workspace, and serialized shards. The initial local run needed the fresh-worktree fake native-provider binary built and focused reruns for port/socket races and load-related timeouts; all affected tests passed on rerun. Suites skipped by fail-fast exits were run separately and passed. The initial serial `pnpm test:run` was stopped in favor of these shards. - Greptile: 5/5 on commit `8b5a790c2af06a52b5dc76e5f52331966df990b8`, with all review threads resolved. - CI: 31 checks passed and two Storybook checks intentionally skipped. The initial workspace and browser jobs were interrupted by runner shutdowns; both passed on the second attempt. Build, typecheck, canary dry run, all general and serialized tests, all browser shards, security checks, and final verification summaries are green. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34654730783) ## Risks - This changes shared source-build locking for the CLI and plugin SDK commands. - Legacy locks have no owner identity. Recovery uses a two-minute age threshold for empty legacy directories. - Startup reads and hashes source and output files to verify the build cache. Identical direct builds reuse the cache. Changed or partial output requires a rebuild. - A reused process ID can delay recovery. Live owner or compiler processes keep their lock. - No database, API, or UI contract changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, code execution, and process-level testing. A more specific API model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a38ccf9972 |
fix: retry transient continuation admission locks (#13290)
## Thinking Path > - Paperclip manages AI agents and the tasks that they execute. > - The run scheduler checks continuation authority before it starts a provider. > - This check uses database locks to order execution against conversation closure. > - A short lock conflict could fail a valid user follow-up before the provider started. > - This pull request retries the admission transaction after the locks are released. > - Valid work can start after normal contention, while closure and cancellation still stop execution. ## Linked Issues or Issue Description Refs #13038. Related continuation work: #13270 and #13239. **What happened?** A user comment started a run through the automation queue. Its source records and admission marker were valid. A database lock conflict at dispatch caused `chat_control_recovery_proof_unresolved` and stopped automatic recovery. The provider received no work. **Expected behavior** Retry short database lock conflicts before failing admission. Read current ownership and conversation-close evidence on each attempt. Do not retry provider execution. **Steps to reproduce** 1. Queue a user follow-up through the automation transport. 2. Hold the task row lock in a separate transaction at the dispatch boundary. 3. Release the lock after 250 ms. 4. Before this fix, the run fails before provider dispatch. With this fix, the run passes admission once the lock is released. **Paperclip version or commit** Reproduced on base commit `1c4bcff2b`. Disabling the new retry reproduces the original error in the regression test. **Deployment mode** Self-hosted server with PostgreSQL. Regression tests use embedded PostgreSQL. ## What Changed - Retry rolled-back admission transactions after lock conflicts, with up to 50 waits of 100 ms. - Keep queue claims nonblocking. Keep provider dispatch outside the retried transaction. - Recheck current run state and committed close evidence after every conflict. - Explain persistent database contention in the exhausted admission error. - Add real database contention tests and bounded retry tests. Document the behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/services/chat-control-admission-retry.test.ts`: 273 passed. This includes task, wake, and run locks, the native runner, close/cancel races, unrelated failures, and retry exhaustion. - Regression proof: disabling retries makes the user-follow-up test fail with `chat_control_recovery_proof_unresolved`. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm test:run`: stopped after all equivalent CI server/workspace shards passed. The local run exposed a missing `fake-codex-app-server` fixture binary in the fresh worktree; after `pnpm --filter @paperclipai/paperclip-runner run build:rust`, the complete affected `native-session-resume.test.ts` suite passes (37 tests). - Greptile: 5/5, no findings, on commit `c974a496a`. - CI: all 31 checks passed on commit `c974a496a`, including all server/workspace test shards, browser tests, typecheck, runner verification, build, and the release dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34654770074). ## Risks - A contended dispatch can wait through 50 short delays, plus transaction time. - Persistent contention still fails closed after the retry budget. Invalid source evidence fails without retrying admission. - No schema, permission, provider retry budget, or queue-claim behavior changes. ## Model Used OpenAI Codex, GPT-6. The exact deployed model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, code editing, and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9031516a7e |
fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy conversation adapters can run in Daytona sandboxes. > - A server restart during provisioning can occur before the invocation event exists. > - Recovery then lacks the old adapter identity and leaves a hold that ordinary user retries cannot clear. > - A remote launch can also fail when its host relay looks for Node in the sandbox PATH. > - This pull request records the adapter at claim time and restores explicit user continuation after verified cleanup. > - Users can recover from the task or inbox while the failed run and uncertain action history remain intact. ## Linked Issues or Issue Description Refs #13237, #13239, #13254. Those changes cover recorded conversation runs, native user continuation, and explicit remote Stop. This change covers legacy failure before `adapter.invoke` and exact task/inbox Retry. Refs #9771 for overlapping generated-command quoting. This change also supplies the absolute host Node executable. Refs #13163 and #13264 for the separate native restart and retained-workspace work. **What happened?** A legacy Daytona run interrupted during provisioning became `process_lost` without an invocation event. Recovery preserved an execution hold, and Retry or a new task reply could not resume it. Cleanup could also run before the Daytona plugin was ready. On a macOS host, a subsequent ACP relay launch failed with `env: node: No such file or directory` because the remote launch environment did not contain the host Node path. **Expected behavior** An interrupted conversation can continue after its previous execution stops. Explicit Retry and new user replies should start a fresh turn with the task history. Cleanup failures must remain visible and recoverable. The host relay must use the host Node executable. **Steps to reproduce** 1. Use a legacy Claude adapter with a Daytona environment. 2. Interrupt the server after it acquires the sandbox lease and before it records `adapter.invoke`. 3. Restart and inspect the task hold. 4. Retry from the task or inbox, or send a new task reply. 5. Confirm the old sandbox has stopped and one new response arrives. **Paperclip version or commit** Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`. The branch is rebased on `51b0e01ea`, including #13261 and #13270. **Deployment mode** Built from source on macOS with a real Daytona sandbox and the legacy Claude ACP adapter. ## What Changed - Count new browser specs with the scheduler's median duration in the shard-balance check. This fixes a false policy failure after new specs arrive from both branches. The balance threshold is unchanged. - Persist server-owned adapter identity in the queued-to-running claim before provisioning starts. - Wait for provider plugin startup before restart cleanup. Keep failed cleanup leases as active ownership blockers. - Admit exact board retries and new user comments after verified termination. Retain the old run, task history, approvals, and unknown action outcomes. - Adopt repeated Retry requests. Permit one scoped cleanup attempt per explicit user Retry after the automatic limit, with an activity record. A later user Retry can recover after a transient provider failure; automatic attempts remain capped. - Resume replies deferred during cleanup, including historical legacy startup failures. - Launch the host ACP relay through the absolute host Node executable. - Add a task-level Retry button and return actionable blockers when retry admission is refused. - Add database regressions and three browser recovery journeys. Exclude installed third-party dependency skills from the shipped-skill audit. ## Verification - Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution retains the saved-message recovery, local stop receipts, and wait reasons from #13270 alongside exact legacy Retry support. - Real Daytona: interrupted the server after lease acquisition and before adapter invocation. Restart cleanup confirmed provider termination. Task Retry cleared a seeded historical hold and a real Claude agent returned `Recovery verified.` in the task. Removed the disposable sandbox and environment after testing. - All three browser recovery journeys passed again after the final rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh successor, completed the task, preserved the failed run, and retained the answer after reload. - All 29 e2e/server shard-partition tests passed. The balance check now uses the scheduler's median fallback for unmeasured specs, with the same balance threshold. - Server typecheck passed after rebuilding the generated runner dependencies. The combined recovery/route run passed 136 of 137 tests. Its remaining route test timed out during the first cold module import at its explicit 10-second limit; an isolated rerun reproduced that timeout and passed the other 51 route cases. The complete CI suite passed on this head. The same route file passed all 52 cases in CI, including the first cold import in 7.5 seconds. - Before the final rebase, recursive typecheck, full build, UI token gates, 132 targeted server tests, and the complete [CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085) passed. The subsequent CI failure was the shard-balance accounting mismatch fixed here. - Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable findings. The complete [current CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949) passed on attempt 2. All test, typecheck, build, and canary jobs passed on the first attempt. Docker setup timed out fetching BuildKit from Docker Hub; retrying that job and its dependent aggregate succeeded. ## Risks - Recovery admission changes executable authority. Company, task, agent, user, approvals, process ownership, and provider termination checks remain required. - Explicit continuation starts a fresh conversation with history. It does not certify unknown external action outcomes or rerun non-conversation adapters automatically. - Changing task status alone does not clear an execution hold. The task now offers an explicit Retry action. - Historical adapter claims and invocation events take precedence over current agent settings. Known process or webhook runs retain their hold. Pre-upgrade rows with no adapter evidence may receive only a new explicit user turn after termination proof; they do not become eligible for automatic replay. - No schema migration or sandbox-image change is required. This branch has not been deployed to production. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution, browser automation, and test execution. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
37d7dfb0e3 |
ci: allow dependency changes in cloud eval verification (#13286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployment requires source verification for the exact merged commit. > - Contributor PRs leave lockfile updates to a separate bot PR. > - Most release checks can refresh an outdated lockfile while installing dependencies. > - Two Runner checks still require a frozen lockfile and fail after dependency changes. > - This PR gives those checks the same install policy as the other release checks. > - A valid dependency change can become deployable without waiting for another merge. ## Linked Issues or Issue Description Refs #13257. The dependency change in #13256 exposed this gap. The separate lockfile update is #13279. Related #12115 addresses the bot PR check trigger; this PR fixes exact-source cloud verification itself. **What happened?** [Cloud readiness forcanary/v2026.911.0-canary.16 |
||
|
|
f2e2d38972 |
feat(ui): summarize completed runner activity groups (#13274)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows tool activity between agent updates. > - Finished groups currently keep the last raw command on screen. > - Failure totals add visual weight to normal retries. > - This pull request summarizes finished groups with short action phrases. > - Users can expand each group to inspect the original activity and output. > - The benefit is a quieter history that still explains the work. ## Linked Issues or Issue Description **What existing behavior does this improve?** Runner activity in live task chat and saved history. **Current behavior** A finished group shows its last command or tool target. A separate failure count remains visible after retries. **Proposed behavior** Show short phrases such as “Ran commands” or “Read files, ran commands” when a commentary group finishes. Remove failure counts in both active and finished groups. Keep original output in expandable history. **Reason and benefit** The activity summary describes the work without filling the conversation with shell arguments or treating each retry as an alarm. Unsuccessful reads and edits use “Checked files” and “Worked on files” to avoid claiming success. **Additional context** Refs: #13255. This builds on the merged rolling activity groups. Related open PR #13246 covers steering and status presentation; this change concerns completed summaries and failure counts. The design was reviewed in desktop and mobile Storybook previews. ## What Changed - Add a deterministic summary shared by builtin and provider-native activity. - Combine repeated categories and retries, preserve category order, and bound long summaries. - Summarize each inactive commentary group while the next group can still be running. - Keep expansion, individual output disclosures, and active rolling rows. - Use the production component in eight review stories, including animated desktop and mobile transitions. - Cover retries, unsuccessful edits, unknown tools, reasoning-only groups, saved history, and resuming work. ## Verification - Focused summary, group, and runner-turn tests pass: 62 tests. - Token gates pass. - The approved Storybook previews passed browser checks for desktop, mobile, keyboard expansion, live-to-completed transitions, and original retry output. - `pnpm -r typecheck`, `pnpm build`, and the production Storybook build pass. - The complete UI suite passes: 582 files and 6,005 tests. The full repository test run is in progress after repairing missing symlinks in the local PostgreSQL dependency. A native workspace suite that hit the setup failure now passes all six tests. - Greptile is 5/5 on the current commit, with no review threads. Security checks pass. All individual CI jobs pass except server shard 4/5, where one plugin-worker output timing assertion failed. The complete plugin-worker suite passes locally (83 tests). The workflow finished. One retry of the failed shard and aggregate check was dispatched and is queued. Squash auto-merge is enabled and remains gated on required checks. - Rechecked the production stories in the browser: mobile history stays on one line, and expansion survives completion while the next group remains active. - Review in Storybook under Tasks → Completed activity preview. Use Next to finish one group while the next group remains active. Open a completed summary to inspect its history. ## Risks - Summaries classify observed activity; “Ran commands” does not mean exit code zero. - Unknown tools use a generic description. Full names and outputs remain in history. - No API, database, or runner protocol contracts change. > Reviewed ROADMAP.md. This is a focused improvement to the existing task-chat presentation. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c4bcff2b1 |
fix(test): await issue lock before retry race assertions (#13273)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Retry decisions must respect changes to an issue owner. > - Database tests verify this with two concurrent transactions. > - The test started the competing operation before its fixture held the lock. > - That race can fail a correct source verification run and block deployment. > - This PR waits for lock acquisition before starting the competing operation. ## Linked Issues or Issue Description Refs #13257 for the deployment verification work that exposed this test race. No duplicate fix was found. **What happened?** [Cloud verification job 103424158921](https://github.com/paperclipai/paperclip/actions/runs/34648268409/job/103424158921) failed because the test did not observe a concurrent issue-row lock waiter. The fixture and retry operation both started without an ordering guarantee. **Expected behavior** The fixture must hold the issue lock before the competing retry operation starts. The test must still prove that the retry waits for the lock and observes the reassignment. **Steps to reproduce** 1. Add a temporary 50 ms delay before the fixture acquires the issue lock. 2. Run the promoteOrCancelDueRetry issue-lock test. 3. The original fixture fails with the same missing-waiter error as CI. 4. The synchronized fixture passes with that delay. The delay is not part of this PR. ## What Changed - Separate fixture readiness from its transaction completion promise. - Wait for readiness at both callers before starting the retry decision. - Propagate transaction failure during setup through Promise.race. ## Verification - The original fixture fails under the temporary delayed-lock probe. The fixed fixture passes the same probe. - All 31 tests in server/src/modules/run-dispatch/adapters/postgres.test.ts pass after removing the probe. - The real concurrent waiter, lock-order, and reassignment assertions remain intact. No timeout was increased. - Full local typecheck and build pass (221s and 50s). The full local test command stopped in its server phase after 10,610 passes, 65 skips, and 14 failures: 13 existing macOS skill-cache rename/permission failures and one unchanged Telegram test assertion. The Telegram test passes in a focused rerun. Later local phases did not run after that failure. Current-head Greptile is 5/5 with no findings. All 32 current-head checks pass, including full Linux typecheck, build, native verification, server suites, browser suites, and canary packaging. The unchanged GitHub browser mock assertion passed its single failed-shard retry. - git diff --check passes. ## Risks - Only test synchronization changes. Production database behavior is unchanged. - Awaiting the transaction itself during setup would deadlock the test. Returning the completion promise inside an object avoids that problem. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — all 31 database adapter tests pass; full local verification limits are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51b0e01ead |
fix: resume saved user messages after execution recovery (#13270)
Preserve verified native process-stop evidence and retry saved user messages through normal continuation admission after recovery cleanup. Show the current wait reason and serialize delivery so a saved message starts one fresh turn. Validated with 410 focused tests, typecheck, build, token gates, all PR CI checks, and Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
250deab910 |
fix(runner): keep healthy native sessions alive (#13261)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions keep a provider process and its work alive across control-plane operations. > - A hidden 15-minute turn deadline stopped work even when the agent timeout was zero. > - A one-hour runner lifetime and fixed connection lease added two more limits. > - Recovery also rejected goal commands because it reconstructed their startup summary with the wrong protocol version. > - This pull request removes implicit duration limits and renews authenticated leases in the harness. > - Healthy sessions can continue without model action or a user interface change. ## Linked Issues or Issue Description Refs #13092 and #12845. Related: #13163 covers sandbox recovery after app restarts; this change covers session duration and lease renewal. **What happened?** A native Codex session stopped after 15 minutes while a tool was still running. The agent had `timeoutSec: 0`. Recovery then rejected a `session.goal.get` startup command with `invalid provider startup ownership fence`. **Expected behavior** An unlimited session keeps working while its provider and authenticated controller remain healthy. Lease maintenance is transparent. Explicit timeouts, cancellation and revoked authority still take effect. **Steps to reproduce** Start a native session with `timeoutSec: 0` and run a tool beyond 15 minutes. Before this fix, the runtime cancels the turn. A recovery startup that uses a goal command also exposes the protocol-version mismatch. ## What Changed - Honor the agent turn timeout. Zero disables the timer. Long explicit durations use timer chunks to avoid Node timer overflow. - Default native runner lifetime to unlimited. Keep bounded startup, reconnect and control-operation deadlines. - Renew leases over the authenticated connection. Persist renewal before the reply. Validate identity, epoch and expiry. Handle duplicate requests and a lost reply on reconnect. - Freeze renewal during warm ownership transitions and terminal handling. - Validate persisted goal startup commands with protocol v2. - Add duration, renewal, ownership, recovery and real-process regression tests. Update runner protocol and recovery docs. - Add no UI components or controls. Renewal requires no model output or user action. ## Verification Current head: `348e369c35c5da8bb8be378f4b35dcf6f40882e7`. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34649767113). - All 32 checks pass on this head. The two Storybook checks are skipped as expected. CI includes full build, typecheck, runner verification, browser suites, server tests and the canary package dry run. - Greptile reports 5/5 on this head. All review threads are resolved, and the security scan passes. - Passed `pnpm -r typecheck` and `pnpm build` locally. - Passed 219 native-runtime and controller tests, including fake-clock tests for three weeks of renewal and 30-day explicit timeouts. Six denial tests confirm that renewal cannot extend expired, revoked or mismatched authority. - Passed 337 executor, cancellation and restart-recovery tests, plus 278 Rust runner-core library tests. - Passed a real runner with a silent fake Codex provider across its original lease expiry. Runner PID, provider PID, thread and active turn stayed unchanged. Warm-attach recovery tests also pass. - Passed all 83 plugin-worker tests and 159 of 161 workspace-runtime tests locally. The two remaining assertions passed with a canonical macOS temporary directory, as did the changed runtime fixture. The full affected server shard passes in CI. - An unchanged GitHub callback-ordering test failed once in CI, passed locally in isolation, and passed its one test-shard retry. The final CI summary is successful. - The full local `pnpm test:run` sweep was interrupted after dependency setup failures and load-related timeouts. Identified failing suites passed in isolated reruns after the dependency repair. The complete test matrix passed remotely in CI. ## Risks - Deploy the controller and runner together to enable renewal. Older peers keep their existing bounded lease behavior. - Unlimited runtime permits long resource use until completion, explicit cancellation, configured timeout or loss of valid authority. - Lease renewal changes authenticated protocol handling. Regression tests cover stale, revoked and mismatched authority, lost replies and warm handoff behavior. - Simulated multi-week tests and a real lease-boundary test do not constitute a weeks-long production soak. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution and TypeScript/Rust test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a12bbd1824 |
fix: stop completion reviews caused by policy upgrades (#13266)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runtime records completion assessments and task decisions. > - An application update can change the assessment policy version. > - The previous code treated that version change as a reason for human review. > - An upgrade alone does not give the user a new decision to make. > - This change keeps existing decisions and withdraws obsolete upgrade review cards. ## Linked Issues or Issue Description **What happened?** A rules version change moved unfinished tasks into review and created a card that said, "Review the superseding native policy assessment." The task did not need new work or a human decision. **Expected behavior** New runs use the current rules. An upgrade leaves existing task decisions alone. The saved policy version remains available in the audit history. **Steps to reproduce** 1. Save a native run assessment and a task decision. 2. Change the native status policy version. 3. Run finalization reconciliation without new task evidence. 4. The previous code created a review card. The corrected code keeps the saved assessment and decision. Related context: #13038 changed the native policy version. Searches found no duplicate fix for upgrade-only completion reviews. ## What Changed - Remove policy-version mismatch as a reconciliation trigger. - Withdraw pending cards only when their source decision, effect ledger, creator, key, and prompt match the old upgrade-only review. - Restore the previous status only while that decision and status version remain current and no other review gate is pending. - Preserve answered cards, real review requests, later task changes, and historical assessments and decisions. - Record cleanup activity and retire any corresponding chat review actions. - Isolate cleanup failures so one old card cannot block other cleanup or normal finalization. - Update the status conformance fixture, regression tests, and architecture documentation. ## Verification - Passed: `pnpm exec vitest run server/src/__tests__/native-status-arbiter-corpus.test.ts` (23 tests, including the 53-fixture status corpus). - Passed: `pnpm -r typecheck`. - Passed: `git diff --check`. - Passed: `pnpm build`. - Passed: fresh Greptile review at 5/5 on `d24e5ecaf`, with no open review threads. - Passed: the 23 focused tests in the isolated full-suite environment. - Local `pnpm test:run`: the server group finished with 10,638 passed and one missing-fixture failure. Built the required `fake-codex-app-server` fixture and reran the entire affected native session-resume suite: 37/37 passed. The initial full command exited on that server-group failure, so remaining groups are covered by CI. - Passed: the workspace-runtime-exposure suite (25 tests, 3 platform skips). - Passed: all CI gates, including every test shard, browser tests, typecheck, and build. The server shard passed on one retry after an unrelated host-port conflict in the unchanged workspace-exposure tests. - Cleanup tests cover repeat runs, real reviews, answered cards, later statuses, newer decision identity, status changes back to review, other pending requests, run scope, cleanup failures, retry, and live publication failures. ## Risks - Cleanup changes stored task state. It checks the exact obsolete decision and status version under database locks before restoring status. - Pending approvals, interactions, and execution stages prevent restoration out of review. - The cleanup handles at most 100 matching cards per reconciliation pass. It does not rewrite old decisions or accept an agent's completion claim. - New evidence and explicit task changes still use the existing reconciliation paths. This change does not reevaluate old work merely because Paperclip was updated. ## Model Used OpenAI Codex, GPT-6. The exact deployment identifier and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.15 |
||
|
|
19c76bfc3f |
fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - CI installs dependencies before it verifies and builds cloud artifacts. > - Install jobs share a pnpm package-store cache with lockfile refresh. > - Lockfile refresh resolves versions without downloading packages. > - That job saved an empty cache before full install jobs could save theirs. > - This PR prevents lockfile refresh from publishing that empty entry. > - Full install jobs can then populate the cache and reuse dependencies. ## Linked Issues or Issue Description Refs #13259 for the related cloud verification cache work. No duplicate empty-cache fix was found. **What happened?** Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm cache at 18:58:08 UTC on September 10. Full install jobs still restore that empty entry. The cache API reports 216 bytes for master and about 703 MB for populated entries with the same key and cache version in PR scopes. **Expected behavior** A job that installs dependencies should populate the shared package-store cache. **Steps to reproduce** 1. Run lockfile refresh with a new lockfile cache key. 2. Its resolution-only command leaves the package store empty. 3. The Node action saves the empty archive before a full install finishes. 4. Later jobs report a cache hit but download packages again. **Paperclip version or commit** Observed on master |
||
|
|
778ac33308 |
feat: accept a base64-encoded Cloud UI snippet (#13245)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip Cloud instances can inject a trusted operator HTML snippet before `</body>` in served UI pages (#13168) > - Operators deliver env vars to managed instances through hosting-provider APIs > - Web application firewalls in front of those APIs reject request bodies that contain raw `<script` markup > - The snippet value always contains a script tag, so the firewall blocks its delivery every time > - This pull request adds `PAPERCLIP_CLOUD_UI_SNIPPET_B64`, which carries the same snippet as standard base64 > - The benefit is that operators can ship the snippet through WAF-fronted delivery pipelines without firewall exceptions ## Linked Issues or Issue Description Refs #13168 (introduced the Cloud UI snippet). **Subsystem affected** Server UI shell serving: `server/src/cloud-ui-snippet.ts`, used by `static-index-html.ts` and `app.ts`. **Current behavior** `PAPERCLIP_CLOUD_UI_SNIPPET` is the only way to configure the Cloud UI snippet. The value is raw HTML. Delivery pipelines that write env vars through provider APIs can fail to deliver it: web application firewalls classify a request body that contains `<script` as an injection attempt and block it. Verified against Railway's GraphQL API, which sits behind Cloudflare: a variable value with a bare `<script></script>` returns HTTP 403 before authentication, and JSON unicode escapes (`<script`) do not bypass the block. The snippet is exactly the kind of value that always contains a script tag, so such pipelines cannot deliver it at all. **Proposed behavior** A new optional `PAPERCLIP_CLOUD_UI_SNIPPET_B64` carries the same snippet as standard base64 of the UTF-8 HTML. The server decodes it and injects the result through the existing path. The plain variable wins when both are set. Whitespace and line wrapping in the value are tolerated. A value that is not canonical base64, or that decodes to blank, is ignored instead of injected as garbage. **Reason and benefit** Operators can deliver the snippet through WAF-fronted APIs without requesting firewall exceptions. Base64 contains no markup, so the firewall has nothing to match. Existing deployments see no behavior change. **Breaking changes** None. The new variable is optional. The existing variable is unchanged and takes precedence. ## What Changed - `server/src/cloud-ui-snippet.ts`: resolve the snippet from `PAPERCLIP_CLOUD_UI_SNIPPET` first, then from base64-decoded `PAPERCLIP_CLOUD_UI_SNIPPET_B64`; ignore non-canonical or blank-decoding values. - `server/src/__tests__/cloud-ui-snippet.test.ts`: cover decode-and-inject, plain-wins precedence, whitespace tolerance, invalid/blank values, and the self-hosted no-op. - `doc/cloud-ui-snippet.md`: document the variant, how to produce the value, and the ignore rules. ## Verification - `pnpm vitest run server/src/__tests__/cloud-ui-snippet.test.ts server/src/__tests__/static-index-html.test.ts` — 13 tests pass. - `tsc --noEmit` in `server/` passes. - Manual check: on a Cloud-managed instance, set `PAPERCLIP_CLOUD_UI_SNIPPET_B64="$(base64 < snippet.html)"`, restart, open `/`, and confirm the decoded snippet appears before `</body>`. On a self-hosted instance, confirm no snippet is injected. ## Risks Low risk. The injection path and the Cloud-managed gate are unchanged. A malformed base64 value is ignored, so the failure mode is a missing widget, not corrupted HTML. The decoded value is trusted operator configuration, the same trust model as the plain variable. ## Model Used - Claude Fable 5 (Anthropic, `claude-fable-5`) via Claude Code CLI, agentic workflow with tool use (code edits, test runs, and live API verification of the WAF behavior). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc68312327 |
ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.
## Linked Issues or Issue Description
Refs #13243.
**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.
**Current behavior**
For merge
|
||
|
|
ad4f0b5867 |
Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtime settings can bind organization secrets to an adapter environment > - Paperclip redacts plain environment values when it returns a saved agent to the UI > - A saved-agent test sent the redacted `CODEX_HOME` value back to the server > - Codex ACP also received the API key without an ACP API-key authentication request > - This pull request restores saved environment values for tests and selects API-key authentication for Codex ACP runs > - The benefit is that Codex agents can test and run with an organization-scoped OpenAI API key ## Linked Issues or Issue Description **What happened?** Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`. Secret normalization rejected that placeholder. Remote Codex ACP runs received `OPENAI_API_KEY`, but session creation stopped with `Authentication required`. **Expected behavior** Paperclip must use the saved `CODEX_HOME` value when it tests an existing agent. Codex ACP must select API-key authentication when `OPENAI_API_KEY` is available. **Steps to reproduce** 1. Create an organization-scoped secret named `OPENAI_API_KEY`. 2. Give a Codex agent access to the secret. 3. Save the agent runtime settings. 4. Test the saved agent again. 5. Run the agent in a remote sandbox through ACP. **Paperclip version or commit** Reproduced on master before commit `68c17709d7c051a804a416263e2e08920f1dfcb1`. **Deployment mode** Self-hosted server with a remote sandbox environment. **Installation method** Built from source. **Agent adapter(s) involved** Codex. ## What Changed - Send the saved agent ID with adapter environment tests. - Restore redacted plain environment values from the saved agent before test-time secret resolution. - Select the Codex ACP `api-key` authentication method when `OPENAI_API_KEY` is present. - Add focused regression coverage for saved-agent tests and remote ACP launch configuration. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/test-agent-setup.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `git diff --check` ## Risks - Low risk. The test route reads saved configuration only when the request supplies a compatible agent ID and the caller can update that agent. - The Codex ACP change applies only when `OPENAI_API_KEY` exists and no explicit `DEFAULT_AUTH_REQUEST` exists. - There are no schema migrations or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. The context-window size is not exposed in this runtime. The model used reasoning, repository search, file editing, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3bafac12f7 |
refactor: remove automatic productivity reviews (#13263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its recovery loop keeps assigned work moving after execution failures. > - Productivity review used run counts, comment counts, and elapsed time to create management tasks. > - Infrastructure failures could satisfy those rules and create more tasks without evidence that the source work needed management review. > - This pull request removes that detector and its continuation holds. > - Bounded recovery, budgets, explicit blockers, and normal review stages remain in place. > - Existing task records stay readable and unchanged. ## Linked Issues or Issue Description Refs #5897. That request describes unwanted automatic productivity reviews and asks to preserve existing tasks. This change retires the feature instead of adding another configuration switch. Related prior approaches: Refs #9191, Refs #12489. Those changes excluded infrastructure failures or bounded review creation. This removal replaces the detector rather than tuning its thresholds. ## What Changed - Delete the scheduled detector, automatic task creation, evidence refresh, and productivity continuation holds. - Remove computed productivity fields, special attention items, badges, and Storybook fixtures. - Retain historical origin values, decision compatibility, and recovery recursion exclusions. Add no migration and change no existing task data. - Update the execution contract. Replace feature tests with regressions for legacy task reads, ordinary attention, and bounded continuation in the presence of an old review. ## Verification - Targeted attention, issue-route, startup, and UI tests: 4 files and 101 tests passed. - Updated issue-route and UI tests: 2 files and 61 tests passed. - Bounded continuation regression: 2 cases passed, including a legacy review plus pre-dispatch cancellation churn. - `pnpm check:token-gates`: all four gates passed. - `git diff --check`: passed. - `pnpm build-storybook`: passed. - Greptile: 5/5 on `a5a612eea`, with no actionable findings. - Scheduler and historical recovery regressions: 2 files and 28 tests passed. - Repository `pnpm -r typecheck` and `pnpm build`: passed. - The complete `pnpm test:run` suite passed across the CI server, serialized-server, and workspace shards on `a5a612eea`. Stopped the duplicate local monolithic run after the full CI suite passed; no completed local full-suite result is claimed. The targeted local suites above passed. - CI serialized shard 5 initially hit a 10-second timeout in the first interaction-route test. The complete file passed locally (78 tests), then the single CI rerun passed. - All CI gates are green, including the build and end-to-end suites. - A local merge check against current `master` (`ce09ea40b`) completed without conflicts. ## Risks - API responses no longer include the computed `productivityReview` field. Consumers must stop using it. - The scheduler no longer creates management work from elapsed time, run counts, or missing comments. This is the intended behavior change. - Existing review tasks and explicit dependencies remain in place. Historical origins still prevent recursive recovery treatment. No task cleanup or data migration occurs. - The native review handoff repair is separate from this removal. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
canary/v2026.911.0-canary.14
|
||
|
|
a23ae894a5 |
ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud images become deployable only after source verification passes. > - The typecheck job builds the native Runner binary through the server package. > - Fresh runners repeatedly compile Rust dependencies for that binary. > - This PR caches those dependencies for exact-source master verification. > - Workspace code and every typecheck still rebuild or run as before. ## Linked Issues or Issue Description Refs #13243 and #13257. **What existing behavior does this improve?** The typecheck portion of post-merge cloud source verification. **Current behavior** The typecheck job has no Rust dependency cache. An observed release build in this job took 4m 13s, including dependency compilation. **Proposed behavior** Restore dependency build outputs for canonical master pushes with the exact source SHA. Use a separate cache key from the Runner verification job, which builds other profiles. **Reason and benefit** A warm cache should remove roughly 2–3 minutes of dependency compilation from this job. Overall deployment gains depend on the remaining critical path. The first cache population still compiles from scratch. **Breaking changes** None. All checks remain enabled. Non-master callers compile without restoring or saving this cache. ## What Changed - Select the pinned Rust toolchain before the typecheck cache lookup. - Reuse the existing pinned Rust cache action with a typecheck-specific key. - Exclude workspace crates and installed cargo executables. - Test restore/save trust boundaries and document cache behavior. ## Verification - All workflow script tests pass locally, including nine new cache trust/contract cases. - actionlint passes for release-verify.yml. - Full local typecheck and build pass on the same source base (167s and 206s); `pnpm test:run` is still running and is recorded with #13257. This PR changes only the workflow, cache guard tests, and documentation. - All 32 current-head checks are successful or intentionally skipped, including the complete Linux test matrix, build, and Greptile 5/5 with no unresolved findings. Verify cache population and subsequent restore on actual master runs. ## Risks - The first run and any toolchain/dependency invalidation compile from scratch. - Cache restore/save overhead reduces the benefit for small dependency graphs. - Disable the cache step to roll back; the existing uncached build remains valid. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. Exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — focused change tests pass; full-suite local permission failures are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a27dd722a2 |
fix(ui): restore task page archive shortcut (#13253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail page gives operators keyboard shortcuts for fast task work. > - The `y` shortcut must archive the open task from the inbox and return the operator to the inbox. > - A streamlined UI gate limited this shortcut to tasks opened from a selected inbox row. > - Direct task pages no longer ran the shortcut. > - This pull request removes that gate and adds regression coverage. > - The benefit is consistent keyboard operation from every visible task page. ## Linked Issues or Issue Description **What happened?** The `y` keyboard shortcut did nothing on a task page opened outside the inbox. **Expected behavior** The `y` shortcut archives the open task from the inbox and returns the operator to `/inbox`. **Steps to reproduce** 1. Enable keyboard shortcuts. 2. Open a visible task from the Tasks page. 3. Press `y` while focus is outside an editor. 4. Observe that the task does not archive. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local development. ## What Changed - Enable the `y` archive shortcut on every visible task detail page when keyboard shortcuts are enabled. - Keep the existing input, dialog, modifier, hidden-task, and pending-request guards. - Verify that `y` archives the task and navigates to `/inbox` for a task opened from the Tasks page. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/keyboardShortcuts.test.ts ui/src/lib/issueDetailBreadcrumb.test.ts` (139 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build` - The full workspace typecheck and build reached the Rust runner. This environment does not include `cargo`, so the Rust steps did not run. - The full test suite entered unrelated serialized external-chat integration tests. It was stopped after the affected UI suite passed. ## Risks - Low risk. The change restores the shortcut behavior that existed before the streamlined UI gate. - The shortcut still ignores text inputs, open dialogs, modifier keys, and hidden tasks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex. The Paperclip runtime did not expose the exact model ID or context window. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce09ea40b0 |
feat(ui): show one rolling runner activity per commentary group (#13255)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows commentary and runner activity as work proceeds. > - The expanded activity list uses multiple text lines and an indented rail. > - The folded view can repeat reasoning and the current tool in separate places. > - This pull request shows one rolling activity between commentary messages. > - Users can open a group to follow its full history and inspect each item. > - The benefit is a compact feed with aligned icons and stable row height. ## Linked Issues or Issue Description **What existing behavior does this improve?** The runner activity groups in task chat, during live work and in saved history. **Current behavior** Activity is spread across a group summary, a reasoning ticker, and a current-tool row. Expanded rows have different alignment and can span several lines. **Proposed behavior** Keep one latest activity per commentary group. Roll new activities upward. Keep each expanded history row on one line. Use neutral failure text without an X icon. Keep the full details available by opening an individual row. **Additional context** Refs: #13246. That open PR covers related folded activity, steering timestamps, and status alignment. This change implements the separately reviewed activity-group design, including animated replacement, single-line expanded history, and neutral failures. It does not change steering timestamps or status pills. ## What Changed - Add a shared runner activity group with persistent expansion and reduced-motion support. - Use the group in the production runner turn and remove duplicate narration and current-tool rows. - Preserve commentary, plan cards, request receipts, and final-response placement. - Keep tool output, reasoning, research sources, file previews, and failure details accessible. - Render the production turn in desktop and mobile animated Storybook fixtures. - Add tests for animation identity, failure visibility, expansion persistence, and timeline order. - Validate resource link schemes and omit disclosure controls for items with no details. ## Verification - Focused activity, runner turn, protocol detail, and task-thread tests pass: 161 tests. - `pnpm check:token-gates` passes. - Browser checks confirm a 32-pixel compact row, centered icons, and one-line expanded rows. - `pnpm -r typecheck` and `pnpm build` pass. The final UI build also passes. - Two real native Codex runs completed in an isolated local instance. Each run created the expected file, performed separate delayed tool calls, reported a deliberate missing-file failure, recovered, and finished the task. - Browser verification confirmed that the open three-row activity group remained expanded after the second run moved into saved history. Saved error output remained accessible without red styling or an X. - The complete UI suite passes: 581 files and 5,979 tests. - `pnpm --filter @paperclipai/ui build-storybook` passes for the final production fixtures. - Greptile rates the latest commit 5/5, with zero new comments and no unresolved review threads. The security scan passes. - All repository test shards pass in CI. The duplicate local `pnpm test:run` was stopped after CI completed the full test coverage. The CI build failed twice because the worker ran out of disk space compiling the Rust runner tests (`No space left on device`, error 28). The build and its aggregate verify gate remain blocked by CI capacity. GitHub squash auto-merge is enabled and will wait for required checks to pass. ## Risks - This changes presentation and disclosure behavior. Users open a group to see earlier activity. - Animation uses stable item IDs. A provider that replaces IDs can trigger another transition. - No API, database, or runner protocol contracts change. > Reviewed `ROADMAP.md`. This improves an existing task-chat surface. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2904a3a6cc |
fix(ui): hide retry countdown after execution starts (#13258)
## Thinking Path
> - Paperclip helps operators manage AI agents and their tasks.
> - Task pages show countdowns for deferred checks and automatic
retries.
> - A retry keeps its scheduled start time after it enters the queue or
starts running.
> - The countdown treated that historical time as a pending deadline and
showed an overdue warning beside active work.
> - This pull request limits retry countdowns to retries that are still
scheduled and hides waiting surfaces on terminal tasks.
> - Operators now see a warning only when the displayed retry is still
waiting to start.
## Linked Issues or Issue Description
Refs #9783, which added the monitor surfaces. Searched related PRs and
issues; no duplicate fix was found.
**What happened?**
After a service restart resumed a task through an automatic retry, the
task showed an overdue retry banner while the agent was running. The
banner also remained when the task became done.
**Expected behavior**
A queued or running retry must not show a countdown against its past
scheduled start time. Done and cancelled tasks must not show waiting
banners.
**Steps to reproduce**
1. Open a task with an automatic retry scheduled for a known time.
2. Let the retry enter the queue and start running.
3. Wait until its scheduled time is more than one minute in the past.
4. Observe the overdue banner and Check now button while the agent is
working.
**Paperclip version or commit**
Reproduced on source commit
|
||
|
|
2fc5e3e43b |
test: isolate native controller takeover fixture (#13212)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native runtime recovery protects controller takeover with process identity checks > - The live-controller test mocked process liveness but left process start time to host state > - This pull request supplies the matching process start time in the fixture > - The benefit is deterministic coverage of the live-controller fence ## Linked Issues or Issue Description Closes #13172 ## What Changed - Complete the live-controller takeover fixture with the recorded process start time. ## Verification - Contributor reports focused native restart-recovery tests, related native tests, integration tests, and typecheck passed locally. Maintainer reran the 20 native restart-recovery tests successfully and verified current-head CI, Greptile 5/5, and zero unresolved review threads. ## Risks Low risk. This changes test setup only. Production controller fencing is unchanged. ## Model Used Original patch by @lorenzozanee; the original author did not record model details. Maintainer review and verification used OpenAI GPT-6 through Codex with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change is needed for this test fixture) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
847d00bdc3 |
fix(ui): use full mobile width for task tree (#13250)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The task detail page gives operators a task tree in a mobile side panel. > - The mobile sheet did not set an explicit full width. > - A maximum width rule could make the task tree narrower than the phone viewport. > - This pull request makes the mobile task tree sheet use the full viewport width. > - The benefit is a consistent and readable task tree on phones. ## Linked Issues or Issue Description **What happened?** The task tree sheet could use less than the full viewport width on a mobile task page. **Expected behavior** The task tree sheet must use the full available phone width. **Steps to reproduce** 1. Open a task on a phone viewport. 2. Open the task side panel. 3. Select the Tasks tab. 4. Compare the sheet width with the viewport width. **Paperclip version or commit** `7b829efdf` **Deployment mode** Built from source. ## What Changed - Set the mobile task side panel to `w-full` and `max-w-none`. - Add regression assertions for both width classes. ## Verification - `vitest run src/pages/IssueDetail.test.tsx --config vitest.config.ts` passes 103 tests. - `node scripts/check-token-gates.mjs` passes all four token gates. - `git diff --check` passes. ## Risks - Low risk. The change only affects the mobile task side panel width. - The desktop panel and other sheet variants do not change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, agentic coding with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.13 |
||
|
|
d0b7ba4194 |
ci: route approved master cloud builds to AWS (#13243)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images built from master commits. > - Cloud image builds share GitHub-hosted capacity with other workflows. > - The organization already operates AWS runners through RunsOn Fleet. > - This pull request allows approved master builds to use a dedicated cloud Fleet. > - The benefit is separate build capacity with a quick operator rollback. ## Linked Issues or Issue Description Refs #13189, #13192. **What existing behavior does this improve?** Placement of the Docker cloud build after a master merge. **Current behavior** Every Docker cloud build uses a GitHub-hosted runner. Busy periods delay the job. **Proposed behavior** An operator variable enables the approved cloud Fleet for canonical master pushes and manual master builds. Other events, refs, and repositories use GitHub-hosted runners. ## What Changed - Add a guarded AWS runner selector to the Docker cloud job. - Keep the existing image cache, verification, and publication steps. - Test the selector against master, branch, tag, PR, fork, and disabled contexts. - Document provisioning requirements, placement checks, and rollback. ## Verification - 29 focused Node tests pass for routing, readiness, and disk handling. - The full workflow-script Node suite passes. - `pnpm -r typecheck` passes locally. - The pinned PR routing regression suite passes. The first live PR run assigned 21 jobs to the approved AWS PR group. AWS then reclaimed 16 Spot instances. The failed run is being repeated on GitHub-hosted runners while the Fleet moves to On-Demand. - Actionlint passes with existing shellcheck findings excluded (SC2012, SC2016, SC2129). - `git diff --check` passes. - Greptile reports 5/5 on commit `764d505a41dd2023751c3f361906fa9ea35bf0c6`, with no review threads. - All 30 current-head CI checks pass, including typecheck, build, all server/workspace test shards, Runner verification, and browser tests. Two Storybook checks are intentionally skipped for this change. Run: https://github.com/paperclipai/paperclip/actions/runs/34630550799 - The broader local test/build sequence is still running. This Mac has reported failures in unchanged application suites; their complete Linux CI shards pass. Local targeted workflow tests and typecheck pass. - Both On-Demand Fleets are deployed and healthy. Live master cloud-build verification follows the merge. ## Risks - Missing Fleet capacity or runner-group authorization can leave an AWS job queued. Disable `AWS_CLOUD_BUILDS_ENABLED` and rerun the workflow to use GitHub-hosted capacity. - The runner group must restrict access to this repository and the master version of `docker-cloud.yml`. - Docker needs more disk space than the PR Fleet. Provision 120 GiB disks and retain the free-space check. - This changes image build placement only. Source verification and migrator publication remain separate prerequisites. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.12 |
||
|
|
063ba59ae3 |
Fix workspace-ready notice rendering (#12716)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task thread shows user, agent, and control-plane messages. > - Workspace-ready comments keep agent attribution for audit and authorization. > - These comments also have an explicit system-notice presentation. > - The client used attribution before the presentation contract. > - As a result, workspace-ready comments appeared as agent bubbles. > - This pull request gives the system-notice presentation priority. > - The benefit is a compact workspace notice with expandable details. ## Linked Issues or Issue Description **What happened?** The task thread showed a workspace-ready control-plane comment as a normal agent message. The comment kept agent attribution for audit and authorization, so the client did not use its system-notice presentation. **Expected behavior** The task thread must show a comment with the `system_notice` presentation as a compact system notice. The user must be able to expand the notice to inspect its details. **Steps to reproduce** 1. Open a task with a workspace that becomes ready. 2. Wait for the control plane to add the workspace-ready comment. 3. Observe that the task thread shows the comment as an agent bubble instead of a system notice. **Paperclip version or commit** `master` at `8eaa5caa0`. **Deployment mode** Local dev (`pnpm dev`). **Agent adapter(s) involved** Not adapter-specific. This is a core UI bug. ## What Changed - Give the `system_notice` presentation priority when the task-chat adapter selects the message kind. - Pass the notice presentation data to the task-chat model. - Improve the compact and expanded workspace notice details. - Add component tests, adapter tests, and Storybook cases for workspace notices. ## Verification - `pnpm exec vitest run ui/src/components/task-chat/task-chat-adapter.test.ts ui/src/components/task-chat/TaskChatSystemNotice.test.tsx` passes 15 tests. - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` passes 5,670 tests and reports four pre-existing workspace-runtime failures. The same four failures reproduce when the two unrelated server test files run alone. They do not use the changed task-chat files. ## Risks - Risk is low. Normal agent messages still use the agent presentation. - A comment with an explicit system-notice presentation now uses the compact notice renderer. - Regression tests cover both paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 (`gpt-5`) through Codex. The agent used reasoning, repository tools, and code execution. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b829efdf6 |
feat: show tasks created from a task by project (#13241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can cause an agent to create more tasks. > - Those tasks can belong to other projects or have another parent. > - The subtask view does not show all work created from the current task. > - This pull request adds a Tasks tab with separate subtask and creation groups. > - Operators can follow created work without changing its parent or project. ## Linked Issues or Issue Description **Problem or motivation** Operators need to see all work that an agent creates while running for a task. Parentage alone does not describe this relationship. Legacy and native runs must follow the same rules. **Proposed solution** Show all subtasks in one section. Separately group tasks created from the source task by their current project, with a No project group when needed. A created subtask appears in both sections. Use saved run context and recorded creation activity to find the source task. **Alternatives considered** Making all created tasks children would change their hierarchy. Removing overlap between sections would hide the creation relationship. This change keeps the two memberships separate. **Roadmap alignment** This extends the existing Activity log & action attribution capability. It does not add a new roadmap area. Related PR: #9727 adds a stored source-task field and inbound attribution UI. This PR adds the outgoing task list using existing run and activity records and does not require that schema change. ## What Changed - Add a company-scoped createdFromIssueId filter to issue lists. - Save the actor run during task creation, including legacy child-helper calls. - Recover historical run attribution from creation activity when the origin run is absent. - Render the production Tasks panel with all subtasks and independently grouped created work. - Keep progress only for subtasks. Add folding, hover fades and project links. - Fetch all result pages and refresh on issue activity. Show load failures with Retry. - Add database, API, UI and pagination tests, design-guide examples and Storybook pages. ## Verification - Before rebase: 158 targeted tests passed. Workspace typecheck, UI/server builds, Storybook build and token gates passed. - After rebase: full workspace typecheck and build passed. The cursor fix passes 26 focused tests and UI/server typechecks. - The full local test command completed its general-server group with 10,563 passing tests and two failures: a missing native-runner fixture and a concurrency-test timeout. Building the fixture and rerunning both affected files passed all 41 tests. The local command stopped before its remaining groups; all corresponding GitHub test shards passed on the submitted head. - GitHub checks on commit |
||
|
|
5545f6d166 |
feat: let user messages continue stopped native tasks (#13239)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A failed native run can leave a durable execution hold. > - The hold prevents automatic replay of actions with unknown outcomes. > - It can also prevent the agent from answering a new user message. > - A new user message should authorize a fresh turn after the prior execution stops. > - This pull request adds that admission path and preserves the existing execution gates. > - Users can continue the conversation without certifying every past action. ## Linked Issues or Issue Description **Subsystem affected** Server task wake admission and native execution recovery. **Problem or motivation** A native task can remain blocked after automatic recovery stops. A new user message is saved, but its run is cancelled before the agent can answer. **Proposed solution** Use a new authenticated user comment to authorize a fresh turn. Check stopped predecessor ownership and available history. Retain uncertain action outcomes. Commit the new run and hold retirement together. **Roadmap alignment** This is a focused improvement to the existing self-healing runs and recovery behavior. Builds on merged #13237, which covers legacy conversation continuation. This PR adds native admission and preserves native automatic-recovery eligibility and budgets. ## What Changed - Admit a fresh native turn for a new user comment after every held predecessor has stopped. - Validate the comment author, task, timing, process ownership, controller, and cleanup leases. - Preserve failed runs, unknown action outcomes, and the failed incident's attempt count. - Record the new comment and run in the existing recovery audit history. - Validate the saved continuation source and discard consumed user-wake authority from later automatic replacements. - Keep pause, approval, budget, ownership, and dependency interaction rules. - Add database and actual wake-path regressions. Update the execution contract. ## Verification - [Full CI run 34626750213](https://github.com/paperclipai/paperclip/actions/runs/34626750213) passed on `c58e6c087e0df7530c747d80b27d491da925a9c4`: all 31 reported checks passed, including all server/workspace suites, browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. - Greptile reviewed this exact head at 5/5. All review threads are resolved, and security checks passed. - All 218 local targeted tests passed across explicit native continuation, continuation history, safe replacement, durable chat wakeups, wake queue, issue liveness, native session resume, and run dispatch. The native implementation is unchanged by the final rebase onto master. - Full workspace `pnpm -r typecheck` and `pnpm build` passed on the final head. Complete test coverage is supplied by the green CI suites; local tests used the targeted suites above. - Regressions cover scoped authorization, concurrent delivery, live ownership, later admission gates, retained message receipts, and automatic replacement after terminal-task or reviewer changes. ## Risks - A fresh model turn can choose to repeat an action. Paperclip preserves prior history and does not replay recorded calls. - Missing process identity and remote ownership without a target-aware stop proof retain the hold. A terminal database row alone does not prove that execution stopped. - No schema or dependency changes. Existing historical tasks are not awakened by deployment. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. This session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.11 |