mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
30c63af0e6197e4d903e3f62572cc34472ee9aba
4346
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
30c63af0e6 |
fix(cli): recover abandoned workspace build locks (#13288)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The test-drive command starts an isolated instance for local testing. > - Source startup builds shared packages before it starts the server. > - An interrupted build can leave an empty lock directory. > - Later starts wait without output and fail after 60 seconds. > - This pull request recovers abandoned locks and shows build progress. > - Local testing can start again without manual lock removal. ## Linked Issues or Issue Description **What happened?** `pnpm paperclipai test-drive` stopped at “Starting Paperclip server…” in a source checkout. A leftover plugin build lock caused a silent 60-second wait and then a timeout. **Expected behavior** Startup should recover an abandoned build lock. It should show when it waits for a live build. An interrupted or failed build should not leave partial output that the next start accepts as complete. **Steps to reproduce** 1. Leave an empty `node_modules/.cache/paperclip-plugin-build-deps.lock` directory after an interrupted build. 2. Make the shared or plugin SDK build output out of date. 3. Run `pnpm paperclipai test-drive --api-key placeholder --no-browser` with a fresh data directory. 4. Observe the silent wait at server startup. **Paperclip version or commit** Reproduced at `2083bf6f9`. **Deployment mode** Local source checkout with an isolated embedded PostgreSQL instance. Related work: #12894 added test-drive. #12898 restored its credential inputs. Neither change handles abandoned workspace build locks. No duplicate fix was found. ## What Changed - Publish a lock directory with an owner record in one rename. - Recover locks after their owner and compiler exit. Recover legacy empty locks after two minutes. - Keep the lock until the compiler stops on SIGINT or SIGTERM. - Print build and lock-wait progress. - Record source, dependency, compiler-config, and output content fingerprints only after a successful compile. Recover partial output even when modification times are unchanged. - Add 12 process-level regression tests and update the development guide. ## Verification - `node --test scripts/__tests__/ensure-plugin-build-deps.test.mjs`: 12 tests pass. - `pnpm exec vitest run --config cli/vitest.config.ts cli/src/__tests__/test-drive.test.ts`: 32 tests pass. - `pnpm --filter paperclipai typecheck`: passed. - `pnpm --filter paperclipai build`: passed. - Live smoke tests: fresh startup and startup with an abandoned lock both reach ready state. The API and UI respond. The command creates the company and CEO and enables worktree execution. Test instances stop cleanly. - Full repository `pnpm -r typecheck` and `pnpm build`: passed. - Full Vitest suite coverage completed using the repository-supported server, chat, workspace, and serialized shards. The initial local run needed the fresh-worktree fake native-provider binary built and focused reruns for port/socket races and load-related timeouts; all affected tests passed on rerun. Suites skipped by fail-fast exits were run separately and passed. The initial serial `pnpm test:run` was stopped in favor of these shards. - Greptile: 5/5 on commit `8b5a790c2af06a52b5dc76e5f52331966df990b8`, with all review threads resolved. - CI: 31 checks passed and two Storybook checks intentionally skipped. The initial workspace and browser jobs were interrupted by runner shutdowns; both passed on the second attempt. Build, typecheck, canary dry run, all general and serialized tests, all browser shards, security checks, and final verification summaries are green. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34654730783) ## Risks - This changes shared source-build locking for the CLI and plugin SDK commands. - Legacy locks have no owner identity. Recovery uses a two-minute age threshold for empty legacy directories. - Startup reads and hashes source and output files to verify the build cache. Identical direct builds reuse the cache. Changed or partial output requires a rebuild. - A reused process ID can delay recovery. Live owner or compiler processes keep their lock. - No database, API, or UI contract changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, code execution, and process-level testing. A more specific API model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a38ccf9972 |
fix: retry transient continuation admission locks (#13290)
## Thinking Path > - Paperclip manages AI agents and the tasks that they execute. > - The run scheduler checks continuation authority before it starts a provider. > - This check uses database locks to order execution against conversation closure. > - A short lock conflict could fail a valid user follow-up before the provider started. > - This pull request retries the admission transaction after the locks are released. > - Valid work can start after normal contention, while closure and cancellation still stop execution. ## Linked Issues or Issue Description Refs #13038. Related continuation work: #13270 and #13239. **What happened?** A user comment started a run through the automation queue. Its source records and admission marker were valid. A database lock conflict at dispatch caused `chat_control_recovery_proof_unresolved` and stopped automatic recovery. The provider received no work. **Expected behavior** Retry short database lock conflicts before failing admission. Read current ownership and conversation-close evidence on each attempt. Do not retry provider execution. **Steps to reproduce** 1. Queue a user follow-up through the automation transport. 2. Hold the task row lock in a separate transaction at the dispatch boundary. 3. Release the lock after 250 ms. 4. Before this fix, the run fails before provider dispatch. With this fix, the run passes admission once the lock is released. **Paperclip version or commit** Reproduced on base commit `1c4bcff2b`. Disabling the new retry reproduces the original error in the regression test. **Deployment mode** Self-hosted server with PostgreSQL. Regression tests use embedded PostgreSQL. ## What Changed - Retry rolled-back admission transactions after lock conflicts, with up to 50 waits of 100 ms. - Keep queue claims nonblocking. Keep provider dispatch outside the retried transaction. - Recheck current run state and committed close evidence after every conflict. - Explain persistent database contention in the exhausted admission error. - Add real database contention tests and bounded retry tests. Document the behavior. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/services/chat-control-admission-retry.test.ts`: 273 passed. This includes task, wake, and run locks, the native runner, close/cancel races, unrelated failures, and retry exhaustion. - Regression proof: disabling retries makes the user-follow-up test fail with `chat_control_recovery_proof_unresolved`. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm test:run`: stopped after all equivalent CI server/workspace shards passed. The local run exposed a missing `fake-codex-app-server` fixture binary in the fresh worktree; after `pnpm --filter @paperclipai/paperclip-runner run build:rust`, the complete affected `native-session-resume.test.ts` suite passes (37 tests). - Greptile: 5/5, no findings, on commit `c974a496a`. - CI: all 31 checks passed on commit `c974a496a`, including all server/workspace test shards, browser tests, typecheck, runner verification, build, and the release dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34654770074). ## Risks - A contended dispatch can wait through 50 short delays, plus transaction time. - Persistent contention still fails closed after the retry budget. Invalid source evidence fails without retrying admission. - No schema, permission, provider retry budget, or queue-claim behavior changes. ## Model Used OpenAI Codex, GPT-6. The exact deployed model ID and context-window size are not exposed in this session. Used reasoning, repository inspection, code editing, and command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9031516a7e |
fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy conversation adapters can run in Daytona sandboxes. > - A server restart during provisioning can occur before the invocation event exists. > - Recovery then lacks the old adapter identity and leaves a hold that ordinary user retries cannot clear. > - A remote launch can also fail when its host relay looks for Node in the sandbox PATH. > - This pull request records the adapter at claim time and restores explicit user continuation after verified cleanup. > - Users can recover from the task or inbox while the failed run and uncertain action history remain intact. ## Linked Issues or Issue Description Refs #13237, #13239, #13254. Those changes cover recorded conversation runs, native user continuation, and explicit remote Stop. This change covers legacy failure before `adapter.invoke` and exact task/inbox Retry. Refs #9771 for overlapping generated-command quoting. This change also supplies the absolute host Node executable. Refs #13163 and #13264 for the separate native restart and retained-workspace work. **What happened?** A legacy Daytona run interrupted during provisioning became `process_lost` without an invocation event. Recovery preserved an execution hold, and Retry or a new task reply could not resume it. Cleanup could also run before the Daytona plugin was ready. On a macOS host, a subsequent ACP relay launch failed with `env: node: No such file or directory` because the remote launch environment did not contain the host Node path. **Expected behavior** An interrupted conversation can continue after its previous execution stops. Explicit Retry and new user replies should start a fresh turn with the task history. Cleanup failures must remain visible and recoverable. The host relay must use the host Node executable. **Steps to reproduce** 1. Use a legacy Claude adapter with a Daytona environment. 2. Interrupt the server after it acquires the sandbox lease and before it records `adapter.invoke`. 3. Restart and inspect the task hold. 4. Retry from the task or inbox, or send a new task reply. 5. Confirm the old sandbox has stopped and one new response arrives. **Paperclip version or commit** Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`. The branch is rebased on `51b0e01ea`, including #13261 and #13270. **Deployment mode** Built from source on macOS with a real Daytona sandbox and the legacy Claude ACP adapter. ## What Changed - Count new browser specs with the scheduler's median duration in the shard-balance check. This fixes a false policy failure after new specs arrive from both branches. The balance threshold is unchanged. - Persist server-owned adapter identity in the queued-to-running claim before provisioning starts. - Wait for provider plugin startup before restart cleanup. Keep failed cleanup leases as active ownership blockers. - Admit exact board retries and new user comments after verified termination. Retain the old run, task history, approvals, and unknown action outcomes. - Adopt repeated Retry requests. Permit one scoped cleanup attempt per explicit user Retry after the automatic limit, with an activity record. A later user Retry can recover after a transient provider failure; automatic attempts remain capped. - Resume replies deferred during cleanup, including historical legacy startup failures. - Launch the host ACP relay through the absolute host Node executable. - Add a task-level Retry button and return actionable blockers when retry admission is refused. - Add database regressions and three browser recovery journeys. Exclude installed third-party dependency skills from the shipped-skill audit. ## Verification - Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution retains the saved-message recovery, local stop receipts, and wait reasons from #13270 alongside exact legacy Retry support. - Real Daytona: interrupted the server after lease acquisition and before adapter invocation. Restart cleanup confirmed provider termination. Task Retry cleared a seeded historical hold and a real Claude agent returned `Recovery verified.` in the task. Removed the disposable sandbox and environment after testing. - All three browser recovery journeys passed again after the final rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh successor, completed the task, preserved the failed run, and retained the answer after reload. - All 29 e2e/server shard-partition tests passed. The balance check now uses the scheduler's median fallback for unmeasured specs, with the same balance threshold. - Server typecheck passed after rebuilding the generated runner dependencies. The combined recovery/route run passed 136 of 137 tests. Its remaining route test timed out during the first cold module import at its explicit 10-second limit; an isolated rerun reproduced that timeout and passed the other 51 route cases. The complete CI suite passed on this head. The same route file passed all 52 cases in CI, including the first cold import in 7.5 seconds. - Before the final rebase, recursive typecheck, full build, UI token gates, 132 targeted server tests, and the complete [CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085) passed. The subsequent CI failure was the shard-balance accounting mismatch fixed here. - Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable findings. The complete [current CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949) passed on attempt 2. All test, typecheck, build, and canary jobs passed on the first attempt. Docker setup timed out fetching BuildKit from Docker Hub; retrying that job and its dependent aggregate succeeded. ## Risks - Recovery admission changes executable authority. Company, task, agent, user, approvals, process ownership, and provider termination checks remain required. - Explicit continuation starts a fresh conversation with history. It does not certify unknown external action outcomes or rerun non-conversation adapters automatically. - Changing task status alone does not clear an execution hold. The task now offers an explicit Retry action. - Historical adapter claims and invocation events take precedence over current agent settings. Known process or webhook runs retain their hold. Pre-upgrade rows with no adapter evidence may receive only a new explicit user turn after termination proof; they do not become eligible for automatic replay. - No schema migration or sandbox-image change is required. This branch has not been deployed to production. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution, browser automation, and test execution. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
37d7dfb0e3 |
ci: allow dependency changes in cloud eval verification (#13286)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployment requires source verification for the exact merged commit. > - Contributor PRs leave lockfile updates to a separate bot PR. > - Most release checks can refresh an outdated lockfile while installing dependencies. > - Two Runner checks still require a frozen lockfile and fail after dependency changes. > - This PR gives those checks the same install policy as the other release checks. > - A valid dependency change can become deployable without waiting for another merge. ## Linked Issues or Issue Description Refs #13257. The dependency change in #13256 exposed this gap. The separate lockfile update is #13279. Related #12115 addresses the bot PR check trigger; this PR fixes exact-source cloud verification itself. **What happened?** [Cloud readiness forcanary/v2026.911.0-canary.16 |
||
|
|
f2e2d38972 |
feat(ui): summarize completed runner activity groups (#13274)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows tool activity between agent updates. > - Finished groups currently keep the last raw command on screen. > - Failure totals add visual weight to normal retries. > - This pull request summarizes finished groups with short action phrases. > - Users can expand each group to inspect the original activity and output. > - The benefit is a quieter history that still explains the work. ## Linked Issues or Issue Description **What existing behavior does this improve?** Runner activity in live task chat and saved history. **Current behavior** A finished group shows its last command or tool target. A separate failure count remains visible after retries. **Proposed behavior** Show short phrases such as “Ran commands” or “Read files, ran commands” when a commentary group finishes. Remove failure counts in both active and finished groups. Keep original output in expandable history. **Reason and benefit** The activity summary describes the work without filling the conversation with shell arguments or treating each retry as an alarm. Unsuccessful reads and edits use “Checked files” and “Worked on files” to avoid claiming success. **Additional context** Refs: #13255. This builds on the merged rolling activity groups. Related open PR #13246 covers steering and status presentation; this change concerns completed summaries and failure counts. The design was reviewed in desktop and mobile Storybook previews. ## What Changed - Add a deterministic summary shared by builtin and provider-native activity. - Combine repeated categories and retries, preserve category order, and bound long summaries. - Summarize each inactive commentary group while the next group can still be running. - Keep expansion, individual output disclosures, and active rolling rows. - Use the production component in eight review stories, including animated desktop and mobile transitions. - Cover retries, unsuccessful edits, unknown tools, reasoning-only groups, saved history, and resuming work. ## Verification - Focused summary, group, and runner-turn tests pass: 62 tests. - Token gates pass. - The approved Storybook previews passed browser checks for desktop, mobile, keyboard expansion, live-to-completed transitions, and original retry output. - `pnpm -r typecheck`, `pnpm build`, and the production Storybook build pass. - The complete UI suite passes: 582 files and 6,005 tests. The full repository test run is in progress after repairing missing symlinks in the local PostgreSQL dependency. A native workspace suite that hit the setup failure now passes all six tests. - Greptile is 5/5 on the current commit, with no review threads. Security checks pass. All individual CI jobs pass except server shard 4/5, where one plugin-worker output timing assertion failed. The complete plugin-worker suite passes locally (83 tests). The workflow finished. One retry of the failed shard and aggregate check was dispatched and is queued. Squash auto-merge is enabled and remains gated on required checks. - Rechecked the production stories in the browser: mobile history stays on one line, and expansion survives completion while the next group remains active. - Review in Storybook under Tasks → Completed activity preview. Use Next to finish one group while the next group remains active. Open a completed summary to inspect its history. ## Risks - Summaries classify observed activity; “Ran commands” does not mean exit code zero. - Unknown tools use a generic description. Full names and outputs remain in history. - No API, database, or runner protocol contracts change. > Reviewed ROADMAP.md. This is a focused improvement to the existing task-chat presentation. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c4bcff2b1 |
fix(test): await issue lock before retry race assertions (#13273)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Retry decisions must respect changes to an issue owner. > - Database tests verify this with two concurrent transactions. > - The test started the competing operation before its fixture held the lock. > - That race can fail a correct source verification run and block deployment. > - This PR waits for lock acquisition before starting the competing operation. ## Linked Issues or Issue Description Refs #13257 for the deployment verification work that exposed this test race. No duplicate fix was found. **What happened?** [Cloud verification job 103424158921](https://github.com/paperclipai/paperclip/actions/runs/34648268409/job/103424158921) failed because the test did not observe a concurrent issue-row lock waiter. The fixture and retry operation both started without an ordering guarantee. **Expected behavior** The fixture must hold the issue lock before the competing retry operation starts. The test must still prove that the retry waits for the lock and observes the reassignment. **Steps to reproduce** 1. Add a temporary 50 ms delay before the fixture acquires the issue lock. 2. Run the promoteOrCancelDueRetry issue-lock test. 3. The original fixture fails with the same missing-waiter error as CI. 4. The synchronized fixture passes with that delay. The delay is not part of this PR. ## What Changed - Separate fixture readiness from its transaction completion promise. - Wait for readiness at both callers before starting the retry decision. - Propagate transaction failure during setup through Promise.race. ## Verification - The original fixture fails under the temporary delayed-lock probe. The fixed fixture passes the same probe. - All 31 tests in server/src/modules/run-dispatch/adapters/postgres.test.ts pass after removing the probe. - The real concurrent waiter, lock-order, and reassignment assertions remain intact. No timeout was increased. - Full local typecheck and build pass (221s and 50s). The full local test command stopped in its server phase after 10,610 passes, 65 skips, and 14 failures: 13 existing macOS skill-cache rename/permission failures and one unchanged Telegram test assertion. The Telegram test passes in a focused rerun. Later local phases did not run after that failure. Current-head Greptile is 5/5 with no findings. All 32 current-head checks pass, including full Linux typecheck, build, native verification, server suites, browser suites, and canary packaging. The unchanged GitHub browser mock assertion passed its single failed-shard retry. - git diff --check passes. ## Risks - Only test synchronization changes. Production database behavior is unchanged. - Awaiting the transaction itself during setup would deadlock the test. Returning the completion promise inside an object avoids that problem. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — all 31 database adapter tests pass; full local verification limits are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51b0e01ead |
fix: resume saved user messages after execution recovery (#13270)
Preserve verified native process-stop evidence and retry saved user messages through normal continuation admission after recovery cleanup. Show the current wait reason and serialize delivery so a saved message starts one fresh turn. Validated with 410 focused tests, typecheck, build, token gates, all PR CI checks, and Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
250deab910 |
fix(runner): keep healthy native sessions alive (#13261)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions keep a provider process and its work alive across control-plane operations. > - A hidden 15-minute turn deadline stopped work even when the agent timeout was zero. > - A one-hour runner lifetime and fixed connection lease added two more limits. > - Recovery also rejected goal commands because it reconstructed their startup summary with the wrong protocol version. > - This pull request removes implicit duration limits and renews authenticated leases in the harness. > - Healthy sessions can continue without model action or a user interface change. ## Linked Issues or Issue Description Refs #13092 and #12845. Related: #13163 covers sandbox recovery after app restarts; this change covers session duration and lease renewal. **What happened?** A native Codex session stopped after 15 minutes while a tool was still running. The agent had `timeoutSec: 0`. Recovery then rejected a `session.goal.get` startup command with `invalid provider startup ownership fence`. **Expected behavior** An unlimited session keeps working while its provider and authenticated controller remain healthy. Lease maintenance is transparent. Explicit timeouts, cancellation and revoked authority still take effect. **Steps to reproduce** Start a native session with `timeoutSec: 0` and run a tool beyond 15 minutes. Before this fix, the runtime cancels the turn. A recovery startup that uses a goal command also exposes the protocol-version mismatch. ## What Changed - Honor the agent turn timeout. Zero disables the timer. Long explicit durations use timer chunks to avoid Node timer overflow. - Default native runner lifetime to unlimited. Keep bounded startup, reconnect and control-operation deadlines. - Renew leases over the authenticated connection. Persist renewal before the reply. Validate identity, epoch and expiry. Handle duplicate requests and a lost reply on reconnect. - Freeze renewal during warm ownership transitions and terminal handling. - Validate persisted goal startup commands with protocol v2. - Add duration, renewal, ownership, recovery and real-process regression tests. Update runner protocol and recovery docs. - Add no UI components or controls. Renewal requires no model output or user action. ## Verification Current head: `348e369c35c5da8bb8be378f4b35dcf6f40882e7`. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34649767113). - All 32 checks pass on this head. The two Storybook checks are skipped as expected. CI includes full build, typecheck, runner verification, browser suites, server tests and the canary package dry run. - Greptile reports 5/5 on this head. All review threads are resolved, and the security scan passes. - Passed `pnpm -r typecheck` and `pnpm build` locally. - Passed 219 native-runtime and controller tests, including fake-clock tests for three weeks of renewal and 30-day explicit timeouts. Six denial tests confirm that renewal cannot extend expired, revoked or mismatched authority. - Passed 337 executor, cancellation and restart-recovery tests, plus 278 Rust runner-core library tests. - Passed a real runner with a silent fake Codex provider across its original lease expiry. Runner PID, provider PID, thread and active turn stayed unchanged. Warm-attach recovery tests also pass. - Passed all 83 plugin-worker tests and 159 of 161 workspace-runtime tests locally. The two remaining assertions passed with a canonical macOS temporary directory, as did the changed runtime fixture. The full affected server shard passes in CI. - An unchanged GitHub callback-ordering test failed once in CI, passed locally in isolation, and passed its one test-shard retry. The final CI summary is successful. - The full local `pnpm test:run` sweep was interrupted after dependency setup failures and load-related timeouts. Identified failing suites passed in isolated reruns after the dependency repair. The complete test matrix passed remotely in CI. ## Risks - Deploy the controller and runner together to enable renewal. Older peers keep their existing bounded lease behavior. - Unlimited runtime permits long resource use until completion, explicit cancellation, configured timeout or loss of valid authority. - Lease renewal changes authenticated protocol handling. Regression tests cover stale, revoked and mismatched authority, lost replies and warm handoff behavior. - Simulated multi-week tests and a real lease-boundary test do not constitute a weeks-long production soak. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution and TypeScript/Rust test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a12bbd1824 |
fix: stop completion reviews caused by policy upgrades (#13266)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runtime records completion assessments and task decisions. > - An application update can change the assessment policy version. > - The previous code treated that version change as a reason for human review. > - An upgrade alone does not give the user a new decision to make. > - This change keeps existing decisions and withdraws obsolete upgrade review cards. ## Linked Issues or Issue Description **What happened?** A rules version change moved unfinished tasks into review and created a card that said, "Review the superseding native policy assessment." The task did not need new work or a human decision. **Expected behavior** New runs use the current rules. An upgrade leaves existing task decisions alone. The saved policy version remains available in the audit history. **Steps to reproduce** 1. Save a native run assessment and a task decision. 2. Change the native status policy version. 3. Run finalization reconciliation without new task evidence. 4. The previous code created a review card. The corrected code keeps the saved assessment and decision. Related context: #13038 changed the native policy version. Searches found no duplicate fix for upgrade-only completion reviews. ## What Changed - Remove policy-version mismatch as a reconciliation trigger. - Withdraw pending cards only when their source decision, effect ledger, creator, key, and prompt match the old upgrade-only review. - Restore the previous status only while that decision and status version remain current and no other review gate is pending. - Preserve answered cards, real review requests, later task changes, and historical assessments and decisions. - Record cleanup activity and retire any corresponding chat review actions. - Isolate cleanup failures so one old card cannot block other cleanup or normal finalization. - Update the status conformance fixture, regression tests, and architecture documentation. ## Verification - Passed: `pnpm exec vitest run server/src/__tests__/native-status-arbiter-corpus.test.ts` (23 tests, including the 53-fixture status corpus). - Passed: `pnpm -r typecheck`. - Passed: `git diff --check`. - Passed: `pnpm build`. - Passed: fresh Greptile review at 5/5 on `d24e5ecaf`, with no open review threads. - Passed: the 23 focused tests in the isolated full-suite environment. - Local `pnpm test:run`: the server group finished with 10,638 passed and one missing-fixture failure. Built the required `fake-codex-app-server` fixture and reran the entire affected native session-resume suite: 37/37 passed. The initial full command exited on that server-group failure, so remaining groups are covered by CI. - Passed: the workspace-runtime-exposure suite (25 tests, 3 platform skips). - Passed: all CI gates, including every test shard, browser tests, typecheck, and build. The server shard passed on one retry after an unrelated host-port conflict in the unchanged workspace-exposure tests. - Cleanup tests cover repeat runs, real reviews, answered cards, later statuses, newer decision identity, status changes back to review, other pending requests, run scope, cleanup failures, retry, and live publication failures. ## Risks - Cleanup changes stored task state. It checks the exact obsolete decision and status version under database locks before restoring status. - Pending approvals, interactions, and execution stages prevent restoration out of review. - The cleanup handles at most 100 matching cards per reconciliation pass. It does not rewrite old decisions or accept an agent's completion claim. - New evidence and explicit task changes still use the existing reconciliation paths. This change does not reevaluate old work merely because Paperclip was updated. ## Model Used OpenAI Codex, GPT-6. The exact deployment identifier and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.15 |
||
|
|
19c76bfc3f |
fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - CI installs dependencies before it verifies and builds cloud artifacts. > - Install jobs share a pnpm package-store cache with lockfile refresh. > - Lockfile refresh resolves versions without downloading packages. > - That job saved an empty cache before full install jobs could save theirs. > - This PR prevents lockfile refresh from publishing that empty entry. > - Full install jobs can then populate the cache and reuse dependencies. ## Linked Issues or Issue Description Refs #13259 for the related cloud verification cache work. No duplicate empty-cache fix was found. **What happened?** Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm cache at 18:58:08 UTC on September 10. Full install jobs still restore that empty entry. The cache API reports 216 bytes for master and about 703 MB for populated entries with the same key and cache version in PR scopes. **Expected behavior** A job that installs dependencies should populate the shared package-store cache. **Steps to reproduce** 1. Run lockfile refresh with a new lockfile cache key. 2. Its resolution-only command leaves the package store empty. 3. The Node action saves the empty archive before a full install finishes. 4. Later jobs report a cache hit but download packages again. **Paperclip version or commit** Observed on master |
||
|
|
778ac33308 |
feat: accept a base64-encoded Cloud UI snippet (#13245)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip Cloud instances can inject a trusted operator HTML snippet before `</body>` in served UI pages (#13168) > - Operators deliver env vars to managed instances through hosting-provider APIs > - Web application firewalls in front of those APIs reject request bodies that contain raw `<script` markup > - The snippet value always contains a script tag, so the firewall blocks its delivery every time > - This pull request adds `PAPERCLIP_CLOUD_UI_SNIPPET_B64`, which carries the same snippet as standard base64 > - The benefit is that operators can ship the snippet through WAF-fronted delivery pipelines without firewall exceptions ## Linked Issues or Issue Description Refs #13168 (introduced the Cloud UI snippet). **Subsystem affected** Server UI shell serving: `server/src/cloud-ui-snippet.ts`, used by `static-index-html.ts` and `app.ts`. **Current behavior** `PAPERCLIP_CLOUD_UI_SNIPPET` is the only way to configure the Cloud UI snippet. The value is raw HTML. Delivery pipelines that write env vars through provider APIs can fail to deliver it: web application firewalls classify a request body that contains `<script` as an injection attempt and block it. Verified against Railway's GraphQL API, which sits behind Cloudflare: a variable value with a bare `<script></script>` returns HTTP 403 before authentication, and JSON unicode escapes (`<script`) do not bypass the block. The snippet is exactly the kind of value that always contains a script tag, so such pipelines cannot deliver it at all. **Proposed behavior** A new optional `PAPERCLIP_CLOUD_UI_SNIPPET_B64` carries the same snippet as standard base64 of the UTF-8 HTML. The server decodes it and injects the result through the existing path. The plain variable wins when both are set. Whitespace and line wrapping in the value are tolerated. A value that is not canonical base64, or that decodes to blank, is ignored instead of injected as garbage. **Reason and benefit** Operators can deliver the snippet through WAF-fronted APIs without requesting firewall exceptions. Base64 contains no markup, so the firewall has nothing to match. Existing deployments see no behavior change. **Breaking changes** None. The new variable is optional. The existing variable is unchanged and takes precedence. ## What Changed - `server/src/cloud-ui-snippet.ts`: resolve the snippet from `PAPERCLIP_CLOUD_UI_SNIPPET` first, then from base64-decoded `PAPERCLIP_CLOUD_UI_SNIPPET_B64`; ignore non-canonical or blank-decoding values. - `server/src/__tests__/cloud-ui-snippet.test.ts`: cover decode-and-inject, plain-wins precedence, whitespace tolerance, invalid/blank values, and the self-hosted no-op. - `doc/cloud-ui-snippet.md`: document the variant, how to produce the value, and the ignore rules. ## Verification - `pnpm vitest run server/src/__tests__/cloud-ui-snippet.test.ts server/src/__tests__/static-index-html.test.ts` — 13 tests pass. - `tsc --noEmit` in `server/` passes. - Manual check: on a Cloud-managed instance, set `PAPERCLIP_CLOUD_UI_SNIPPET_B64="$(base64 < snippet.html)"`, restart, open `/`, and confirm the decoded snippet appears before `</body>`. On a self-hosted instance, confirm no snippet is injected. ## Risks Low risk. The injection path and the Cloud-managed gate are unchanged. A malformed base64 value is ignored, so the failure mode is a missing widget, not corrupted HTML. The decoded value is trusted operator configuration, the same trust model as the plain variable. ## Model Used - Claude Fable 5 (Anthropic, `claude-fable-5`) via Claude Code CLI, agentic workflow with tool use (code edits, test runs, and live API verification of the WAF behavior). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc68312327 |
ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.
## Linked Issues or Issue Description
Refs #13243.
**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.
**Current behavior**
For merge
|
||
|
|
ad4f0b5867 |
Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtime settings can bind organization secrets to an adapter environment > - Paperclip redacts plain environment values when it returns a saved agent to the UI > - A saved-agent test sent the redacted `CODEX_HOME` value back to the server > - Codex ACP also received the API key without an ACP API-key authentication request > - This pull request restores saved environment values for tests and selects API-key authentication for Codex ACP runs > - The benefit is that Codex agents can test and run with an organization-scoped OpenAI API key ## Linked Issues or Issue Description **What happened?** Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`. Secret normalization rejected that placeholder. Remote Codex ACP runs received `OPENAI_API_KEY`, but session creation stopped with `Authentication required`. **Expected behavior** Paperclip must use the saved `CODEX_HOME` value when it tests an existing agent. Codex ACP must select API-key authentication when `OPENAI_API_KEY` is available. **Steps to reproduce** 1. Create an organization-scoped secret named `OPENAI_API_KEY`. 2. Give a Codex agent access to the secret. 3. Save the agent runtime settings. 4. Test the saved agent again. 5. Run the agent in a remote sandbox through ACP. **Paperclip version or commit** Reproduced on master before commit `68c17709d7c051a804a416263e2e08920f1dfcb1`. **Deployment mode** Self-hosted server with a remote sandbox environment. **Installation method** Built from source. **Agent adapter(s) involved** Codex. ## What Changed - Send the saved agent ID with adapter environment tests. - Restore redacted plain environment values from the saved agent before test-time secret resolution. - Select the Codex ACP `api-key` authentication method when `OPENAI_API_KEY` is present. - Add focused regression coverage for saved-agent tests and remote ACP launch configuration. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/test-agent-setup.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `git diff --check` ## Risks - Low risk. The test route reads saved configuration only when the request supplies a compatible agent ID and the caller can update that agent. - The Codex ACP change applies only when `OPENAI_API_KEY` exists and no explicit `DEFAULT_AUTH_REQUEST` exists. - There are no schema migrations or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. The context-window size is not exposed in this runtime. The model used reasoning, repository search, file editing, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3bafac12f7 |
refactor: remove automatic productivity reviews (#13263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its recovery loop keeps assigned work moving after execution failures. > - Productivity review used run counts, comment counts, and elapsed time to create management tasks. > - Infrastructure failures could satisfy those rules and create more tasks without evidence that the source work needed management review. > - This pull request removes that detector and its continuation holds. > - Bounded recovery, budgets, explicit blockers, and normal review stages remain in place. > - Existing task records stay readable and unchanged. ## Linked Issues or Issue Description Refs #5897. That request describes unwanted automatic productivity reviews and asks to preserve existing tasks. This change retires the feature instead of adding another configuration switch. Related prior approaches: Refs #9191, Refs #12489. Those changes excluded infrastructure failures or bounded review creation. This removal replaces the detector rather than tuning its thresholds. ## What Changed - Delete the scheduled detector, automatic task creation, evidence refresh, and productivity continuation holds. - Remove computed productivity fields, special attention items, badges, and Storybook fixtures. - Retain historical origin values, decision compatibility, and recovery recursion exclusions. Add no migration and change no existing task data. - Update the execution contract. Replace feature tests with regressions for legacy task reads, ordinary attention, and bounded continuation in the presence of an old review. ## Verification - Targeted attention, issue-route, startup, and UI tests: 4 files and 101 tests passed. - Updated issue-route and UI tests: 2 files and 61 tests passed. - Bounded continuation regression: 2 cases passed, including a legacy review plus pre-dispatch cancellation churn. - `pnpm check:token-gates`: all four gates passed. - `git diff --check`: passed. - `pnpm build-storybook`: passed. - Greptile: 5/5 on `a5a612eea`, with no actionable findings. - Scheduler and historical recovery regressions: 2 files and 28 tests passed. - Repository `pnpm -r typecheck` and `pnpm build`: passed. - The complete `pnpm test:run` suite passed across the CI server, serialized-server, and workspace shards on `a5a612eea`. Stopped the duplicate local monolithic run after the full CI suite passed; no completed local full-suite result is claimed. The targeted local suites above passed. - CI serialized shard 5 initially hit a 10-second timeout in the first interaction-route test. The complete file passed locally (78 tests), then the single CI rerun passed. - All CI gates are green, including the build and end-to-end suites. - A local merge check against current `master` (`ce09ea40b`) completed without conflicts. ## Risks - API responses no longer include the computed `productivityReview` field. Consumers must stop using it. - The scheduler no longer creates management work from elapsed time, run counts, or missing comments. This is the intended behavior change. - Existing review tasks and explicit dependencies remain in place. Historical origins still prevent recursive recovery treatment. No task cleanup or data migration occurs. - The native review handoff repair is separate from this removal. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
canary/v2026.911.0-canary.14
|
||
|
|
a23ae894a5 |
ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud images become deployable only after source verification passes. > - The typecheck job builds the native Runner binary through the server package. > - Fresh runners repeatedly compile Rust dependencies for that binary. > - This PR caches those dependencies for exact-source master verification. > - Workspace code and every typecheck still rebuild or run as before. ## Linked Issues or Issue Description Refs #13243 and #13257. **What existing behavior does this improve?** The typecheck portion of post-merge cloud source verification. **Current behavior** The typecheck job has no Rust dependency cache. An observed release build in this job took 4m 13s, including dependency compilation. **Proposed behavior** Restore dependency build outputs for canonical master pushes with the exact source SHA. Use a separate cache key from the Runner verification job, which builds other profiles. **Reason and benefit** A warm cache should remove roughly 2–3 minutes of dependency compilation from this job. Overall deployment gains depend on the remaining critical path. The first cache population still compiles from scratch. **Breaking changes** None. All checks remain enabled. Non-master callers compile without restoring or saving this cache. ## What Changed - Select the pinned Rust toolchain before the typecheck cache lookup. - Reuse the existing pinned Rust cache action with a typecheck-specific key. - Exclude workspace crates and installed cargo executables. - Test restore/save trust boundaries and document cache behavior. ## Verification - All workflow script tests pass locally, including nine new cache trust/contract cases. - actionlint passes for release-verify.yml. - Full local typecheck and build pass on the same source base (167s and 206s); `pnpm test:run` is still running and is recorded with #13257. This PR changes only the workflow, cache guard tests, and documentation. - All 32 current-head checks are successful or intentionally skipped, including the complete Linux test matrix, build, and Greptile 5/5 with no unresolved findings. Verify cache population and subsequent restore on actual master runs. ## Risks - The first run and any toolchain/dependency invalidation compile from scratch. - Cache restore/save overhead reduces the benefit for small dependency graphs. - Disable the cache step to roll back; the existing uncached build remains valid. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. Exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — focused change tests pass; full-suite local permission failures are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a27dd722a2 |
fix(ui): restore task page archive shortcut (#13253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail page gives operators keyboard shortcuts for fast task work. > - The `y` shortcut must archive the open task from the inbox and return the operator to the inbox. > - A streamlined UI gate limited this shortcut to tasks opened from a selected inbox row. > - Direct task pages no longer ran the shortcut. > - This pull request removes that gate and adds regression coverage. > - The benefit is consistent keyboard operation from every visible task page. ## Linked Issues or Issue Description **What happened?** The `y` keyboard shortcut did nothing on a task page opened outside the inbox. **Expected behavior** The `y` shortcut archives the open task from the inbox and returns the operator to `/inbox`. **Steps to reproduce** 1. Enable keyboard shortcuts. 2. Open a visible task from the Tasks page. 3. Press `y` while focus is outside an editor. 4. Observe that the task does not archive. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local development. ## What Changed - Enable the `y` archive shortcut on every visible task detail page when keyboard shortcuts are enabled. - Keep the existing input, dialog, modifier, hidden-task, and pending-request guards. - Verify that `y` archives the task and navigates to `/inbox` for a task opened from the Tasks page. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/keyboardShortcuts.test.ts ui/src/lib/issueDetailBreadcrumb.test.ts` (139 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build` - The full workspace typecheck and build reached the Rust runner. This environment does not include `cargo`, so the Rust steps did not run. - The full test suite entered unrelated serialized external-chat integration tests. It was stopped after the affected UI suite passed. ## Risks - Low risk. The change restores the shortcut behavior that existed before the streamlined UI gate. - The shortcut still ignores text inputs, open dialogs, modifier keys, and hidden tasks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex. The Paperclip runtime did not expose the exact model ID or context window. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce09ea40b0 |
feat(ui): show one rolling runner activity per commentary group (#13255)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows commentary and runner activity as work proceeds. > - The expanded activity list uses multiple text lines and an indented rail. > - The folded view can repeat reasoning and the current tool in separate places. > - This pull request shows one rolling activity between commentary messages. > - Users can open a group to follow its full history and inspect each item. > - The benefit is a compact feed with aligned icons and stable row height. ## Linked Issues or Issue Description **What existing behavior does this improve?** The runner activity groups in task chat, during live work and in saved history. **Current behavior** Activity is spread across a group summary, a reasoning ticker, and a current-tool row. Expanded rows have different alignment and can span several lines. **Proposed behavior** Keep one latest activity per commentary group. Roll new activities upward. Keep each expanded history row on one line. Use neutral failure text without an X icon. Keep the full details available by opening an individual row. **Additional context** Refs: #13246. That open PR covers related folded activity, steering timestamps, and status alignment. This change implements the separately reviewed activity-group design, including animated replacement, single-line expanded history, and neutral failures. It does not change steering timestamps or status pills. ## What Changed - Add a shared runner activity group with persistent expansion and reduced-motion support. - Use the group in the production runner turn and remove duplicate narration and current-tool rows. - Preserve commentary, plan cards, request receipts, and final-response placement. - Keep tool output, reasoning, research sources, file previews, and failure details accessible. - Render the production turn in desktop and mobile animated Storybook fixtures. - Add tests for animation identity, failure visibility, expansion persistence, and timeline order. - Validate resource link schemes and omit disclosure controls for items with no details. ## Verification - Focused activity, runner turn, protocol detail, and task-thread tests pass: 161 tests. - `pnpm check:token-gates` passes. - Browser checks confirm a 32-pixel compact row, centered icons, and one-line expanded rows. - `pnpm -r typecheck` and `pnpm build` pass. The final UI build also passes. - Two real native Codex runs completed in an isolated local instance. Each run created the expected file, performed separate delayed tool calls, reported a deliberate missing-file failure, recovered, and finished the task. - Browser verification confirmed that the open three-row activity group remained expanded after the second run moved into saved history. Saved error output remained accessible without red styling or an X. - The complete UI suite passes: 581 files and 5,979 tests. - `pnpm --filter @paperclipai/ui build-storybook` passes for the final production fixtures. - Greptile rates the latest commit 5/5, with zero new comments and no unresolved review threads. The security scan passes. - All repository test shards pass in CI. The duplicate local `pnpm test:run` was stopped after CI completed the full test coverage. The CI build failed twice because the worker ran out of disk space compiling the Rust runner tests (`No space left on device`, error 28). The build and its aggregate verify gate remain blocked by CI capacity. GitHub squash auto-merge is enabled and will wait for required checks to pass. ## Risks - This changes presentation and disclosure behavior. Users open a group to see earlier activity. - Animation uses stable item IDs. A provider that replaces IDs can trigger another transition. - No API, database, or runner protocol contracts change. > Reviewed `ROADMAP.md`. This improves an existing task-chat surface. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2904a3a6cc |
fix(ui): hide retry countdown after execution starts (#13258)
## Thinking Path
> - Paperclip helps operators manage AI agents and their tasks.
> - Task pages show countdowns for deferred checks and automatic
retries.
> - A retry keeps its scheduled start time after it enters the queue or
starts running.
> - The countdown treated that historical time as a pending deadline and
showed an overdue warning beside active work.
> - This pull request limits retry countdowns to retries that are still
scheduled and hides waiting surfaces on terminal tasks.
> - Operators now see a warning only when the displayed retry is still
waiting to start.
## Linked Issues or Issue Description
Refs #9783, which added the monitor surfaces. Searched related PRs and
issues; no duplicate fix was found.
**What happened?**
After a service restart resumed a task through an automatic retry, the
task showed an overdue retry banner while the agent was running. The
banner also remained when the task became done.
**Expected behavior**
A queued or running retry must not show a countdown against its past
scheduled start time. Done and cancelled tasks must not show waiting
banners.
**Steps to reproduce**
1. Open a task with an automatic retry scheduled for a known time.
2. Let the retry enter the queue and start running.
3. Wait until its scheduled time is more than one minute in the past.
4. Observe the overdue banner and Check now button while the agent is
working.
**Paperclip version or commit**
Reproduced on source commit
|
||
|
|
2fc5e3e43b |
test: isolate native controller takeover fixture (#13212)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native runtime recovery protects controller takeover with process identity checks > - The live-controller test mocked process liveness but left process start time to host state > - This pull request supplies the matching process start time in the fixture > - The benefit is deterministic coverage of the live-controller fence ## Linked Issues or Issue Description Closes #13172 ## What Changed - Complete the live-controller takeover fixture with the recorded process start time. ## Verification - Contributor reports focused native restart-recovery tests, related native tests, integration tests, and typecheck passed locally. Maintainer reran the 20 native restart-recovery tests successfully and verified current-head CI, Greptile 5/5, and zero unresolved review threads. ## Risks Low risk. This changes test setup only. Production controller fencing is unchanged. ## Model Used Original patch by @lorenzozanee; the original author did not record model details. Maintainer review and verification used OpenAI GPT-6 through Codex with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change is needed for this test fixture) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
847d00bdc3 |
fix(ui): use full mobile width for task tree (#13250)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The task detail page gives operators a task tree in a mobile side panel. > - The mobile sheet did not set an explicit full width. > - A maximum width rule could make the task tree narrower than the phone viewport. > - This pull request makes the mobile task tree sheet use the full viewport width. > - The benefit is a consistent and readable task tree on phones. ## Linked Issues or Issue Description **What happened?** The task tree sheet could use less than the full viewport width on a mobile task page. **Expected behavior** The task tree sheet must use the full available phone width. **Steps to reproduce** 1. Open a task on a phone viewport. 2. Open the task side panel. 3. Select the Tasks tab. 4. Compare the sheet width with the viewport width. **Paperclip version or commit** `7b829efdf` **Deployment mode** Built from source. ## What Changed - Set the mobile task side panel to `w-full` and `max-w-none`. - Add regression assertions for both width classes. ## Verification - `vitest run src/pages/IssueDetail.test.tsx --config vitest.config.ts` passes 103 tests. - `node scripts/check-token-gates.mjs` passes all four token gates. - `git diff --check` passes. ## Risks - Low risk. The change only affects the mobile task side panel width. - The desktop panel and other sheet variants do not change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, agentic coding with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.13 |
||
|
|
d0b7ba4194 |
ci: route approved master cloud builds to AWS (#13243)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images built from master commits. > - Cloud image builds share GitHub-hosted capacity with other workflows. > - The organization already operates AWS runners through RunsOn Fleet. > - This pull request allows approved master builds to use a dedicated cloud Fleet. > - The benefit is separate build capacity with a quick operator rollback. ## Linked Issues or Issue Description Refs #13189, #13192. **What existing behavior does this improve?** Placement of the Docker cloud build after a master merge. **Current behavior** Every Docker cloud build uses a GitHub-hosted runner. Busy periods delay the job. **Proposed behavior** An operator variable enables the approved cloud Fleet for canonical master pushes and manual master builds. Other events, refs, and repositories use GitHub-hosted runners. ## What Changed - Add a guarded AWS runner selector to the Docker cloud job. - Keep the existing image cache, verification, and publication steps. - Test the selector against master, branch, tag, PR, fork, and disabled contexts. - Document provisioning requirements, placement checks, and rollback. ## Verification - 29 focused Node tests pass for routing, readiness, and disk handling. - The full workflow-script Node suite passes. - `pnpm -r typecheck` passes locally. - The pinned PR routing regression suite passes. The first live PR run assigned 21 jobs to the approved AWS PR group. AWS then reclaimed 16 Spot instances. The failed run is being repeated on GitHub-hosted runners while the Fleet moves to On-Demand. - Actionlint passes with existing shellcheck findings excluded (SC2012, SC2016, SC2129). - `git diff --check` passes. - Greptile reports 5/5 on commit `764d505a41dd2023751c3f361906fa9ea35bf0c6`, with no review threads. - All 30 current-head CI checks pass, including typecheck, build, all server/workspace test shards, Runner verification, and browser tests. Two Storybook checks are intentionally skipped for this change. Run: https://github.com/paperclipai/paperclip/actions/runs/34630550799 - The broader local test/build sequence is still running. This Mac has reported failures in unchanged application suites; their complete Linux CI shards pass. Local targeted workflow tests and typecheck pass. - Both On-Demand Fleets are deployed and healthy. Live master cloud-build verification follows the merge. ## Risks - Missing Fleet capacity or runner-group authorization can leave an AWS job queued. Disable `AWS_CLOUD_BUILDS_ENABLED` and rerun the workflow to use GitHub-hosted capacity. - The runner group must restrict access to this repository and the master version of `docker-cloud.yml`. - Docker needs more disk space than the PR Fleet. Provision 120 GiB disks and retain the free-space check. - This changes image build placement only. Source verification and migrator publication remain separate prerequisites. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.12 |
||
|
|
063ba59ae3 |
Fix workspace-ready notice rendering (#12716)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task thread shows user, agent, and control-plane messages. > - Workspace-ready comments keep agent attribution for audit and authorization. > - These comments also have an explicit system-notice presentation. > - The client used attribution before the presentation contract. > - As a result, workspace-ready comments appeared as agent bubbles. > - This pull request gives the system-notice presentation priority. > - The benefit is a compact workspace notice with expandable details. ## Linked Issues or Issue Description **What happened?** The task thread showed a workspace-ready control-plane comment as a normal agent message. The comment kept agent attribution for audit and authorization, so the client did not use its system-notice presentation. **Expected behavior** The task thread must show a comment with the `system_notice` presentation as a compact system notice. The user must be able to expand the notice to inspect its details. **Steps to reproduce** 1. Open a task with a workspace that becomes ready. 2. Wait for the control plane to add the workspace-ready comment. 3. Observe that the task thread shows the comment as an agent bubble instead of a system notice. **Paperclip version or commit** `master` at `8eaa5caa0`. **Deployment mode** Local dev (`pnpm dev`). **Agent adapter(s) involved** Not adapter-specific. This is a core UI bug. ## What Changed - Give the `system_notice` presentation priority when the task-chat adapter selects the message kind. - Pass the notice presentation data to the task-chat model. - Improve the compact and expanded workspace notice details. - Add component tests, adapter tests, and Storybook cases for workspace notices. ## Verification - `pnpm exec vitest run ui/src/components/task-chat/task-chat-adapter.test.ts ui/src/components/task-chat/TaskChatSystemNotice.test.tsx` passes 15 tests. - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` passes 5,670 tests and reports four pre-existing workspace-runtime failures. The same four failures reproduce when the two unrelated server test files run alone. They do not use the changed task-chat files. ## Risks - Risk is low. Normal agent messages still use the agent presentation. - A comment with an explicit system-notice presentation now uses the compact notice renderer. - Regression tests cover both paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 (`gpt-5`) through Codex. The agent used reasoning, repository tools, and code execution. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b829efdf6 |
feat: show tasks created from a task by project (#13241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can cause an agent to create more tasks. > - Those tasks can belong to other projects or have another parent. > - The subtask view does not show all work created from the current task. > - This pull request adds a Tasks tab with separate subtask and creation groups. > - Operators can follow created work without changing its parent or project. ## Linked Issues or Issue Description **Problem or motivation** Operators need to see all work that an agent creates while running for a task. Parentage alone does not describe this relationship. Legacy and native runs must follow the same rules. **Proposed solution** Show all subtasks in one section. Separately group tasks created from the source task by their current project, with a No project group when needed. A created subtask appears in both sections. Use saved run context and recorded creation activity to find the source task. **Alternatives considered** Making all created tasks children would change their hierarchy. Removing overlap between sections would hide the creation relationship. This change keeps the two memberships separate. **Roadmap alignment** This extends the existing Activity log & action attribution capability. It does not add a new roadmap area. Related PR: #9727 adds a stored source-task field and inbound attribution UI. This PR adds the outgoing task list using existing run and activity records and does not require that schema change. ## What Changed - Add a company-scoped createdFromIssueId filter to issue lists. - Save the actor run during task creation, including legacy child-helper calls. - Recover historical run attribution from creation activity when the origin run is absent. - Render the production Tasks panel with all subtasks and independently grouped created work. - Keep progress only for subtasks. Add folding, hover fades and project links. - Fetch all result pages and refresh on issue activity. Show load failures with Retry. - Add database, API, UI and pagination tests, design-guide examples and Storybook pages. ## Verification - Before rebase: 158 targeted tests passed. Workspace typecheck, UI/server builds, Storybook build and token gates passed. - After rebase: full workspace typecheck and build passed. The cursor fix passes 26 focused tests and UI/server typechecks. - The full local test command completed its general-server group with 10,563 passing tests and two failures: a missing native-runner fixture and a concurrency-test timeout. Building the fixture and rerunning both affected files passed all 41 tests. The local command stopped before its remaining groups; all corresponding GitHub test shards passed on the submitted head. - GitHub checks on commit |
||
|
|
5545f6d166 |
feat: let user messages continue stopped native tasks (#13239)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A failed native run can leave a durable execution hold. > - The hold prevents automatic replay of actions with unknown outcomes. > - It can also prevent the agent from answering a new user message. > - A new user message should authorize a fresh turn after the prior execution stops. > - This pull request adds that admission path and preserves the existing execution gates. > - Users can continue the conversation without certifying every past action. ## Linked Issues or Issue Description **Subsystem affected** Server task wake admission and native execution recovery. **Problem or motivation** A native task can remain blocked after automatic recovery stops. A new user message is saved, but its run is cancelled before the agent can answer. **Proposed solution** Use a new authenticated user comment to authorize a fresh turn. Check stopped predecessor ownership and available history. Retain uncertain action outcomes. Commit the new run and hold retirement together. **Roadmap alignment** This is a focused improvement to the existing self-healing runs and recovery behavior. Builds on merged #13237, which covers legacy conversation continuation. This PR adds native admission and preserves native automatic-recovery eligibility and budgets. ## What Changed - Admit a fresh native turn for a new user comment after every held predecessor has stopped. - Validate the comment author, task, timing, process ownership, controller, and cleanup leases. - Preserve failed runs, unknown action outcomes, and the failed incident's attempt count. - Record the new comment and run in the existing recovery audit history. - Validate the saved continuation source and discard consumed user-wake authority from later automatic replacements. - Keep pause, approval, budget, ownership, and dependency interaction rules. - Add database and actual wake-path regressions. Update the execution contract. ## Verification - [Full CI run 34626750213](https://github.com/paperclipai/paperclip/actions/runs/34626750213) passed on `c58e6c087e0df7530c747d80b27d491da925a9c4`: all 31 reported checks passed, including all server/workspace suites, browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. - Greptile reviewed this exact head at 5/5. All review threads are resolved, and security checks passed. - All 218 local targeted tests passed across explicit native continuation, continuation history, safe replacement, durable chat wakeups, wake queue, issue liveness, native session resume, and run dispatch. The native implementation is unchanged by the final rebase onto master. - Full workspace `pnpm -r typecheck` and `pnpm build` passed on the final head. Complete test coverage is supplied by the green CI suites; local tests used the targeted suites above. - Regressions cover scoped authorization, concurrent delivery, live ownership, later admission gates, retained message receipts, and automatic replacement after terminal-task or reviewer changes. ## Risks - A fresh model turn can choose to repeat an action. Paperclip preserves prior history and does not replay recorded calls. - Missing process identity and remote ownership without a target-aware stop proof retain the hold. A terminal database row alone does not prove that execution stopped. - No schema or dependency changes. Existing historical tasks are not awakened by deployment. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. This session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.11 |
||
|
|
42a4f5b15b |
feat(ui): advance single-choice questions on selection (#13234)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask structured questions in task threads and the composer. > - The shared form shows one question at a time. > - A single choice already completes an answer, but the form requires another click on Next. > - This pull request moves to the next question when the user selects one option. > - Multi-select and custom answers keep their Next step. The last page keeps explicit submission. ## Linked Issues or Issue Description **What existing behavior does this improve?** The paged question form in the task composer and task interaction cards. **Subsystem affected** ui/ — React board UI. **Current behavior** The user clicks a single choice, then clicks Next to reach the next question. **Proposed behavior** A single choice shows its checked state for 160 ms, then opens the next question with an 80 ms fade. Reduced-motion mode skips the animation. Multi-select stays on the current page until Next. Other stays open for typing. The last question waits for Submit answers. **Reason and benefit** Remove an extra click from each single-choice question while preserving explicit submission. **Breaking changes** Single-choice selection now changes the page. Answer payloads and APIs do not change. The user can return to earlier answers with the previous arrow. **Additional context** Related work: #12640 introduced the task workspace. No duplicate auto-advance change was found. ## What Changed - Confirm a single-choice answer with a brief radio animation and row highlight, then fade into the next page and focus the new question. - Use motion tokens for the 160 ms confirmation and 80 ms page fade. Honor reduced motion and cancel pending advances when the user changes direction or closes the form. - Lightly highlight every selected row with a foreground tint that remains visible against the composer in both themes, for single-select and multi-select answers. - Preserve multi-select, custom answers, back navigation, and final submission. - Ignore repeated number-key events and selection during an upload or submission. - Update composer and card tests. Add interactive and verified composer stories to the existing interaction Storybook group. - Document the Storybook scenario in the developer guide. ## Verification - Focused composer, card, and motion-catalog tests: 121 passed, including animation timing and cancellation. - `pnpm check:token-gates`: passed. - `pnpm build-storybook`: passed. Browser interaction story: passed with the selection animation enabled. - Selected-row refinement: 121 focused tests, token gates, and Storybook build passed. Browser inspection confirmed row highlighting for single-select and multiple selected checkboxes, plus light-theme contrast. - Manual browser check: select SQLite, select two features, click Next, select Now, then submit. The summary contains every answer. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - All 31 remote checks passed on the latest commit (`04a3f3983`), including typecheck, general and serialized tests, production build/native-runner verification, canary validation, and browser tests. Two optional Storybook deployment/visual jobs were skipped. Greptile reviewed the same commit at 5/5 with no actionable findings. The selected-row highlight refinement is included in that verification. - Animation refinement: UI typecheck, UI build, token gates, and 121 focused tests passed. - The broader UI run had 5,940 passing tests and three failures in unchanged Inbox/IssuesList tests. Both affected suites passed in isolation (73/73), including all three previously failing cases. - The broad local `pnpm test:run` was stopped after reproducing the native-runner failure below and after the full remote suite passed. It is not a clean local full-suite result. - Local runner limitation: `server/src/services/native-runtime/native-session-resume.test.ts` has one reproducible failure in unchanged code. The damaged-epoch recovery test expects `run.attach requires a settled Codex provider session`; the runner instead reports `semantic tool input content digest does not match its transmitted input` at line 1019. After building the missing fake provider with `pnpm --filter @paperclipai/paperclip-runner build:rust`, the isolated suite has 36 passing tests and this one failure. No server or runner files changed in this PR. ## Risks - Selecting a single choice changes the visible question after a brief checked-state confirmation. Back navigation preserves the choice so it can be edited. - The final page still requires Submit answers. Selecting Other still requires text and explicit progress. - No database, server, API, or dependency changes. ## Model Used - OpenAI GPT-6 in Codex, with reasoning, code editing, terminal tools, and browser verification. Exact serving revision and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb640ec129 |
fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path > - Paperclip manages AI agents and their work. > - Wake admission decides when a task can create an execution run. > - Recovery can prohibit replay while the previous execution needs review. > - Dependency reconciliation kept creating runs before dispatch rejected that same hold. > - Each rejected run added another startup notice without doing useful work. > - This change checks the hold during admission and records repeated automatic waits once. > - Tasks keep their messages and can resume when the current gates permit execution. ## Linked Issues or Issue Description Related changes: Refs #13173 (stale completed-task continuations). Refs #12651 (dependency waits during recovery). **What happened?** A blocked task with completed dependencies can remain under a durable execution reconciliation hold. Each scheduler pass created a queued run. Dispatch then cancelled it before the adapter started. The skipped wake did not satisfy dependency wake deduplication, so this repeated and filled the conversation with “Couldn't start” notices. **Expected behavior** A known execution hold creates a waiting diagnostic without a run. Repeated automatic observations share that diagnostic. Clearing the hold permits a new wake only after the other gates pass. New comments remain available for the next eligible execution. **Steps to reproduce** 1. Assign a blocked task with a completed blocker. 2. Give the task an active reconciliation action, or a resolved action whose automatic recovery evidence still prohibits replay. 3. Run dependency reconciliation repeatedly. 4. Observe repeated cancelled pre-start runs on the base branch. This branch creates no runs while held and admits work after the effective hold clears. ## What Changed - Check effective execution holds under the issue admission lock before inserting runs. Keep the final dispatch check for races. - Share automatic wait diagnostics across producers, wake keys, and service restarts. Apply the helper to reconciliation, dependencies, pause holds, availability, budgets, and disabled heartbeats. - Preserve ordinary comment and interaction receipts during execution holds. Prevent release from draining them while replay is blocked. Keep external-chat receipt authorization intact. - Group empty pre-start reconciliation cancellations into a neutral waiting notice. Keep started runs and the full run history. - Document the waiting contract and add database-backed, UI, and browser regressions. - Stabilize two existing verification tests: allow the asynchronous chat lease transition a bounded five-second wait, and accept either legitimate damaged-session refusal while retaining exact archive-evidence assertions. ## Verification Passed targeted tests: - `pnpm exec vitest run server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests. - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` — covered in the initial combined test run; UI suite passed. - Run-dispatch adapter tests passed in the combined gate regression run. - `pnpm exec vitest run server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including held receipt replay, promotion, and revoked access. - `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass. Repeated held messages create no additional runs or provider prompts and do not replay writes. - `pnpm check:token-gates` - `pnpm check:module-boundaries` - `git diff --check` `pnpm -r typecheck` and `pnpm build` pass. The full local `pnpm test:run` invocation did not finish green: it encountered exhausted local PostgreSQL shared-memory slots, a missing fresh-worktree runner test binary, and tests loaded across in-flight edits. The affected chat/database suites passed on rerun (81 tests), and the targeted lifecycle/recovery verification passed (3 tests). After building the runner test binary, the full native session suite also passed (37 tests). Final-head [CI run 34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475) passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all server/workspace test shards, all three browser shards, runner verification, typecheck, build, release dry run, and the aggregate verification gates. All 31 reported checks passed; the two conditional Storybook checks were skipped as intended. Greptile reviewed that exact commit at 5/5 with no unresolved review threads. ## Risks The wait record is diagnostic only. It must never count as a delivered wake or bypass a current gate. Tests cover repeated and concurrent admission, resolved no-replay evidence, a remaining dependency after hold clearance, deferred comments, and release gating. Explicit user requests and authorized chat receipts do not share automatic diagnostics. No migration or historical data deletion is required. Existing provider retry budgets remain unchanged. ## Model Used OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted snapshot ID or context-window size. Used repository inspection, code editing, command execution, tests, and review tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1d26ae965e |
fix(ui): keep active runner status current and say Working (#13238)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task transcript shows a running agent's progress.
> - The active-run query stops polling when the live-run list has data.
> - The transcript still preferred that initial snapshot, so an old
execution-confirmation state could remain after work resumed.
> - This pull request uses the refreshed snapshot for the same run and
keeps active status text at Working.
> - Operators can see current activity without connection-state jargon.
## Linked Issues or Issue Description
**What happened?**
The task transcript said Reconnecting while the runner continued sending
messages and calling tools. The stale projection could also hide the
Thinking tail or stop the status spinner and timer.
**Expected behavior**
The selected run uses its current live snapshot. Active transcripts say
Working and show current activity. Completed and failed runs say Worked
and Stopped.
**Steps to reproduce**
1. Open a running task before its execution confirmation arrives.
2. Let the active-run query stop polling when the live-run list returns
the run.
3. Let the list refresh to working while the cached active-run snapshot
still says reconnecting.
4. Inspect the transcript status and activity tail.
**Paperclip version or commit**
Base commit:
|
||
|
|
4fde92107e |
fix(ci): reuse cloud source verification for npm canaries (#13233)
Reuse the exact master source-verification result before npm canary publication, removing a duplicate verification matrix while preserving fail-closed release checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
52811c6ce6 |
fix(tasks): require resume before sending to paused tasks (#13232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task execution controls let board users pause a task or its subtree. > - The composer still accepted messages while a pause hold was active. > - A paused task must require an explicit resume before the user can send another message. > - This pull request replaces the composer with an amber pause card and checks board comment writes on the server. > - The user keeps their draft and resumes through the existing task controls. ## Linked Issues or Issue Description Refs #13104. Refs #13119. **What existing behavior does this improve?** The task composer and existing task/subtree pause controls. **Current behavior** A paused task can still receive a board message. The pause notice sits outside the composer, which leaves the send action available. **Proposed behavior** Show an amber takeover in both task chat and the classic composer. Preserve the draft. Require the user to resume the task or the ancestor subtree before sending. Reject board comment writes through either supported write route while the pause hold is active. **Breaking changes** Board comment writes to a paused task now return HTTP 409. Agent run reports remain supported during a pause. There is no schema migration. ## What Changed - Add a shared amber composer takeover with task, subtree, saved draft, pending, and error states. - Use effective ancestor pause state in both composer interfaces. Refresh it after pause events, task updates, and rejected sends. - Preserve draft text and attachments. Hide editor, send, queued edit, and pending question controls while paused. - Check active pause holds before board comment writes can mutate tasks, store comments, or wake agents. - Connect the approved Storybook examples to the production component and update the design and behavior docs. - Add browser coverage for both composers, draft persistence, resume, inherited holds, and rejected writes. Update ACP continuation coverage for the explicit resume requirement. ## Verification - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: `pnpm build-storybook`. - Passed: `pnpm check:token-gates` and `git diff --check`. - Passed: focused UI tests (398 tests) and server route tests (127 tests). - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/paused-composer.spec.ts tests/e2e/acp-stop-continuation.spec.ts` (5 tests). - Passed: manual browser walkthrough in a disposable local instance. Pause with a draft, refresh while paused, resume, send, and reopen. The draft returned, and one message persisted. The amber card and resume dialog were readable with no clipping. - Full local `pnpm test:run` did not pass: the general-server stage recorded 9,072 passing tests, 6 database setup failures from macOS shared-memory exhaustion, and 4 failed tests. This stopped the script before its later groups. Latest-head CI runs those groups independently. - Local follow-up: the Git file-resource load test passed on rerun (4 tests); native finalization migration passed after clearing the abandoned browser-test database allocation. Building the native debug fixtures fixed the missing fake provider. The remaining native-session recovery assertion also reproduces on untouched base commit `87b3e5fc6` (36 pass, 1 fail on both base and PR). It expects a settled-session error but receives a semantic-input-digest error. - The final UI build, UI typecheck, token gates, both thread suites (182 tests), and all five browser tests passed after the queued-action review fix. All 31 latest-head CI checks passed, including all server, workspace, browser, build, release, and security gates. Two optional Storybook jobs were skipped by workflow policy. Greptile reviewed `32d8fb5f5` at 5/5 with no open findings. - Review the Paused Composer and Tasks / Execution Controls stories. Pause a task with a draft, verify the amber card, resume, and verify the draft can be sent once. ## Risks - Clients that used board comments to continue paused work must resume first. The response is an explicit HTTP 409. - Pause state can change while a page is open. Live updates refresh the composer, and the server rejects stale sends before their side effects. - Resume keeps the existing dialog and optional agent wake behavior. Agent reports from interrupted runs remain allowed. ## Model Used OpenAI Codex, based on GPT-6, assisted with design, implementation, code execution, and browser verification. The exact runtime model ID and context window are not exposed in this session. The agent used reasoning and tool calls. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the relevant tests locally and they pass; the full local-suite limits are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.10 |
||
|
|
96bba78fba |
feat: add readable Storybook branch bookmarks (#13231)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Maintainers use Storybook previews to review the board UI. > - Branch previews need stable bookmarks that people can read. > - The current publisher only provides a hashed branch path. > - This pull request adds a readable branch bookmark after each successful upload. > - Existing branch and build links keep working. ## Linked Issues or Issue Description Refs #13226. **What existing behavior does this improve?** Manual Storybook publication for repository branches. **Current behavior** The stable branch path contains a hash. The expected `/storybook/branches/master/` URL does not exist. **Proposed behavior** Each publication updates a readable bookmark. The action summary and Markdown artifact link it. Master uses `/storybook/branches/master/`. Other names use a safe path segment that preserves case and escapes special characters. **Reason and benefit** Maintainers can save and share a readable URL that opens the latest published branch build. **Breaking changes** None. Existing hashed branch entries still update. Existing build URLs remain valid. **Additional context** This follows the publisher in #13226. A duplicate search found no related bookmark change. It does not overlap planned core work in ROADMAP.md. ## What Changed - Generate readable branch bookmarks without collisions with existing build directories. - Upload the bookmark only after the full build and compatibility entry uploads succeed. - Link the bookmark in the existing summary and Markdown artifact. - Document branch-name escaping and test path isolation, stable links, and upload order. ## Verification - `node --test scripts/__tests__/storybook-deploy.test.mjs`: 20 tests pass. - `actionlint .github/workflows/storybook-deploy.yml .github/workflows/storybook-visual.yml`: passes. - `git diff --check`: passes. - [Master bookmark publication](https://github.com/paperclipai/paperclip/actions/runs/34613344758): passed. Opened `/storybook/branches/master/` in the browser and confirmed a story renders. Downloaded the Markdown report and verified its bookmark link. - [Feature branch bookmark publication](https://github.com/paperclipai/paperclip/actions/runs/34613449034): passed. Its separate bookmark uses `codex~2Fstorybook-bookmarks`. - Greptile: 5/5 on `dccaf10413ecf447cb34e622b6b3c505791abb51`, with no unresolved review threads. All current-head Paperclip CI gates pass, including typecheck, tests, build, browser suites, and the canary dry run. - Full local repository checks were not repeated for this focused publisher change. The preceding run passed typecheck but encountered unrelated native-session test failures. ## Risks - Special characters in branch names use `~HH` byte escapes. For example, `feature/foo` becomes `feature~2Ffoo`. - Names that could overlap an existing hashed build directory escape the final hyphen. Very long names retain a hash suffix. - The two branch entries update separately. If the final upload fails, the workflow fails and a rerun can repair the bookmark. ## Model Used OpenAI GPT-6 via Codex, with reasoning, shell tools, and live deployment verification. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b03c4b9eb |
perf(ui): keep long task chat responsive during streaming (#13229)
## Thinking Path > - Paperclip lets operators manage agent work through tasks. > - Task chat keeps responses and run history together. > - A live update rendered every historical bubble and hidden tool row again. > - New image callbacks also forced unchanged markdown to parse again. > - Long conversations saturated the browser main thread. > - This change reuses unchanged history and mounts folded tools on first inspection. > - Operators can read and reply while work continues. ## Linked Issues or Issue Description **What happened?** Chat-style tasks with substantial scrollback became almost unusable. A deterministic browser reproduction with 200 long responses and 4,000 tools consumed 97.8% of the main thread during live updates. It delivered only 8 updates during the sample. **Expected behavior** The task should remain responsive during streaming. Historical markdown and unopened run details should not repeat expensive render work. **Steps to reproduce** 1. Install dependencies with `pnpm install`. 2. Run `pnpm exec playwright test --config tests/perf/task-chat/playwright.config.ts`. 3. Compare the attached performance JSON. The new test fails against the original rendering code. **Paperclip version or commit** Reproduced at `a05b828bc`. The branch is rebased on current master. **Deployment mode** Local Chromium and Vite with deterministic fixtures. No database or agent credentials are required. Related work: #10463 reduces the issue-page bundle. This change addresses repeated rendering after the page loads. No duplicate scrollback fix was found. ## What Changed - Keep the bubble image callback stable so unchanged markdown can skip parsing. - Memoize the settled history separately from the header and streaming tail. Keep the brief renderer and default attachment array stable. - Mount folded tool history on first expansion. Keep it mounted afterward to preserve child state and closing motion. Runtime request receipts remain visible. - Add deterministic rendering tests to the normal Vitest suite and an opt-in Chromium regression fixture. - Document the reproduction, commands, scope, and local measurements. ## Verification - Browser reproduction: 97.8% main-thread utilization before; 11.6% after for tail-only updates; 31.5% after when projection recreates history objects. Both fixed cases delivered 32 updates. - Browser checks pass for scroll-position retention, typing, return to latest, tool inspection, and retained expansion state. - Focused component suite: 155 tests passed. Post-rebase thread/performance rerun: 101 tests passed. - Recursive typecheck, build, Storybook build, and token gates passed. UI typecheck passed after the final edits. - Local UI/CLI lane: 5,910 tests passed; ten files hit worker-start timeouts, then all ten passed with two workers (20 tests). Shared/adapter lane: 3,128 tests passed. - Full local `pnpm test:run` was attempted and is **not green**: its server lane recorded 10,517 passes and 11 failures plus fixture/setup errors. Queue (31 tests), Cursor/Git-load (9 tests), and missing-binary failures cleared on isolated reruns / building the runner test binaries. Two native suites still cannot initialize embedded PostgreSQL on this host. - The remaining native-session recovery assertion was reproduced in a clean worktree at base `a20ecce40` (1 failed, 67 passed across the native/queue suites). It expects a settled-session error but receives a semantic-tool-input digest error. No server or runner files changed in this PR. - All substantive CI jobs have passed, including build, typecheck, all server/workspace test shards, all three e2e shards, and the canary dry run. The unchanged Slack ordering test exhausted its one-second wait on the first run; its shard passed on rerun. Final aggregate verification passed: **31 passing checks**, no failures or pending checks; two optional Storybook deployment/visual checks were skipped. Greptile is **5/5**, with no unresolved review threads. - The supplemental local serialized route run was stopped after CI passed all five serialized shards; it had reported no failures. - The browser fixture uses the real chat rendering components with a plain textarea. It does not test the full composer or server transport. Timing results are local samples; deterministic render-count tests provide normal CI coverage. ## Risks - Closed tool content becomes available to DOM search only after first expansion. Visible run summaries and runtime request receipts remain available immediately. - Opened run history remains mounted to preserve child state. The first full markdown render still scales with conversation size. - Memo dependencies must stay current when adding render inputs. Tests check content edits and replacement gallery callbacks. - No API, database, or permission changes. ## Model Used OpenAI GPT-6 (Codex). Exact serving variant and context-window size are not exposed in this session. Used reasoning, repository tools, code execution, and browser testing. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted/UI/workspace checks pass; the full local server suite has the baseline/host failures documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1616046c24 |
fix(runner): prevent trace scans from delaying live events (#13228)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner sends events that keep task state and steering
controls current.
> - Debug trace correlation reread and parsed the full trace for each
pending event.
> - Repeated scans blocked event delivery and left task state minutes
behind the provider.
> - This pull request indexes appended trace records once and reuses
those correlations.
> - Operators get current run state while native trace evidence remains
available.
## Linked Issues or Issue Description
**What happened?**
Native runs appeared live after their provider turn ended. Steering was
unavailable while the board still showed an active run. A CPU profile
attributed 93% of sampled time to trace lookup and file reads.
**Expected behavior**
Debug trace correlation must not delay run events or task state by
minutes.
**Steps to reproduce**
1. Enable native provider trace capture for a Codex run.
2. Produce a long event stream with pending correlations.
3. Compare provider event times with persisted run event times.
**Paperclip version**
Reproduced on source build
canary/v2026.911.0-canary.9
|
||
|
|
87b3e5fc61 |
fix(adapter-utils): stage selected skills into the sandbox for a remote Claude ACP run (#13196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A user selects skills for an agent, and the host materializes those skills into a bundle the agent reads > - An agent can run in a remote sandbox, where the host must stage every file the agent needs > - On the Agent Client Protocol lane the host built that bundle and then named its host path in the prompt, but it never staged the bundle into the sandbox > - The agent therefore read a path that does not exist inside the sandbox, and the run failed on the missing skill file > - The command-line lane of the same adapter already stages a `skills` asset and reads the in-sandbox directory back from the staged runtime > - This pull request carries that proven pattern to the Agent Client Protocol lane, so a selected skill reaches the agent in a remote run ## Linked Issues or Issue Description No public issue exists for this change. The description follows. **What happened?** A remote run of the Claude adapter on the Agent Client Protocol lane could not read any selected skill. The host builds the skill bundle in its own state directory, then writes that host path into the prompt as `Skill root: <path>`. The remote seam of that lane staged one asset only, the configuration seed. It staged no skills asset, so no skill file crossed into the sandbox. The agent then tried to read the skill file at the host path, and the read failed with a missing-file error. **Expected behavior** A remote run receives the skills the user selected, and the prompt names the directory that holds those skills inside the sandbox. **Steps to reproduce** 1. Select one or more skills for an agent that uses the Claude adapter. 2. Start a run for that agent in a remote sandbox on the Agent Client Protocol lane. 3. Ask the agent to read the skill file at the path the prompt names. The file is not there. **Agent adapter(s) involved** The Claude local adapter, on its Agent Client Protocol lane. The shared engine in `packages/adapter-utils` carries the prompt rewrite. **Additional context** The command-line lane of the same adapter already stages a `skills` asset and remaps onto the staged directory. This change reuses that mechanism instead of adding a new transport. One other adapter shows the same host-path shape on its own Agent Client Protocol lane. That lane is tracked separately and this pull request does not change it. ## What Changed - Return the host skill bundle directory from the Claude skill runtime step, and carry it through the remote managed-home context to the staging seam. The value is null for a non-Claude agent, for a run that selects no skill, and for a run whose selected skills all fail to materialize. - Stage that bundle as a `skills` asset on the Claude Agent Client Protocol remote seam, and only when the run selected a skill. - **Stage that asset with `followSymlinks: false`.** The bundle holds an owned copy of each selected skill, and the copy step never copies a symbolic link at the root or at any depth. So the bundle contains no symbolic link, and staging has none to follow. Refusing to follow one also stops a link planted in the bundle directory after the copy from pulling an unrelated host file into the sandbox. A regression test walks the real adapter sources and pins the reviewed `followSymlinks` value at every skills staging site, so a new or changed site fails the test. - **Drop a skill whose staged copy has no usable `SKILL.md`** from the prompt, the skill identity, the command notes, and the bundle, and log which skill was dropped and why. Without this, a skill whose copy failed, or whose `SKILL.md` is a symbolic link the copy step skips, stayed advertised in the prompt while its file was absent — the same missing-file symptom this change exists to fix. - Rewrite the `Skill root:` prompt line, the skill identity, and the command notes onto the in-sandbox directory. The rewrite runs in the engine, after the workspace placement returns the staged runtime. A compatible session resume reuses the cached staged runtime, so the rewrite runs on that path too. - Keep the session fingerprint on the host-independent skill identity. A change to the selected skill set still invalidates a warm session, and the volatile sandbox path stays out of the hash. - A local run, and a run with no selected skill, keep their current behaviour. ## Verification - `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` — 28 of 28 pass. - `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` — the new engine tests and the staging-site tests pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`, and the same check on the adapter package — both exit 0. - The end-to-end test drives the lane against a local sandbox stand-in. It reads the skill root out of the prompt the runtime received, and then opens the skill file at that path. That is the reported symptom, proved closed. - The new tests carry a sensitivity control. Restoring only the production files to their previous content fails 7 of the 9 new tests. The other 2 do not depend on production code: one is a parser unit test for the source scanner. ## Risks Low risk, and the change is a two-way door. A revert restores the previous behaviour exactly. - **Scope.** The change touches one adapter lane. It does not change the local lane, and it does not change any other adapter. No existing staging site changes its `followSymlinks` value. - **The staged bundle and the workspace.** The staged skills land under the runtime directory inside the workspace. The workspace restore excludes that whole runtime directory, so the staged skills never return to the host worktree. A test proves the exclusion end to end. - **Session reuse.** The rewritten path never enters the session fingerprint, so it cannot invalidate a warm session, and a compatible resume applies the same staged path. - **Direction of data.** Files move from the host into the sandbox only. The change adds no path that writes sandbox content onto the host. - **A dropped skill.** A skill with no usable `SKILL.md` is now absent from the prompt instead of named but unreadable. The run logs the skill and the reason, so the cause is visible. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking and tool use, through Paperclip agents. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` passes — 28 tests. - [x] `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` passes. - [x] `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - [x] All continuous-integration gates are green. - [x] The automated review reports no open finding against the current head. ## Required Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant bug report template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect this change - [x] I have considered and documented the risks above - [x] All continuous-integration gates are green - [x] The automated review score is 5/5 with no open current-head findings - [x] I have addressed every reviewer comment that applies to the current head **Note on the branch history.** This branch first carried a different change: a filename-based admission filter that refused to stage files such as `.env` from a skill directory, together with a switch from symbolic-link bundles to copied bundles. That approach was rejected and **reverted** on this branch. It does not match the documented trust boundary, because the host already delivers credentials into the sandbox on purpose, and replacing the symbolic-link bundles broke live editing of a skill. The revert is in this branch's history. The file that work changed, `packages/adapter-utils/src/server-utils.ts`, is byte-for-byte identical to `master` here and is not part of this diff. Earlier review findings that name that file target the reverted code. All of them are resolved, and the automated review passes on the current head. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
974949a39b |
ci: spread cloud server verification across ten runners (#13227)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud waits for source verification before deploying a new image. > - The slowest server verification job spends about ten minutes running tests. > - Each job uses one test worker to preserve test isolation. > - This pull request distributes those suites across ten standard hosted runners. > - The benefit is a shorter verification path with the same test coverage. ## Linked Issues or Issue Description **Current behavior** In [readiness run 34572340764](https://github.com/paperclipai/paperclip/actions/runs/34572340764), the slowest server job ran for 638 seconds. Test execution used 594 seconds. This held readiness behind the image job. **Proposed behavior** Use ten general server jobs in the reusable release verification workflow. Keep the three chat jobs and every existing prerequisite. The complete partition test verifies that no server suite is omitted or duplicated. **Reason and benefit** Reduce merge-to-deployable time on the existing runner type. The next longest prerequisite was Runner verification at 526 seconds, so the initial expected total gain is about two minutes rather than a halving of readiness time. Measure actual queue and execution time before claiming a result. Related: #13198 introduced the separate chat lane. #12577 refreshes duration estimates; this change leaves that manifest alone. ## What Changed - Increase the general server matrix from five jobs to ten. - Verify the ten-way partition covers the complete server suite when combined with the chat lane. - Document runner demand and the unchanged local and PR grouping. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/__tests__/run-vitest-stable-shard.test.mjs`: 29 passed. - `actionlint .github/workflows/release-verify.yml`: passed. - Full local `pnpm -r typecheck` and `pnpm build`: passed. - All latest-head GitHub CI checks passed, including the complete Linux test partition, build, typecheck, and browser gates. Greptile: 5/5 with zero open findings. - [Ten-shard timing probe](https://github.com/paperclipai/paperclip/actions/runs/34606772388): all 16 jobs passed; slowest server job 6m 23s versus 10m 38s in the earlier five-shard sample. This compares the server lane, not total readiness, and is not a controlled same-source A/B. - The full local `pnpm test:run` is also running. It has reproduced previously observed macOS-only failures in unchanged skill-cache and native-session suites; the corresponding Linux CI suites passed. Final local results will be attached separately. No affected-workflow test failed. ## Risks Five additional concurrent jobs per release verification run increase runner demand and repeated setup work. Queueing can offset the gain. Test workers, timeouts, permissions, and readiness requirements stay unchanged. Revert the matrix and its partition test to restore the previous split. ## Model Used OpenAI GPT-6 / Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected workflow tests locally and they pass; full-suite macOS limitations are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20ecce409 |
feat: publish CODEOWNER-approved Storybook branch previews (#13226)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Maintainers use Storybook to review the board UI. > - Reviews need public previews of selected repository branches. > - Each branch needs its own URL so previews do not replace each other. > - This pull request adds manual, CODEOWNER-controlled publishing to S3 and CloudFront. > - The action returns stable branch links and permanent build links in its summary and a Markdown artifact. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Storybook build and manual visual-review workflow. **Current behavior** The repository has no manual branch-preview publisher. A single GitHub Pages site cannot support independent publishers without combining their output. **Proposed behavior** A CODEOWNER selects a source branch and approves publication. Each branch has a stable CloudFront URL. A completed build becomes the branch target only after its upload succeeds. The action attaches `storybook-deployment.md` with the preview links and source commit. **Reason and benefit** Maintainers can share multiple branch previews at the same time. Branch builds have no repository token permissions or AWS credentials. Dependency caching and install hooks are disabled. The publisher cannot write runner dashboard files or delete objects. **Breaking changes** None. Normal visual checks keep their existing behavior. This does not change application code or GitHub Pages settings. **Additional context** Searched public issues and PRs for Storybook deployment work. No duplicate deployment proposal was found. This is maintainer infrastructure, not a roadmap-level core feature. ## What Changed - Add `Storybook Deploy` with a source-branch input and a manual entry through `Storybook Visual`. - Check the original actor and rerunner against default-branch CODEOWNERS. Require a protected deployment environment with CODEOWNER reviewers. - Separate public-source builds with no repository permissions from an OIDC publisher restricted to the Storybook S3 prefix. - Publish distinct branch URLs and retain build URLs. Preserve Storybook deep links across the branch redirect. - Add the run summary, a downloadable Markdown deployment report, focused tests, and operator setup docs and IAM policies. ## Verification - `node --test scripts/__tests__/storybook-deploy.test.mjs`: 19 tests pass. - `actionlint .github/workflows/storybook-deploy.yml .github/workflows/storybook-visual.yml`: passes. - [Feature branch live publication and deployment-only rerun](https://github.com/paperclipai/paperclip/actions/runs/34533202273): passed. - [Master branch live publication](https://github.com/paperclipai/paperclip/actions/runs/34533204743): passed. - Both public branch URLs render a component story without browser errors. A deployment-only rerun updates only the selected branch entry and preserves the previous build URL. - AWS policy simulation allows Storybook uploads and denies dashboard writes and object deletion. - Full local typechecking passes. Full local tests, build, and current-head PR checks are running. - [Revised build and Markdown artifact validation](https://github.com/paperclipai/paperclip/actions/runs/34605623088): passed. Downloaded the report and verified its branch URL, build URL, and source commit. - The public verifier also checks that the stable branch URL points to this build and rejects stale targets. ## Risks - Storybook previews are public. Maintainers must publish only public UI fixtures. - Retained builds accumulate until an operator prunes them. - Environment reviewers must stay synchronized with CODEOWNERS. The workflow fails closed if its environment loses required protection. - The existing CloudFront distribution is shared with runner reports. Separate S3 prefixes and a dedicated role prevent the publisher from overwriting those reports. ## Model Used OpenAI GPT-6 via Codex, with reasoning, shell tools, and browser verification. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d10cbde815 |
fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path > - Paperclip manages agent work through tasks and runs. > - Recovery continues assigned work when no live execution path remains. > - A recovery sweep can read an in-progress task before its run completes. > - The sweep can then observe the successful run after completion has changed the task status. > - This pull request checks current status and assignment under the existing enqueue lock. > - A stale continuation leaves a skipped wake receipt and creates no run. > - Task chat also omits an empty continuation cancelled before it started because its task had become terminal. ## Linked Issues or Issue Description Related public work: #10779 and #8419. Those older open changes address terminal disposition across other recovery paths. This change uses the existing scheduler guard for productive successful-run continuation and adds real database lock contention coverage. **What happened?** Recovery could combine an old in-progress task snapshot with a newer successful run. It queued an automatic continuation after the task was done. Dispatch cancelled that run before it started, but task chat displayed “Couldn't start” below the successful answer. This can happen after the native runner's finish result has already been accepted. It does not require a missing comment. **Expected behavior** Productive continuation must remain eligible when enqueueing acquires the task lock. Completion, cancellation, reassignment, or a move away from in-progress must prevent creation of the run. Actual execution stops must remain visible. **Steps to reproduce** 1. Let recovery select an assigned in-progress task whose latest run succeeded with productive progress. 2. Hold the task row lock in another transaction and change the task to done. 3. Let recovery attempt to enqueue while that transaction holds the lock. 4. Commit completion. Before this fix, recovery creates a redundant run from the stale snapshot. **Paperclip version or commit** Reproduced against master at `4042eb1c4` with deterministic integration tests. **Deployment mode** Built from source with PostgreSQL. The bug is in core recovery and is not adapter-specific. ## What Changed - Pass the existing status-and-assignee guard for productive terminal continuation recovery. - Preserve a skipped wake receipt with the expected and actual task state, without creating a run. - Test actual PostgreSQL lock contention for native and legacy completion, cancellation, backlog, review, blocked state, and reassignment. - Omit empty redundant pre-start cancellations from native and legacy task chat. Preserve stop markers for runs that started. - Document recovery eligibility at enqueue time. ## Verification - All seven new race cases failed before the guard was connected. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Task chat suite: 98 tests passed. - Recovery integration suites: 290 tests passed, including 31 stale-queue tests. - Full CI verification passed on `7ea71f04d`: all 31 active checks succeeded, including all test shards, browser tests, build, typecheck, and canary release dry run. Storybook visual regression was skipped by its path filter. - Greptile reviewed this commit at 5/5 with no review threads. - The local `pnpm test:run` aggregate reported a setup failure in the unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed all 231 tests without edits. The duplicate aggregate was stopped after the complete CI matrix passed; it is not counted as a successful local full-suite run. ## Risks Low risk. The backend guard applies only to productive successful-run recovery. It requires the task to remain in-progress with the same agent. Other wake sources keep their current policy. The UI change only suppresses empty redundant cancellations; run records remain available. No schema change or migration is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code edits, and local test execution. The exact deployment snapshot and context window are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a05b828bcd |
Reduce run polling and workspace inspection amplification (#13174)
## Thinking Path > - Paperclip manages agent work and shows run progress to operators. > - Run lists, live events, transcripts, and workspace details must remain responsive as usage grows. > - Run-list redaction rereads the full context for every run. Hidden tabs can still trigger requests through live events and manual timers. > - Workspace detail reads repeat Git inspection even when concurrent callers request the same state. > - This pull request batches registry reads, pauses hidden-tab refreshes, and caches Git inspection for display. > - Cleanup keeps fresh Git checks, and redaction keeps company and run boundaries. ## Linked Issues **What happened?** Run-list responses perform one extra database read per run and parse full context JSON to obtain small secret registries. Hidden tabs continue transcript reads and event-triggered refetches. Workspace detail requests repeat Git scans. **Expected behavior** A run list reads registries once. Hidden tabs stop recurring run reads and reconcile when visible. Concurrent workspace detail reads share a short-lived Git result. **Steps to reproduce** 1. Open run lists and task transcripts in several tabs while agents run. 2. Hide some tabs and observe transcript and event-triggered requests. 3. Request a 200-run list and count redaction database queries. 4. Request the same workspace detail concurrently and count Git inspections. Related: #5255 adjusts polling cadence. This change addresses hidden-tab lifecycle, batched registry reads, and workspace inspection reuse. No duplicate with this scope was found. ## What Changed - Batch heartbeat and live-run redaction into one company-scoped registry query. Select only registry JSON for run and issue redaction. - Resolve duplicate secret values once per request. Preserve each run's registry and remove registry material from responses. - Suspend company event sockets and transcript reads while hidden. Refresh active queries and resume transcript offsets on return. - Prevent queued event invalidations and developer health polling from fetching in hidden tabs. Gate legacy run-log readers in both UI variants. - Exclude legacy plugin placeholder connections from remote health probes. Select only due connection IDs in SQL before the sweep limit. Preserve existing plugin records. - Cache concurrent Git display inspections for five seconds, with at most 256 entries. Leave close-readiness and cleanup checks uncached. - Add regression coverage and document the performance behavior. - Stabilize the existing Rust descendant-lineage fixture: allow a bounded 30 seconds for 300 durable notifications under concurrent test load, retaining every correctness assertion and adding timeout diagnostics. ## Verification - Regression coverage verifies one registry query for 200 runs, per-run isolation, request-local secret resolution, decryption failures, Git cache expiry/bounds, hidden-tab pause, and visibility recovery. - Real PostgreSQL redaction/run-route suites passed all 57 tests; workspace-service coverage passed. The health-sweep regression verifies plugin placeholders and chat connections remain untouched and do not consume the sweep limit. - Both legacy transcript viewers retain history and resume their byte offset after visibility changes. The related visibility/progress/chunk suites passed all 29 tests. Other focused UI suites and token gates passed. - Full `pnpm -r typecheck` and `pnpm build` passed. Affected-package typechecks/builds passed after review fixes. The concurrent Rust provider suite passed 84 tests (two ignored), and Rust formatting passed. - Full local `pnpm test:run` stopped after the general-server group: 10,538 passed, 65 skipped, four failed. Fresh chat-delivery and health-sweep reruns passed; building the debug runner fixture cleared the native-event test. One unchanged native-session recovery assertion still fails locally with a semantic-digest error instead of the expected settled-session message. The full local command is therefore not green. CI runs the later groups separately and skips the two native-session tests requiring a prebuilt runner binary (confirmed in its 37-test native-session suite). - All CI gates pass on final head `ee610e737`: typechecking, general and serialized tests, browser tests, runner verification, build, and canary dry run. One server shard passed on its single retry after exposure fixtures encountered port 42001 where they assumed 42000; that suite also passed locally (25 passed, three platform-specific skips). - Greptile reviewed the final head at 5/5 with no actionable findings. ## Risks - Workspace delivery display can lag local Git changes by five seconds. Destructive operations still inspect current state. - Hidden tabs do not receive company live-event notifications until visible. Active queries refresh on return. - This change preserves legacy plugin records and does not repair instance-specific workspace rows. There is no database migration. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser inspection. The exact model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted regressions; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.8 |
||
|
|
932c8bec56 |
fix(ci): bake the managed runtime identity into cloud images (#13210)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed deployments start from the image built by the Cloud workflow. > - The managed runtime requests user and group 1001. > - The image currently builds the node user as 1000. > - Startup must remap that user, which can walk a large mounted home directory. > - This pull request uses the existing Docker build arguments to bake user and group 1001 into Cloud images. > - Matching the runtime identity removes that startup work and helps avoid health-check retries. ## Linked Issues or Issue Description Refs #13208, #1923, and #7861. Searched open and closed PRs for the Cloud UID change. The older #7861 addresses build context and volume ownership repair. This change uses the existing identity arguments in the Cloud workflow and preserves ownership repair. **What happened?** A measured rollout had a container log `Updating node UID to 1001` after startup. The container stayed at this step for at least 2 minutes 55 seconds before rollback stopped it. The baked node identity was 1000, while the managed runtime requested 1001. A health check timed out and the target required a second deployment attempt. **Expected behavior** Cloud images should already have the managed runtime identity. A matching image should skip user and group remapping. Fresh or mismatched volumes must still receive ownership repair. **Steps to reproduce** 1. Build the current Cloud image with its default build arguments. 2. Start it with `USER_UID=1001`, `USER_GID=1001`, and a populated home volume. 3. Observe the startup user remap before the application starts. **Paperclip version or commit** `fc06f7f05f42c675be71ff0927b6334405d520ed` **Deployment mode** Docker on managed hosts. ## What Changed - Pass `USER_UID=1001` and `USER_GID=1001` to the Cloud image build. - Check the pushed digest's baked identity before the entrypoint can repair it. Then check the normal entrypoint's effective identity and writable home before publishing the verified full-SHA tag. - Add a workflow regression and two entrypoint cases for a matching Cloud identity, including a mismatched volume. - Document the runtime identity and the first-build cache cost. ## Verification - Focused workflow and artifact tests: 27 passed. - Entrypoint tests: 11 passed. Actionlint passed. Full local `pnpm -r typecheck` passed. Full local `pnpm build` passed. The manual [Cloud image build](https://github.com/paperclipai/paperclip/actions/runs/34575473213) passed on the exact PR head. It checked Sentry, baked and effective identity, writable home, orphan reaping, and full-SHA publication. The new identity check took one second. All 30 PR checks passed; the Storybook workflow was intentionally skipped. Greptile reviewed commit `114d408f637a0b53e2e2b1339c263779b1e4ae54` at 5/5 with no findings or open threads. - The full local suite for the same application source was already run in #13205. Its macOS general-server phase had 10,471 passes and 70 failures in seven unchanged files. Those failures included missing Runner fixtures, filesystem errors, timeouts, a port conflict, and a load-count mismatch. After configuring Cargo and rebuilding fixtures, 37 of 38 native tests passed; one unchanged native-resume assertion still failed. Linux PR CI passed. This change adds entrypoint tests and does not change application code. ## Risks - The first build must rebuild layers that depend on the base image identity. Later builds can reuse them. - A future managed runtime identity change must update these build arguments and checks together. - The Dockerfile's self-hosted defaults remain 1000. Runtime overrides and mounted-volume ownership repair remain supported. - The observed startup delay supports this change, but fleet timing also includes provider startup, image pull, canary order, and retries. No fixed end-to-end gain is claimed before a live rollout. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and tool use. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused workflow tests; full-suite limitations are listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.7 |
||
|
|
fc06f7f05f |
fix(ci): isolate chaos verification by caller workflow (#13208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments require verified artifacts for the merged source commit. > - Cloud readiness and the npm release independently run the same source checks. > - Their shared chaos workflow used only the source ref as its concurrency key. > - One caller could cancel the other caller's required job for the same commit. > - This pull request scopes that key to the caller workflow and source ref. > - Both callers can finish their checks without blocking deployment readiness. ## Linked Issues or Issue Description Refs #13192 and #13205. Searched for related open issues and PRs; no duplicate fix was found. **What happened?** The master push for `398d304e15739d1ee6105633bd8a0e42c929d33f` started Cloud readiness and Release together. GitHub cancelled the Cloud readiness chaos job before it acquired a runner. Its annotation reported a higher-priority waiting request for the same concurrency group. The required readiness gate cannot pass after that cancellation. **Expected behavior** Cloud readiness and Release must each finish source verification for the same SHA. Standalone chaos evals must also have a separate group. **Steps to reproduce** Merge a commit to master while the npm release queue is empty. Both callers reach the reusable chaos workflow with the same source SHA. See [the cancelled job](https://github.com/paperclipai/paperclip/actions/runs/34569569760/job/103168603926). **Paperclip version or commit** `398d304e15739d1ee6105633bd8a0e42c929d33f`. **Deployment mode** GitHub Actions on master. ## What Changed - Add the caller workflow name to the chaos workflow concurrency group. Retain source isolation and cancellation of duplicate calls within the same workflow. - Add a regression test that evaluates the group for Cloud readiness, Release, and standalone evals at the same source SHA. - Document the concurrency boundary in the readiness runbook. ## Verification - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 26 tests. - The new regression test fails against the previous concurrency key and passes with this fix. - `actionlint -shellcheck= -pyflakes= .github/workflows/runner-chaos-evals.yml .github/workflows/release-verify.yml .github/workflows/cloud-readiness.yml` passed. - `git diff --check` passed. - The full local typecheck passed for the same application source in #13205. Its macOS general-server test phase had 10,471 passes and 70 failures in seven unchanged application test files: missing Cargo/Runner test binaries, filesystem permissions, timeouts, a port conflict, and a load-test count mismatch. Linux CI test checks passed. The full local build passed with Cargo on PATH. This PR changes workflow configuration, its test, and documentation only. - All CI checks pass on the final head, including typecheck, tests, browser suites, build, and canary dry run. Greptile is 5/5 with no open findings. After merge, verify both callers' chaos jobs complete for the same master SHA and record the resulting readiness time. ## Risks - Two callers may now run chaos tests at the same time. This uses two existing GitHub runners, which is the intended cost of independent verification. - Renaming a caller changes its concurrency group. The fixed prefix keeps this child group separate from caller-level concurrency groups. - The readiness gate continues to require every verification prerequisite. No gate is bypassed. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (26 focused workflow/artifact tests) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.6 nightly/v2026.911.0-nightly.0 |
||
|
|
398d304e15 |
docs: measure cloud deployment through target health (#13205)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Hosted deployments need a verified image and migrator for the same source commit. > - The cloud readiness workflow certifies those inputs before a deployment consumer acts. > - Its completion time does not show when a tenant runs the new commit. > - This pull request documents each milestone from merge through target health and fleet completion. > - Operators can use the evidence to find the slow stage and measure a complete deployment. ## Linked Issues or Issue Description Refs #13192, #13188, and #13189. Searched related issues and PRs; no duplicate timing documentation change was found. **Issue type** Missing documentation. **Where is the issue?** `doc/cloud-build-readiness.md`, Timing and rollout. **What's wrong?** The timing instructions stop at the readiness job. That omits consumer queues, artifact resolution, and target deployment. An image can be ready while the tenant still runs an older commit. **Suggested fix** Record separate merge, image, readiness, canary health, and fleet completion timestamps for the same full source SHA. Keep preparation-only runs out of deployment results. ## What Changed - Define the evidence needed for each merge-to-deployment milestone. - Explain how consumer queues can hide upstream build gains. - Require target source identity as well as health, and report exclusions, retries, cache state, and queue conditions. ## Verification - `git diff --check` passed. - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 25 tests. - Cross-checked the readiness identity and artifact prerequisites against the current workflows and consumer contract. - Full local `pnpm -r typecheck` passed using the session's installed Rust toolchain. The full local test suite and subsequent build are still running. - All CI checks pass and Greptile is 5/5 on the exact head, with no unresolved findings. This changes one documentation file and adds no runtime behavior. ## Risks - Low risk: documentation only. Timing must still use trusted run evidence and the actual target commit. A single measured run is not a latency guarantee. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (25 focused workflow/artifact tests; full checks pending) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.5 |
||
|
|
d56be3f3fc |
fix(ci): verify deployable cloud artifacts independently (#13192)
Verify source, build the cloud image, and wait for exact-source migrator packages concurrently. Emit Cloud deployable v1 only when every prerequisite succeeds for the merged full SHA. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5cc51fad06 |
fix(release): publish exact-source cloud migrators on merge (#13188)
Publish exact-source shared and database migrator packages for each master merge through the existing trusted Release workflow, independently of the full release and image build. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6a7025ebe3 |
ci: skip cloud runner cleanup when disk headroom is ample (#13191)
Skip cloud runner disk cleanup when both the Docker and workspace filesystems have at least 64 GiB free. Preserve the existing cleanup for low, unavailable, or invalid measurements and verify the actual shell behavior across eight scenarios. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5c660a32f3 |
ci: build cloud images independently for each merge (#13189)
Build cloud images independently for each master commit through a reusable workflow. Preserve production release dependencies and image runtime checks, and write cloud registry caches per commit with bounded ancestor imports to prevent overlapping builds from replacing each other's cache. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
59d74b68b2 |
ci: cache the native Runner in a separate Docker stage (#13195)
Compile the native Runner from its complete Cargo and protocol inputs in a separate cached Docker stage. Preserve Cargo validation and generated-contract checks during the normal application build, normalize input timestamps across checkouts, and compile the isolated target in PR CI. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
42961b6ef1 |
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e970df4c5 |
ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push. Co-Authored-By: Paperclip <noreply@paperclip.ing> |