mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 21:03:51 +02:00
a12bbd18249b676aa2cff055c2b6ef4e3cbb5ebe
4337
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a12bbd1824 |
fix: stop completion reviews caused by policy upgrades (#13266)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runtime records completion assessments and task decisions. > - An application update can change the assessment policy version. > - The previous code treated that version change as a reason for human review. > - An upgrade alone does not give the user a new decision to make. > - This change keeps existing decisions and withdraws obsolete upgrade review cards. ## Linked Issues or Issue Description **What happened?** A rules version change moved unfinished tasks into review and created a card that said, "Review the superseding native policy assessment." The task did not need new work or a human decision. **Expected behavior** New runs use the current rules. An upgrade leaves existing task decisions alone. The saved policy version remains available in the audit history. **Steps to reproduce** 1. Save a native run assessment and a task decision. 2. Change the native status policy version. 3. Run finalization reconciliation without new task evidence. 4. The previous code created a review card. The corrected code keeps the saved assessment and decision. Related context: #13038 changed the native policy version. Searches found no duplicate fix for upgrade-only completion reviews. ## What Changed - Remove policy-version mismatch as a reconciliation trigger. - Withdraw pending cards only when their source decision, effect ledger, creator, key, and prompt match the old upgrade-only review. - Restore the previous status only while that decision and status version remain current and no other review gate is pending. - Preserve answered cards, real review requests, later task changes, and historical assessments and decisions. - Record cleanup activity and retire any corresponding chat review actions. - Isolate cleanup failures so one old card cannot block other cleanup or normal finalization. - Update the status conformance fixture, regression tests, and architecture documentation. ## Verification - Passed: `pnpm exec vitest run server/src/__tests__/native-status-arbiter-corpus.test.ts` (23 tests, including the 53-fixture status corpus). - Passed: `pnpm -r typecheck`. - Passed: `git diff --check`. - Passed: `pnpm build`. - Passed: fresh Greptile review at 5/5 on `d24e5ecaf`, with no open review threads. - Passed: the 23 focused tests in the isolated full-suite environment. - Local `pnpm test:run`: the server group finished with 10,638 passed and one missing-fixture failure. Built the required `fake-codex-app-server` fixture and reran the entire affected native session-resume suite: 37/37 passed. The initial full command exited on that server-group failure, so remaining groups are covered by CI. - Passed: the workspace-runtime-exposure suite (25 tests, 3 platform skips). - Passed: all CI gates, including every test shard, browser tests, typecheck, and build. The server shard passed on one retry after an unrelated host-port conflict in the unchanged workspace-exposure tests. - Cleanup tests cover repeat runs, real reviews, answered cards, later statuses, newer decision identity, status changes back to review, other pending requests, run scope, cleanup failures, retry, and live publication failures. ## Risks - Cleanup changes stored task state. It checks the exact obsolete decision and status version under database locks before restoring status. - Pending approvals, interactions, and execution stages prevent restoration out of review. - The cleanup handles at most 100 matching cards per reconciliation pass. It does not rewrite old decisions or accept an agent's completion claim. - New evidence and explicit task changes still use the existing reconciliation paths. This change does not reevaluate old work merely because Paperclip was updated. ## Model Used OpenAI Codex, GPT-6. The exact deployment identifier and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.15 |
||
|
|
19c76bfc3f |
fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - CI installs dependencies before it verifies and builds cloud artifacts. > - Install jobs share a pnpm package-store cache with lockfile refresh. > - Lockfile refresh resolves versions without downloading packages. > - That job saved an empty cache before full install jobs could save theirs. > - This PR prevents lockfile refresh from publishing that empty entry. > - Full install jobs can then populate the cache and reuse dependencies. ## Linked Issues or Issue Description Refs #13259 for the related cloud verification cache work. No duplicate empty-cache fix was found. **What happened?** Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm cache at 18:58:08 UTC on September 10. Full install jobs still restore that empty entry. The cache API reports 216 bytes for master and about 703 MB for populated entries with the same key and cache version in PR scopes. **Expected behavior** A job that installs dependencies should populate the shared package-store cache. **Steps to reproduce** 1. Run lockfile refresh with a new lockfile cache key. 2. Its resolution-only command leaves the package store empty. 3. The Node action saves the empty archive before a full install finishes. 4. Later jobs report a cache hit but download packages again. **Paperclip version or commit** Observed on master |
||
|
|
778ac33308 |
feat: accept a base64-encoded Cloud UI snippet (#13245)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip Cloud instances can inject a trusted operator HTML snippet before `</body>` in served UI pages (#13168) > - Operators deliver env vars to managed instances through hosting-provider APIs > - Web application firewalls in front of those APIs reject request bodies that contain raw `<script` markup > - The snippet value always contains a script tag, so the firewall blocks its delivery every time > - This pull request adds `PAPERCLIP_CLOUD_UI_SNIPPET_B64`, which carries the same snippet as standard base64 > - The benefit is that operators can ship the snippet through WAF-fronted delivery pipelines without firewall exceptions ## Linked Issues or Issue Description Refs #13168 (introduced the Cloud UI snippet). **Subsystem affected** Server UI shell serving: `server/src/cloud-ui-snippet.ts`, used by `static-index-html.ts` and `app.ts`. **Current behavior** `PAPERCLIP_CLOUD_UI_SNIPPET` is the only way to configure the Cloud UI snippet. The value is raw HTML. Delivery pipelines that write env vars through provider APIs can fail to deliver it: web application firewalls classify a request body that contains `<script` as an injection attempt and block it. Verified against Railway's GraphQL API, which sits behind Cloudflare: a variable value with a bare `<script></script>` returns HTTP 403 before authentication, and JSON unicode escapes (`<script`) do not bypass the block. The snippet is exactly the kind of value that always contains a script tag, so such pipelines cannot deliver it at all. **Proposed behavior** A new optional `PAPERCLIP_CLOUD_UI_SNIPPET_B64` carries the same snippet as standard base64 of the UTF-8 HTML. The server decodes it and injects the result through the existing path. The plain variable wins when both are set. Whitespace and line wrapping in the value are tolerated. A value that is not canonical base64, or that decodes to blank, is ignored instead of injected as garbage. **Reason and benefit** Operators can deliver the snippet through WAF-fronted APIs without requesting firewall exceptions. Base64 contains no markup, so the firewall has nothing to match. Existing deployments see no behavior change. **Breaking changes** None. The new variable is optional. The existing variable is unchanged and takes precedence. ## What Changed - `server/src/cloud-ui-snippet.ts`: resolve the snippet from `PAPERCLIP_CLOUD_UI_SNIPPET` first, then from base64-decoded `PAPERCLIP_CLOUD_UI_SNIPPET_B64`; ignore non-canonical or blank-decoding values. - `server/src/__tests__/cloud-ui-snippet.test.ts`: cover decode-and-inject, plain-wins precedence, whitespace tolerance, invalid/blank values, and the self-hosted no-op. - `doc/cloud-ui-snippet.md`: document the variant, how to produce the value, and the ignore rules. ## Verification - `pnpm vitest run server/src/__tests__/cloud-ui-snippet.test.ts server/src/__tests__/static-index-html.test.ts` — 13 tests pass. - `tsc --noEmit` in `server/` passes. - Manual check: on a Cloud-managed instance, set `PAPERCLIP_CLOUD_UI_SNIPPET_B64="$(base64 < snippet.html)"`, restart, open `/`, and confirm the decoded snippet appears before `</body>`. On a self-hosted instance, confirm no snippet is injected. ## Risks Low risk. The injection path and the Cloud-managed gate are unchanged. A malformed base64 value is ignored, so the failure mode is a missing widget, not corrupted HTML. The decoded value is trusted operator configuration, the same trust model as the plain variable. ## Model Used - Claude Fable 5 (Anthropic, `claude-fable-5`) via Claude Code CLI, agentic workflow with tool use (code edits, test runs, and live API verification of the WAF behavior). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
bc68312327 |
ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.
## Linked Issues or Issue Description
Refs #13243.
**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.
**Current behavior**
For merge
|
||
|
|
ad4f0b5867 |
Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtime settings can bind organization secrets to an adapter environment > - Paperclip redacts plain environment values when it returns a saved agent to the UI > - A saved-agent test sent the redacted `CODEX_HOME` value back to the server > - Codex ACP also received the API key without an ACP API-key authentication request > - This pull request restores saved environment values for tests and selects API-key authentication for Codex ACP runs > - The benefit is that Codex agents can test and run with an organization-scoped OpenAI API key ## Linked Issues or Issue Description **What happened?** Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`. Secret normalization rejected that placeholder. Remote Codex ACP runs received `OPENAI_API_KEY`, but session creation stopped with `Authentication required`. **Expected behavior** Paperclip must use the saved `CODEX_HOME` value when it tests an existing agent. Codex ACP must select API-key authentication when `OPENAI_API_KEY` is available. **Steps to reproduce** 1. Create an organization-scoped secret named `OPENAI_API_KEY`. 2. Give a Codex agent access to the secret. 3. Save the agent runtime settings. 4. Test the saved agent again. 5. Run the agent in a remote sandbox through ACP. **Paperclip version or commit** Reproduced on master before commit `68c17709d7c051a804a416263e2e08920f1dfcb1`. **Deployment mode** Self-hosted server with a remote sandbox environment. **Installation method** Built from source. **Agent adapter(s) involved** Codex. ## What Changed - Send the saved agent ID with adapter environment tests. - Restore redacted plain environment values from the saved agent before test-time secret resolution. - Select the Codex ACP `api-key` authentication method when `OPENAI_API_KEY` is present. - Add focused regression coverage for saved-agent tests and remote ACP launch configuration. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/test-agent-setup.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `git diff --check` ## Risks - Low risk. The test route reads saved configuration only when the request supplies a compatible agent ID and the caller can update that agent. - The Codex ACP change applies only when `OPENAI_API_KEY` exists and no explicit `DEFAULT_AUTH_REQUEST` exists. - There are no schema migrations or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. The context-window size is not exposed in this runtime. The model used reasoning, repository search, file editing, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3bafac12f7 |
refactor: remove automatic productivity reviews (#13263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its recovery loop keeps assigned work moving after execution failures. > - Productivity review used run counts, comment counts, and elapsed time to create management tasks. > - Infrastructure failures could satisfy those rules and create more tasks without evidence that the source work needed management review. > - This pull request removes that detector and its continuation holds. > - Bounded recovery, budgets, explicit blockers, and normal review stages remain in place. > - Existing task records stay readable and unchanged. ## Linked Issues or Issue Description Refs #5897. That request describes unwanted automatic productivity reviews and asks to preserve existing tasks. This change retires the feature instead of adding another configuration switch. Related prior approaches: Refs #9191, Refs #12489. Those changes excluded infrastructure failures or bounded review creation. This removal replaces the detector rather than tuning its thresholds. ## What Changed - Delete the scheduled detector, automatic task creation, evidence refresh, and productivity continuation holds. - Remove computed productivity fields, special attention items, badges, and Storybook fixtures. - Retain historical origin values, decision compatibility, and recovery recursion exclusions. Add no migration and change no existing task data. - Update the execution contract. Replace feature tests with regressions for legacy task reads, ordinary attention, and bounded continuation in the presence of an old review. ## Verification - Targeted attention, issue-route, startup, and UI tests: 4 files and 101 tests passed. - Updated issue-route and UI tests: 2 files and 61 tests passed. - Bounded continuation regression: 2 cases passed, including a legacy review plus pre-dispatch cancellation churn. - `pnpm check:token-gates`: all four gates passed. - `git diff --check`: passed. - `pnpm build-storybook`: passed. - Greptile: 5/5 on `a5a612eea`, with no actionable findings. - Scheduler and historical recovery regressions: 2 files and 28 tests passed. - Repository `pnpm -r typecheck` and `pnpm build`: passed. - The complete `pnpm test:run` suite passed across the CI server, serialized-server, and workspace shards on `a5a612eea`. Stopped the duplicate local monolithic run after the full CI suite passed; no completed local full-suite result is claimed. The targeted local suites above passed. - CI serialized shard 5 initially hit a 10-second timeout in the first interaction-route test. The complete file passed locally (78 tests), then the single CI rerun passed. - All CI gates are green, including the build and end-to-end suites. - A local merge check against current `master` (`ce09ea40b`) completed without conflicts. ## Risks - API responses no longer include the computed `productivityReview` field. Consumers must stop using it. - The scheduler no longer creates management work from elapsed time, run counts, or missing comments. This is the intended behavior change. - Existing review tasks and explicit dependencies remain in place. Historical origins still prevent recursive recovery treatment. No task cleanup or data migration occurs. - The native review handoff repair is separate from this removal. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
canary/v2026.911.0-canary.14
|
||
|
|
a23ae894a5 |
ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud images become deployable only after source verification passes. > - The typecheck job builds the native Runner binary through the server package. > - Fresh runners repeatedly compile Rust dependencies for that binary. > - This PR caches those dependencies for exact-source master verification. > - Workspace code and every typecheck still rebuild or run as before. ## Linked Issues or Issue Description Refs #13243 and #13257. **What existing behavior does this improve?** The typecheck portion of post-merge cloud source verification. **Current behavior** The typecheck job has no Rust dependency cache. An observed release build in this job took 4m 13s, including dependency compilation. **Proposed behavior** Restore dependency build outputs for canonical master pushes with the exact source SHA. Use a separate cache key from the Runner verification job, which builds other profiles. **Reason and benefit** A warm cache should remove roughly 2–3 minutes of dependency compilation from this job. Overall deployment gains depend on the remaining critical path. The first cache population still compiles from scratch. **Breaking changes** None. All checks remain enabled. Non-master callers compile without restoring or saving this cache. ## What Changed - Select the pinned Rust toolchain before the typecheck cache lookup. - Reuse the existing pinned Rust cache action with a typecheck-specific key. - Exclude workspace crates and installed cargo executables. - Test restore/save trust boundaries and document cache behavior. ## Verification - All workflow script tests pass locally, including nine new cache trust/contract cases. - actionlint passes for release-verify.yml. - Full local typecheck and build pass on the same source base (167s and 206s); `pnpm test:run` is still running and is recorded with #13257. This PR changes only the workflow, cache guard tests, and documentation. - All 32 current-head checks are successful or intentionally skipped, including the complete Linux test matrix, build, and Greptile 5/5 with no unresolved findings. Verify cache population and subsequent restore on actual master runs. ## Risks - The first run and any toolchain/dependency invalidation compile from scratch. - Cache restore/save overhead reduces the benefit for small dependency graphs. - Disable the cache step to roll back; the existing uncached build remains valid. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. Exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — focused change tests pass; full-suite local permission failures are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a27dd722a2 |
fix(ui): restore task page archive shortcut (#13253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task detail page gives operators keyboard shortcuts for fast task work. > - The `y` shortcut must archive the open task from the inbox and return the operator to the inbox. > - A streamlined UI gate limited this shortcut to tasks opened from a selected inbox row. > - Direct task pages no longer ran the shortcut. > - This pull request removes that gate and adds regression coverage. > - The benefit is consistent keyboard operation from every visible task page. ## Linked Issues or Issue Description **What happened?** The `y` keyboard shortcut did nothing on a task page opened outside the inbox. **Expected behavior** The `y` shortcut archives the open task from the inbox and returns the operator to `/inbox`. **Steps to reproduce** 1. Enable keyboard shortcuts. 2. Open a visible task from the Tasks page. 3. Press `y` while focus is outside an editor. 4. Observe that the task does not archive. **Paperclip version or commit** Current `master` before this pull request. **Deployment mode** Local development. ## What Changed - Enable the `y` archive shortcut on every visible task detail page when keyboard shortcuts are enabled. - Keep the existing input, dialog, modifier, hidden-task, and pending-request guards. - Verify that `y` archives the task and navigates to `/inbox` for a task opened from the Tasks page. ## Verification - `pnpm exec vitest run ui/src/pages/IssueDetail.test.tsx ui/src/lib/keyboardShortcuts.test.ts ui/src/lib/issueDetailBreadcrumb.test.ts` (139 tests passed) - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/ui build` - The full workspace typecheck and build reached the Rust runner. This environment does not include `cargo`, so the Rust steps did not run. - The full test suite entered unrelated serialized external-chat integration tests. It was stopped after the affected UI suite passed. ## Risks - Low risk. The change restores the shortcut behavior that existed before the streamlined UI gate. - The shortcut still ignores text inputs, open dialogs, modifier keys, and hidden tasks. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex. The Paperclip runtime did not expose the exact model ID or context window. The model used reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ce09ea40b0 |
feat(ui): show one rolling runner activity per commentary group (#13255)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task chat shows commentary and runner activity as work proceeds. > - The expanded activity list uses multiple text lines and an indented rail. > - The folded view can repeat reasoning and the current tool in separate places. > - This pull request shows one rolling activity between commentary messages. > - Users can open a group to follow its full history and inspect each item. > - The benefit is a compact feed with aligned icons and stable row height. ## Linked Issues or Issue Description **What existing behavior does this improve?** The runner activity groups in task chat, during live work and in saved history. **Current behavior** Activity is spread across a group summary, a reasoning ticker, and a current-tool row. Expanded rows have different alignment and can span several lines. **Proposed behavior** Keep one latest activity per commentary group. Roll new activities upward. Keep each expanded history row on one line. Use neutral failure text without an X icon. Keep the full details available by opening an individual row. **Additional context** Refs: #13246. That open PR covers related folded activity, steering timestamps, and status alignment. This change implements the separately reviewed activity-group design, including animated replacement, single-line expanded history, and neutral failures. It does not change steering timestamps or status pills. ## What Changed - Add a shared runner activity group with persistent expansion and reduced-motion support. - Use the group in the production runner turn and remove duplicate narration and current-tool rows. - Preserve commentary, plan cards, request receipts, and final-response placement. - Keep tool output, reasoning, research sources, file previews, and failure details accessible. - Render the production turn in desktop and mobile animated Storybook fixtures. - Add tests for animation identity, failure visibility, expansion persistence, and timeline order. - Validate resource link schemes and omit disclosure controls for items with no details. ## Verification - Focused activity, runner turn, protocol detail, and task-thread tests pass: 161 tests. - `pnpm check:token-gates` passes. - Browser checks confirm a 32-pixel compact row, centered icons, and one-line expanded rows. - `pnpm -r typecheck` and `pnpm build` pass. The final UI build also passes. - Two real native Codex runs completed in an isolated local instance. Each run created the expected file, performed separate delayed tool calls, reported a deliberate missing-file failure, recovered, and finished the task. - Browser verification confirmed that the open three-row activity group remained expanded after the second run moved into saved history. Saved error output remained accessible without red styling or an X. - The complete UI suite passes: 581 files and 5,979 tests. - `pnpm --filter @paperclipai/ui build-storybook` passes for the final production fixtures. - Greptile rates the latest commit 5/5, with zero new comments and no unresolved review threads. The security scan passes. - All repository test shards pass in CI. The duplicate local `pnpm test:run` was stopped after CI completed the full test coverage. The CI build failed twice because the worker ran out of disk space compiling the Rust runner tests (`No space left on device`, error 28). The build and its aggregate verify gate remain blocked by CI capacity. GitHub squash auto-merge is enabled and will wait for required checks to pass. ## Risks - This changes presentation and disclosure behavior. Users open a group to see earlier activity. - Animation uses stable item IDs. A provider that replaces IDs can trigger another transition. - No API, database, or runner protocol contracts change. > Reviewed `ROADMAP.md`. This improves an existing task-chat surface. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, code execution, and browser testing. The runtime does not expose a more specific model version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2904a3a6cc |
fix(ui): hide retry countdown after execution starts (#13258)
## Thinking Path
> - Paperclip helps operators manage AI agents and their tasks.
> - Task pages show countdowns for deferred checks and automatic
retries.
> - A retry keeps its scheduled start time after it enters the queue or
starts running.
> - The countdown treated that historical time as a pending deadline and
showed an overdue warning beside active work.
> - This pull request limits retry countdowns to retries that are still
scheduled and hides waiting surfaces on terminal tasks.
> - Operators now see a warning only when the displayed retry is still
waiting to start.
## Linked Issues or Issue Description
Refs #9783, which added the monitor surfaces. Searched related PRs and
issues; no duplicate fix was found.
**What happened?**
After a service restart resumed a task through an automatic retry, the
task showed an overdue retry banner while the agent was running. The
banner also remained when the task became done.
**Expected behavior**
A queued or running retry must not show a countdown against its past
scheduled start time. Done and cancelled tasks must not show waiting
banners.
**Steps to reproduce**
1. Open a task with an automatic retry scheduled for a known time.
2. Let the retry enter the queue and start running.
3. Wait until its scheduled time is more than one minute in the past.
4. Observe the overdue banner and Check now button while the agent is
working.
**Paperclip version or commit**
Reproduced on source commit
|
||
|
|
2fc5e3e43b |
test: isolate native controller takeover fixture (#13212)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native runtime recovery protects controller takeover with process identity checks > - The live-controller test mocked process liveness but left process start time to host state > - This pull request supplies the matching process start time in the fixture > - The benefit is deterministic coverage of the live-controller fence ## Linked Issues or Issue Description Closes #13172 ## What Changed - Complete the live-controller takeover fixture with the recorded process start time. ## Verification - Contributor reports focused native restart-recovery tests, related native tests, integration tests, and typecheck passed locally. Maintainer reran the 20 native restart-recovery tests successfully and verified current-head CI, Greptile 5/5, and zero unresolved review threads. ## Risks Low risk. This changes test setup only. Production controller fencing is unchanged. ## Model Used Original patch by @lorenzozanee; the original author did not record model details. Maintainer review and verification used OpenAI GPT-6 through Codex with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change is needed for this test fixture) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
847d00bdc3 |
fix(ui): use full mobile width for task tree (#13250)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The task detail page gives operators a task tree in a mobile side panel. > - The mobile sheet did not set an explicit full width. > - A maximum width rule could make the task tree narrower than the phone viewport. > - This pull request makes the mobile task tree sheet use the full viewport width. > - The benefit is a consistent and readable task tree on phones. ## Linked Issues or Issue Description **What happened?** The task tree sheet could use less than the full viewport width on a mobile task page. **Expected behavior** The task tree sheet must use the full available phone width. **Steps to reproduce** 1. Open a task on a phone viewport. 2. Open the task side panel. 3. Select the Tasks tab. 4. Compare the sheet width with the viewport width. **Paperclip version or commit** `7b829efdf` **Deployment mode** Built from source. ## What Changed - Set the mobile task side panel to `w-full` and `max-w-none`. - Add regression assertions for both width classes. ## Verification - `vitest run src/pages/IssueDetail.test.tsx --config vitest.config.ts` passes 103 tests. - `node scripts/check-token-gates.mjs` passes all four token gates. - `git diff --check` passes. ## Risks - Low risk. The change only affects the mobile task side panel width. - The desktop panel and other sheet variants do not change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.6, agentic coding with reasoning and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.13 |
||
|
|
d0b7ba4194 |
ci: route approved master cloud builds to AWS (#13243)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images built from master commits. > - Cloud image builds share GitHub-hosted capacity with other workflows. > - The organization already operates AWS runners through RunsOn Fleet. > - This pull request allows approved master builds to use a dedicated cloud Fleet. > - The benefit is separate build capacity with a quick operator rollback. ## Linked Issues or Issue Description Refs #13189, #13192. **What existing behavior does this improve?** Placement of the Docker cloud build after a master merge. **Current behavior** Every Docker cloud build uses a GitHub-hosted runner. Busy periods delay the job. **Proposed behavior** An operator variable enables the approved cloud Fleet for canonical master pushes and manual master builds. Other events, refs, and repositories use GitHub-hosted runners. ## What Changed - Add a guarded AWS runner selector to the Docker cloud job. - Keep the existing image cache, verification, and publication steps. - Test the selector against master, branch, tag, PR, fork, and disabled contexts. - Document provisioning requirements, placement checks, and rollback. ## Verification - 29 focused Node tests pass for routing, readiness, and disk handling. - The full workflow-script Node suite passes. - `pnpm -r typecheck` passes locally. - The pinned PR routing regression suite passes. The first live PR run assigned 21 jobs to the approved AWS PR group. AWS then reclaimed 16 Spot instances. The failed run is being repeated on GitHub-hosted runners while the Fleet moves to On-Demand. - Actionlint passes with existing shellcheck findings excluded (SC2012, SC2016, SC2129). - `git diff --check` passes. - Greptile reports 5/5 on commit `764d505a41dd2023751c3f361906fa9ea35bf0c6`, with no review threads. - All 30 current-head CI checks pass, including typecheck, build, all server/workspace test shards, Runner verification, and browser tests. Two Storybook checks are intentionally skipped for this change. Run: https://github.com/paperclipai/paperclip/actions/runs/34630550799 - The broader local test/build sequence is still running. This Mac has reported failures in unchanged application suites; their complete Linux CI shards pass. Local targeted workflow tests and typecheck pass. - Both On-Demand Fleets are deployed and healthy. Live master cloud-build verification follows the merge. ## Risks - Missing Fleet capacity or runner-group authorization can leave an AWS job queued. Disable `AWS_CLOUD_BUILDS_ENABLED` and rerun the workflow to use GitHub-hosted capacity. - The runner group must restrict access to this repository and the master version of `docker-cloud.yml`. - Docker needs more disk space than the PR Fleet. Provision 120 GiB disks and retain the free-space check. - This changes image build placement only. Source verification and migrator publication remain separate prerequisites. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.12 |
||
|
|
063ba59ae3 |
Fix workspace-ready notice rendering (#12716)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The task thread shows user, agent, and control-plane messages. > - Workspace-ready comments keep agent attribution for audit and authorization. > - These comments also have an explicit system-notice presentation. > - The client used attribution before the presentation contract. > - As a result, workspace-ready comments appeared as agent bubbles. > - This pull request gives the system-notice presentation priority. > - The benefit is a compact workspace notice with expandable details. ## Linked Issues or Issue Description **What happened?** The task thread showed a workspace-ready control-plane comment as a normal agent message. The comment kept agent attribution for audit and authorization, so the client did not use its system-notice presentation. **Expected behavior** The task thread must show a comment with the `system_notice` presentation as a compact system notice. The user must be able to expand the notice to inspect its details. **Steps to reproduce** 1. Open a task with a workspace that becomes ready. 2. Wait for the control plane to add the workspace-ready comment. 3. Observe that the task thread shows the comment as an agent bubble instead of a system notice. **Paperclip version or commit** `master` at `8eaa5caa0`. **Deployment mode** Local dev (`pnpm dev`). **Agent adapter(s) involved** Not adapter-specific. This is a core UI bug. ## What Changed - Give the `system_notice` presentation priority when the task-chat adapter selects the message kind. - Pass the notice presentation data to the task-chat model. - Improve the compact and expanded workspace notice details. - Add component tests, adapter tests, and Storybook cases for workspace notices. ## Verification - `pnpm exec vitest run ui/src/components/task-chat/task-chat-adapter.test.ts ui/src/components/task-chat/TaskChatSystemNotice.test.tsx` passes 15 tests. - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` passes 5,670 tests and reports four pre-existing workspace-runtime failures. The same four failures reproduce when the two unrelated server test files run alone. They do not use the changed task-chat files. ## Risks - Risk is low. Normal agent messages still use the agent presentation. - A comment with an explicit system-notice presentation now uses the compact notice renderer. - Regression tests cover both paths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5 (`gpt-5`) through Codex. The agent used reasoning, repository tools, and code execution. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b829efdf6 |
feat: show tasks created from a task by project (#13241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can cause an agent to create more tasks. > - Those tasks can belong to other projects or have another parent. > - The subtask view does not show all work created from the current task. > - This pull request adds a Tasks tab with separate subtask and creation groups. > - Operators can follow created work without changing its parent or project. ## Linked Issues or Issue Description **Problem or motivation** Operators need to see all work that an agent creates while running for a task. Parentage alone does not describe this relationship. Legacy and native runs must follow the same rules. **Proposed solution** Show all subtasks in one section. Separately group tasks created from the source task by their current project, with a No project group when needed. A created subtask appears in both sections. Use saved run context and recorded creation activity to find the source task. **Alternatives considered** Making all created tasks children would change their hierarchy. Removing overlap between sections would hide the creation relationship. This change keeps the two memberships separate. **Roadmap alignment** This extends the existing Activity log & action attribution capability. It does not add a new roadmap area. Related PR: #9727 adds a stored source-task field and inbound attribution UI. This PR adds the outgoing task list using existing run and activity records and does not require that schema change. ## What Changed - Add a company-scoped createdFromIssueId filter to issue lists. - Save the actor run during task creation, including legacy child-helper calls. - Recover historical run attribution from creation activity when the origin run is absent. - Render the production Tasks panel with all subtasks and independently grouped created work. - Keep progress only for subtasks. Add folding, hover fades and project links. - Fetch all result pages and refresh on issue activity. Show load failures with Retry. - Add database, API, UI and pagination tests, design-guide examples and Storybook pages. ## Verification - Before rebase: 158 targeted tests passed. Workspace typecheck, UI/server builds, Storybook build and token gates passed. - After rebase: full workspace typecheck and build passed. The cursor fix passes 26 focused tests and UI/server typechecks. - The full local test command completed its general-server group with 10,563 passing tests and two failures: a missing native-runner fixture and a concurrency-test timeout. Building the fixture and rerunning both affected files passed all 41 tests. The local command stopped before its remaining groups; all corresponding GitHub test shards passed on the submitted head. - GitHub checks on commit |
||
|
|
5545f6d166 |
feat: let user messages continue stopped native tasks (#13239)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A failed native run can leave a durable execution hold. > - The hold prevents automatic replay of actions with unknown outcomes. > - It can also prevent the agent from answering a new user message. > - A new user message should authorize a fresh turn after the prior execution stops. > - This pull request adds that admission path and preserves the existing execution gates. > - Users can continue the conversation without certifying every past action. ## Linked Issues or Issue Description **Subsystem affected** Server task wake admission and native execution recovery. **Problem or motivation** A native task can remain blocked after automatic recovery stops. A new user message is saved, but its run is cancelled before the agent can answer. **Proposed solution** Use a new authenticated user comment to authorize a fresh turn. Check stopped predecessor ownership and available history. Retain uncertain action outcomes. Commit the new run and hold retirement together. **Roadmap alignment** This is a focused improvement to the existing self-healing runs and recovery behavior. Builds on merged #13237, which covers legacy conversation continuation. This PR adds native admission and preserves native automatic-recovery eligibility and budgets. ## What Changed - Admit a fresh native turn for a new user comment after every held predecessor has stopped. - Validate the comment author, task, timing, process ownership, controller, and cleanup leases. - Preserve failed runs, unknown action outcomes, and the failed incident's attempt count. - Record the new comment and run in the existing recovery audit history. - Validate the saved continuation source and discard consumed user-wake authority from later automatic replacements. - Keep pause, approval, budget, ownership, and dependency interaction rules. - Add database and actual wake-path regressions. Update the execution contract. ## Verification - [Full CI run 34626750213](https://github.com/paperclipai/paperclip/actions/runs/34626750213) passed on `c58e6c087e0df7530c747d80b27d491da925a9c4`: all 31 reported checks passed, including all server/workspace suites, browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. - Greptile reviewed this exact head at 5/5. All review threads are resolved, and security checks passed. - All 218 local targeted tests passed across explicit native continuation, continuation history, safe replacement, durable chat wakeups, wake queue, issue liveness, native session resume, and run dispatch. The native implementation is unchanged by the final rebase onto master. - Full workspace `pnpm -r typecheck` and `pnpm build` passed on the final head. Complete test coverage is supplied by the green CI suites; local tests used the targeted suites above. - Regressions cover scoped authorization, concurrent delivery, live ownership, later admission gates, retained message receipts, and automatic replacement after terminal-task or reviewer changes. ## Risks - A fresh model turn can choose to repeat an action. Paperclip preserves prior history and does not replay recorded calls. - Missing process identity and remote ownership without a target-aware stop proof retain the hold. A terminal database row alone does not prove that execution stopped. - No schema or dependency changes. Existing historical tasks are not awakened by deployment. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. This session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.11 |
||
|
|
42a4f5b15b |
feat(ui): advance single-choice questions on selection (#13234)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask structured questions in task threads and the composer. > - The shared form shows one question at a time. > - A single choice already completes an answer, but the form requires another click on Next. > - This pull request moves to the next question when the user selects one option. > - Multi-select and custom answers keep their Next step. The last page keeps explicit submission. ## Linked Issues or Issue Description **What existing behavior does this improve?** The paged question form in the task composer and task interaction cards. **Subsystem affected** ui/ — React board UI. **Current behavior** The user clicks a single choice, then clicks Next to reach the next question. **Proposed behavior** A single choice shows its checked state for 160 ms, then opens the next question with an 80 ms fade. Reduced-motion mode skips the animation. Multi-select stays on the current page until Next. Other stays open for typing. The last question waits for Submit answers. **Reason and benefit** Remove an extra click from each single-choice question while preserving explicit submission. **Breaking changes** Single-choice selection now changes the page. Answer payloads and APIs do not change. The user can return to earlier answers with the previous arrow. **Additional context** Related work: #12640 introduced the task workspace. No duplicate auto-advance change was found. ## What Changed - Confirm a single-choice answer with a brief radio animation and row highlight, then fade into the next page and focus the new question. - Use motion tokens for the 160 ms confirmation and 80 ms page fade. Honor reduced motion and cancel pending advances when the user changes direction or closes the form. - Lightly highlight every selected row with a foreground tint that remains visible against the composer in both themes, for single-select and multi-select answers. - Preserve multi-select, custom answers, back navigation, and final submission. - Ignore repeated number-key events and selection during an upload or submission. - Update composer and card tests. Add interactive and verified composer stories to the existing interaction Storybook group. - Document the Storybook scenario in the developer guide. ## Verification - Focused composer, card, and motion-catalog tests: 121 passed, including animation timing and cancellation. - `pnpm check:token-gates`: passed. - `pnpm build-storybook`: passed. Browser interaction story: passed with the selection animation enabled. - Selected-row refinement: 121 focused tests, token gates, and Storybook build passed. Browser inspection confirmed row highlighting for single-select and multiple selected checkboxes, plus light-theme contrast. - Manual browser check: select SQLite, select two features, click Next, select Now, then submit. The summary contains every answer. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - All 31 remote checks passed on the latest commit (`04a3f3983`), including typecheck, general and serialized tests, production build/native-runner verification, canary validation, and browser tests. Two optional Storybook deployment/visual jobs were skipped. Greptile reviewed the same commit at 5/5 with no actionable findings. The selected-row highlight refinement is included in that verification. - Animation refinement: UI typecheck, UI build, token gates, and 121 focused tests passed. - The broader UI run had 5,940 passing tests and three failures in unchanged Inbox/IssuesList tests. Both affected suites passed in isolation (73/73), including all three previously failing cases. - The broad local `pnpm test:run` was stopped after reproducing the native-runner failure below and after the full remote suite passed. It is not a clean local full-suite result. - Local runner limitation: `server/src/services/native-runtime/native-session-resume.test.ts` has one reproducible failure in unchanged code. The damaged-epoch recovery test expects `run.attach requires a settled Codex provider session`; the runner instead reports `semantic tool input content digest does not match its transmitted input` at line 1019. After building the missing fake provider with `pnpm --filter @paperclipai/paperclip-runner build:rust`, the isolated suite has 36 passing tests and this one failure. No server or runner files changed in this PR. ## Risks - Selecting a single choice changes the visible question after a brief checked-state confirmation. Back navigation preserves the choice so it can be edited. - The final page still requires Submit answers. Selecting Other still requires text and explicit progress. - No database, server, API, or dependency changes. ## Model Used - OpenAI GPT-6 in Codex, with reasoning, code editing, terminal tools, and browser verification. Exact serving revision and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb640ec129 |
fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path > - Paperclip manages AI agents and their work. > - Wake admission decides when a task can create an execution run. > - Recovery can prohibit replay while the previous execution needs review. > - Dependency reconciliation kept creating runs before dispatch rejected that same hold. > - Each rejected run added another startup notice without doing useful work. > - This change checks the hold during admission and records repeated automatic waits once. > - Tasks keep their messages and can resume when the current gates permit execution. ## Linked Issues or Issue Description Related changes: Refs #13173 (stale completed-task continuations). Refs #12651 (dependency waits during recovery). **What happened?** A blocked task with completed dependencies can remain under a durable execution reconciliation hold. Each scheduler pass created a queued run. Dispatch then cancelled it before the adapter started. The skipped wake did not satisfy dependency wake deduplication, so this repeated and filled the conversation with “Couldn't start” notices. **Expected behavior** A known execution hold creates a waiting diagnostic without a run. Repeated automatic observations share that diagnostic. Clearing the hold permits a new wake only after the other gates pass. New comments remain available for the next eligible execution. **Steps to reproduce** 1. Assign a blocked task with a completed blocker. 2. Give the task an active reconciliation action, or a resolved action whose automatic recovery evidence still prohibits replay. 3. Run dependency reconciliation repeatedly. 4. Observe repeated cancelled pre-start runs on the base branch. This branch creates no runs while held and admits work after the effective hold clears. ## What Changed - Check effective execution holds under the issue admission lock before inserting runs. Keep the final dispatch check for races. - Share automatic wait diagnostics across producers, wake keys, and service restarts. Apply the helper to reconciliation, dependencies, pause holds, availability, budgets, and disabled heartbeats. - Preserve ordinary comment and interaction receipts during execution holds. Prevent release from draining them while replay is blocked. Keep external-chat receipt authorization intact. - Group empty pre-start reconciliation cancellations into a neutral waiting notice. Keep started runs and the full run history. - Document the waiting contract and add database-backed, UI, and browser regressions. - Stabilize two existing verification tests: allow the asynchronous chat lease transition a bounded five-second wait, and accept either legitimate damaged-session refusal while retaining exact archive-evidence assertions. ## Verification Passed targeted tests: - `pnpm exec vitest run server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests. - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` — covered in the initial combined test run; UI suite passed. - Run-dispatch adapter tests passed in the combined gate regression run. - `pnpm exec vitest run server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including held receipt replay, promotion, and revoked access. - `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass. Repeated held messages create no additional runs or provider prompts and do not replay writes. - `pnpm check:token-gates` - `pnpm check:module-boundaries` - `git diff --check` `pnpm -r typecheck` and `pnpm build` pass. The full local `pnpm test:run` invocation did not finish green: it encountered exhausted local PostgreSQL shared-memory slots, a missing fresh-worktree runner test binary, and tests loaded across in-flight edits. The affected chat/database suites passed on rerun (81 tests), and the targeted lifecycle/recovery verification passed (3 tests). After building the runner test binary, the full native session suite also passed (37 tests). Final-head [CI run 34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475) passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all server/workspace test shards, all three browser shards, runner verification, typecheck, build, release dry run, and the aggregate verification gates. All 31 reported checks passed; the two conditional Storybook checks were skipped as intended. Greptile reviewed that exact commit at 5/5 with no unresolved review threads. ## Risks The wait record is diagnostic only. It must never count as a delivered wake or bypass a current gate. Tests cover repeated and concurrent admission, resolved no-replay evidence, a remaining dependency after hold clearance, deferred comments, and release gating. Explicit user requests and authorized chat receipts do not share automatic diagnostics. No migration or historical data deletion is required. Existing provider retry budgets remain unchanged. ## Model Used OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted snapshot ID or context-window size. Used repository inspection, code editing, command execution, tests, and review tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1d26ae965e |
fix(ui): keep active runner status current and say Working (#13238)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task transcript shows a running agent's progress.
> - The active-run query stops polling when the live-run list has data.
> - The transcript still preferred that initial snapshot, so an old
execution-confirmation state could remain after work resumed.
> - This pull request uses the refreshed snapshot for the same run and
keeps active status text at Working.
> - Operators can see current activity without connection-state jargon.
## Linked Issues or Issue Description
**What happened?**
The task transcript said Reconnecting while the runner continued sending
messages and calling tools. The stale projection could also hide the
Thinking tail or stop the status spinner and timer.
**Expected behavior**
The selected run uses its current live snapshot. Active transcripts say
Working and show current activity. Completed and failed runs say Worked
and Stopped.
**Steps to reproduce**
1. Open a running task before its execution confirmation arrives.
2. Let the active-run query stop polling when the live-run list returns
the run.
3. Let the list refresh to working while the cached active-run snapshot
still says reconnecting.
4. Inspect the transcript status and activity tail.
**Paperclip version or commit**
Base commit:
|
||
|
|
4fde92107e |
fix(ci): reuse cloud source verification for npm canaries (#13233)
Reuse the exact master source-verification result before npm canary publication, removing a duplicate verification matrix while preserving fail-closed release checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
52811c6ce6 |
fix(tasks): require resume before sending to paused tasks (#13232)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task execution controls let board users pause a task or its subtree. > - The composer still accepted messages while a pause hold was active. > - A paused task must require an explicit resume before the user can send another message. > - This pull request replaces the composer with an amber pause card and checks board comment writes on the server. > - The user keeps their draft and resumes through the existing task controls. ## Linked Issues or Issue Description Refs #13104. Refs #13119. **What existing behavior does this improve?** The task composer and existing task/subtree pause controls. **Current behavior** A paused task can still receive a board message. The pause notice sits outside the composer, which leaves the send action available. **Proposed behavior** Show an amber takeover in both task chat and the classic composer. Preserve the draft. Require the user to resume the task or the ancestor subtree before sending. Reject board comment writes through either supported write route while the pause hold is active. **Breaking changes** Board comment writes to a paused task now return HTTP 409. Agent run reports remain supported during a pause. There is no schema migration. ## What Changed - Add a shared amber composer takeover with task, subtree, saved draft, pending, and error states. - Use effective ancestor pause state in both composer interfaces. Refresh it after pause events, task updates, and rejected sends. - Preserve draft text and attachments. Hide editor, send, queued edit, and pending question controls while paused. - Check active pause holds before board comment writes can mutate tasks, store comments, or wake agents. - Connect the approved Storybook examples to the production component and update the design and behavior docs. - Add browser coverage for both composers, draft persistence, resume, inherited holds, and rejected writes. Update ACP continuation coverage for the explicit resume requirement. ## Verification - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: `pnpm build-storybook`. - Passed: `pnpm check:token-gates` and `git diff --check`. - Passed: focused UI tests (398 tests) and server route tests (127 tests). - Passed: `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/paused-composer.spec.ts tests/e2e/acp-stop-continuation.spec.ts` (5 tests). - Passed: manual browser walkthrough in a disposable local instance. Pause with a draft, refresh while paused, resume, send, and reopen. The draft returned, and one message persisted. The amber card and resume dialog were readable with no clipping. - Full local `pnpm test:run` did not pass: the general-server stage recorded 9,072 passing tests, 6 database setup failures from macOS shared-memory exhaustion, and 4 failed tests. This stopped the script before its later groups. Latest-head CI runs those groups independently. - Local follow-up: the Git file-resource load test passed on rerun (4 tests); native finalization migration passed after clearing the abandoned browser-test database allocation. Building the native debug fixtures fixed the missing fake provider. The remaining native-session recovery assertion also reproduces on untouched base commit `87b3e5fc6` (36 pass, 1 fail on both base and PR). It expects a settled-session error but receives a semantic-input-digest error. - The final UI build, UI typecheck, token gates, both thread suites (182 tests), and all five browser tests passed after the queued-action review fix. All 31 latest-head CI checks passed, including all server, workspace, browser, build, release, and security gates. Two optional Storybook jobs were skipped by workflow policy. Greptile reviewed `32d8fb5f5` at 5/5 with no open findings. - Review the Paused Composer and Tasks / Execution Controls stories. Pause a task with a draft, verify the amber card, resume, and verify the draft can be sent once. ## Risks - Clients that used board comments to continue paused work must resume first. The response is an explicit HTTP 409. - Pause state can change while a page is open. Live updates refresh the composer, and the server rejects stale sends before their side effects. - Resume keeps the existing dialog and optional agent wake behavior. Agent reports from interrupted runs remain allowed. ## Model Used OpenAI Codex, based on GPT-6, assisted with design, implementation, code execution, and browser verification. The exact runtime model ID and context window are not exposed in this session. The agent used reasoning and tool calls. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the relevant tests locally and they pass; the full local-suite limits are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.10 |
||
|
|
96bba78fba |
feat: add readable Storybook branch bookmarks (#13231)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Maintainers use Storybook previews to review the board UI. > - Branch previews need stable bookmarks that people can read. > - The current publisher only provides a hashed branch path. > - This pull request adds a readable branch bookmark after each successful upload. > - Existing branch and build links keep working. ## Linked Issues or Issue Description Refs #13226. **What existing behavior does this improve?** Manual Storybook publication for repository branches. **Current behavior** The stable branch path contains a hash. The expected `/storybook/branches/master/` URL does not exist. **Proposed behavior** Each publication updates a readable bookmark. The action summary and Markdown artifact link it. Master uses `/storybook/branches/master/`. Other names use a safe path segment that preserves case and escapes special characters. **Reason and benefit** Maintainers can save and share a readable URL that opens the latest published branch build. **Breaking changes** None. Existing hashed branch entries still update. Existing build URLs remain valid. **Additional context** This follows the publisher in #13226. A duplicate search found no related bookmark change. It does not overlap planned core work in ROADMAP.md. ## What Changed - Generate readable branch bookmarks without collisions with existing build directories. - Upload the bookmark only after the full build and compatibility entry uploads succeed. - Link the bookmark in the existing summary and Markdown artifact. - Document branch-name escaping and test path isolation, stable links, and upload order. ## Verification - `node --test scripts/__tests__/storybook-deploy.test.mjs`: 20 tests pass. - `actionlint .github/workflows/storybook-deploy.yml .github/workflows/storybook-visual.yml`: passes. - `git diff --check`: passes. - [Master bookmark publication](https://github.com/paperclipai/paperclip/actions/runs/34613344758): passed. Opened `/storybook/branches/master/` in the browser and confirmed a story renders. Downloaded the Markdown report and verified its bookmark link. - [Feature branch bookmark publication](https://github.com/paperclipai/paperclip/actions/runs/34613449034): passed. Its separate bookmark uses `codex~2Fstorybook-bookmarks`. - Greptile: 5/5 on `dccaf10413ecf447cb34e622b6b3c505791abb51`, with no unresolved review threads. All current-head Paperclip CI gates pass, including typecheck, tests, build, browser suites, and the canary dry run. - Full local repository checks were not repeated for this focused publisher change. The preceding run passed typecheck but encountered unrelated native-session test failures. ## Risks - Special characters in branch names use `~HH` byte escapes. For example, `feature/foo` becomes `feature~2Ffoo`. - Names that could overlap an existing hashed build directory escape the final hyphen. Very long names retain a hash suffix. - The two branch entries update separately. If the final upload fails, the workflow fails and a rerun can repair the bookmark. ## Model Used OpenAI GPT-6 via Codex, with reasoning, shell tools, and live deployment verification. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b03c4b9eb |
perf(ui): keep long task chat responsive during streaming (#13229)
## Thinking Path > - Paperclip lets operators manage agent work through tasks. > - Task chat keeps responses and run history together. > - A live update rendered every historical bubble and hidden tool row again. > - New image callbacks also forced unchanged markdown to parse again. > - Long conversations saturated the browser main thread. > - This change reuses unchanged history and mounts folded tools on first inspection. > - Operators can read and reply while work continues. ## Linked Issues or Issue Description **What happened?** Chat-style tasks with substantial scrollback became almost unusable. A deterministic browser reproduction with 200 long responses and 4,000 tools consumed 97.8% of the main thread during live updates. It delivered only 8 updates during the sample. **Expected behavior** The task should remain responsive during streaming. Historical markdown and unopened run details should not repeat expensive render work. **Steps to reproduce** 1. Install dependencies with `pnpm install`. 2. Run `pnpm exec playwright test --config tests/perf/task-chat/playwright.config.ts`. 3. Compare the attached performance JSON. The new test fails against the original rendering code. **Paperclip version or commit** Reproduced at `a05b828bc`. The branch is rebased on current master. **Deployment mode** Local Chromium and Vite with deterministic fixtures. No database or agent credentials are required. Related work: #10463 reduces the issue-page bundle. This change addresses repeated rendering after the page loads. No duplicate scrollback fix was found. ## What Changed - Keep the bubble image callback stable so unchanged markdown can skip parsing. - Memoize the settled history separately from the header and streaming tail. Keep the brief renderer and default attachment array stable. - Mount folded tool history on first expansion. Keep it mounted afterward to preserve child state and closing motion. Runtime request receipts remain visible. - Add deterministic rendering tests to the normal Vitest suite and an opt-in Chromium regression fixture. - Document the reproduction, commands, scope, and local measurements. ## Verification - Browser reproduction: 97.8% main-thread utilization before; 11.6% after for tail-only updates; 31.5% after when projection recreates history objects. Both fixed cases delivered 32 updates. - Browser checks pass for scroll-position retention, typing, return to latest, tool inspection, and retained expansion state. - Focused component suite: 155 tests passed. Post-rebase thread/performance rerun: 101 tests passed. - Recursive typecheck, build, Storybook build, and token gates passed. UI typecheck passed after the final edits. - Local UI/CLI lane: 5,910 tests passed; ten files hit worker-start timeouts, then all ten passed with two workers (20 tests). Shared/adapter lane: 3,128 tests passed. - Full local `pnpm test:run` was attempted and is **not green**: its server lane recorded 10,517 passes and 11 failures plus fixture/setup errors. Queue (31 tests), Cursor/Git-load (9 tests), and missing-binary failures cleared on isolated reruns / building the runner test binaries. Two native suites still cannot initialize embedded PostgreSQL on this host. - The remaining native-session recovery assertion was reproduced in a clean worktree at base `a20ecce40` (1 failed, 67 passed across the native/queue suites). It expects a settled-session error but receives a semantic-tool-input digest error. No server or runner files changed in this PR. - All substantive CI jobs have passed, including build, typecheck, all server/workspace test shards, all three e2e shards, and the canary dry run. The unchanged Slack ordering test exhausted its one-second wait on the first run; its shard passed on rerun. Final aggregate verification passed: **31 passing checks**, no failures or pending checks; two optional Storybook deployment/visual checks were skipped. Greptile is **5/5**, with no unresolved review threads. - The supplemental local serialized route run was stopped after CI passed all five serialized shards; it had reported no failures. - The browser fixture uses the real chat rendering components with a plain textarea. It does not test the full composer or server transport. Timing results are local samples; deterministic render-count tests provide normal CI coverage. ## Risks - Closed tool content becomes available to DOM search only after first expansion. Visible run summaries and runtime request receipts remain available immediately. - Opened run history remains mounted to preserve child state. The first full markdown render still scales with conversation size. - Memo dependencies must stay current when adding render inputs. Tests check content edits and replacement gallery callbacks. - No API, database, or permission changes. ## Model Used OpenAI GPT-6 (Codex). Exact serving variant and context-window size are not exposed in this session. Used reasoning, repository tools, code execution, and browser testing. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted/UI/workspace checks pass; the full local server suite has the baseline/host failures documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1616046c24 |
fix(runner): prevent trace scans from delaying live events (#13228)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner sends events that keep task state and steering
controls current.
> - Debug trace correlation reread and parsed the full trace for each
pending event.
> - Repeated scans blocked event delivery and left task state minutes
behind the provider.
> - This pull request indexes appended trace records once and reuses
those correlations.
> - Operators get current run state while native trace evidence remains
available.
## Linked Issues or Issue Description
**What happened?**
Native runs appeared live after their provider turn ended. Steering was
unavailable while the board still showed an active run. A CPU profile
attributed 93% of sampled time to trace lookup and file reads.
**Expected behavior**
Debug trace correlation must not delay run events or task state by
minutes.
**Steps to reproduce**
1. Enable native provider trace capture for a Codex run.
2. Produce a long event stream with pending correlations.
3. Compare provider event times with persisted run event times.
**Paperclip version**
Reproduced on source build
canary/v2026.911.0-canary.9
|
||
|
|
87b3e5fc61 |
fix(adapter-utils): stage selected skills into the sandbox for a remote Claude ACP run (#13196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A user selects skills for an agent, and the host materializes those skills into a bundle the agent reads > - An agent can run in a remote sandbox, where the host must stage every file the agent needs > - On the Agent Client Protocol lane the host built that bundle and then named its host path in the prompt, but it never staged the bundle into the sandbox > - The agent therefore read a path that does not exist inside the sandbox, and the run failed on the missing skill file > - The command-line lane of the same adapter already stages a `skills` asset and reads the in-sandbox directory back from the staged runtime > - This pull request carries that proven pattern to the Agent Client Protocol lane, so a selected skill reaches the agent in a remote run ## Linked Issues or Issue Description No public issue exists for this change. The description follows. **What happened?** A remote run of the Claude adapter on the Agent Client Protocol lane could not read any selected skill. The host builds the skill bundle in its own state directory, then writes that host path into the prompt as `Skill root: <path>`. The remote seam of that lane staged one asset only, the configuration seed. It staged no skills asset, so no skill file crossed into the sandbox. The agent then tried to read the skill file at the host path, and the read failed with a missing-file error. **Expected behavior** A remote run receives the skills the user selected, and the prompt names the directory that holds those skills inside the sandbox. **Steps to reproduce** 1. Select one or more skills for an agent that uses the Claude adapter. 2. Start a run for that agent in a remote sandbox on the Agent Client Protocol lane. 3. Ask the agent to read the skill file at the path the prompt names. The file is not there. **Agent adapter(s) involved** The Claude local adapter, on its Agent Client Protocol lane. The shared engine in `packages/adapter-utils` carries the prompt rewrite. **Additional context** The command-line lane of the same adapter already stages a `skills` asset and remaps onto the staged directory. This change reuses that mechanism instead of adding a new transport. One other adapter shows the same host-path shape on its own Agent Client Protocol lane. That lane is tracked separately and this pull request does not change it. ## What Changed - Return the host skill bundle directory from the Claude skill runtime step, and carry it through the remote managed-home context to the staging seam. The value is null for a non-Claude agent, for a run that selects no skill, and for a run whose selected skills all fail to materialize. - Stage that bundle as a `skills` asset on the Claude Agent Client Protocol remote seam, and only when the run selected a skill. - **Stage that asset with `followSymlinks: false`.** The bundle holds an owned copy of each selected skill, and the copy step never copies a symbolic link at the root or at any depth. So the bundle contains no symbolic link, and staging has none to follow. Refusing to follow one also stops a link planted in the bundle directory after the copy from pulling an unrelated host file into the sandbox. A regression test walks the real adapter sources and pins the reviewed `followSymlinks` value at every skills staging site, so a new or changed site fails the test. - **Drop a skill whose staged copy has no usable `SKILL.md`** from the prompt, the skill identity, the command notes, and the bundle, and log which skill was dropped and why. Without this, a skill whose copy failed, or whose `SKILL.md` is a symbolic link the copy step skips, stayed advertised in the prompt while its file was absent — the same missing-file symptom this change exists to fix. - Rewrite the `Skill root:` prompt line, the skill identity, and the command notes onto the in-sandbox directory. The rewrite runs in the engine, after the workspace placement returns the staged runtime. A compatible session resume reuses the cached staged runtime, so the rewrite runs on that path too. - Keep the session fingerprint on the host-independent skill identity. A change to the selected skill set still invalidates a warm session, and the volatile sandbox path stays out of the hash. - A local run, and a run with no selected skill, keep their current behaviour. ## Verification - `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` — 28 of 28 pass. - `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` — the new engine tests and the staging-site tests pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`, and the same check on the adapter package — both exit 0. - The end-to-end test drives the lane against a local sandbox stand-in. It reads the skill root out of the prompt the runtime received, and then opens the skill file at that path. That is the reported symptom, proved closed. - The new tests carry a sensitivity control. Restoring only the production files to their previous content fails 7 of the 9 new tests. The other 2 do not depend on production code: one is a parser unit test for the source scanner. ## Risks Low risk, and the change is a two-way door. A revert restores the previous behaviour exactly. - **Scope.** The change touches one adapter lane. It does not change the local lane, and it does not change any other adapter. No existing staging site changes its `followSymlinks` value. - **The staged bundle and the workspace.** The staged skills land under the runtime directory inside the workspace. The workspace restore excludes that whole runtime directory, so the staged skills never return to the host worktree. A test proves the exclusion end to end. - **Session reuse.** The rewritten path never enters the session fingerprint, so it cannot invalidate a warm session, and a compatible resume applies the same staged path. - **Direction of data.** Files move from the host into the sandbox only. The change adds no path that writes sandbox content onto the host. - **A dropped skill.** A skill with no usable `SKILL.md` is now absent from the prompt instead of named but unreadable. The run logs the skill and the reason, so the cause is visible. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking and tool use, through Paperclip agents. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` passes — 28 tests. - [x] `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` passes. - [x] `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - [x] All continuous-integration gates are green. - [x] The automated review reports no open finding against the current head. ## Required Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant bug report template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect this change - [x] I have considered and documented the risks above - [x] All continuous-integration gates are green - [x] The automated review score is 5/5 with no open current-head findings - [x] I have addressed every reviewer comment that applies to the current head **Note on the branch history.** This branch first carried a different change: a filename-based admission filter that refused to stage files such as `.env` from a skill directory, together with a switch from symbolic-link bundles to copied bundles. That approach was rejected and **reverted** on this branch. It does not match the documented trust boundary, because the host already delivers credentials into the sandbox on purpose, and replacing the symbolic-link bundles broke live editing of a skill. The revert is in this branch's history. The file that work changed, `packages/adapter-utils/src/server-utils.ts`, is byte-for-byte identical to `master` here and is not part of this diff. Earlier review findings that name that file target the reverted code. All of them are resolved, and the automated review passes on the current head. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
974949a39b |
ci: spread cloud server verification across ten runners (#13227)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud waits for source verification before deploying a new image. > - The slowest server verification job spends about ten minutes running tests. > - Each job uses one test worker to preserve test isolation. > - This pull request distributes those suites across ten standard hosted runners. > - The benefit is a shorter verification path with the same test coverage. ## Linked Issues or Issue Description **Current behavior** In [readiness run 34572340764](https://github.com/paperclipai/paperclip/actions/runs/34572340764), the slowest server job ran for 638 seconds. Test execution used 594 seconds. This held readiness behind the image job. **Proposed behavior** Use ten general server jobs in the reusable release verification workflow. Keep the three chat jobs and every existing prerequisite. The complete partition test verifies that no server suite is omitted or duplicated. **Reason and benefit** Reduce merge-to-deployable time on the existing runner type. The next longest prerequisite was Runner verification at 526 seconds, so the initial expected total gain is about two minutes rather than a halving of readiness time. Measure actual queue and execution time before claiming a result. Related: #13198 introduced the separate chat lane. #12577 refreshes duration estimates; this change leaves that manifest alone. ## What Changed - Increase the general server matrix from five jobs to ten. - Verify the ten-way partition covers the complete server suite when combined with the chat lane. - Document runner demand and the unchanged local and PR grouping. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/__tests__/run-vitest-stable-shard.test.mjs`: 29 passed. - `actionlint .github/workflows/release-verify.yml`: passed. - Full local `pnpm -r typecheck` and `pnpm build`: passed. - All latest-head GitHub CI checks passed, including the complete Linux test partition, build, typecheck, and browser gates. Greptile: 5/5 with zero open findings. - [Ten-shard timing probe](https://github.com/paperclipai/paperclip/actions/runs/34606772388): all 16 jobs passed; slowest server job 6m 23s versus 10m 38s in the earlier five-shard sample. This compares the server lane, not total readiness, and is not a controlled same-source A/B. - The full local `pnpm test:run` is also running. It has reproduced previously observed macOS-only failures in unchanged skill-cache and native-session suites; the corresponding Linux CI suites passed. Final local results will be attached separately. No affected-workflow test failed. ## Risks Five additional concurrent jobs per release verification run increase runner demand and repeated setup work. Queueing can offset the gain. Test workers, timeouts, permissions, and readiness requirements stay unchanged. Revert the matrix and its partition test to restore the previous split. ## Model Used OpenAI GPT-6 / Codex, with reasoning, tool use, and code execution. The exact serving model identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the affected workflow tests locally and they pass; full-suite macOS limitations are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20ecce409 |
feat: publish CODEOWNER-approved Storybook branch previews (#13226)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Maintainers use Storybook to review the board UI. > - Reviews need public previews of selected repository branches. > - Each branch needs its own URL so previews do not replace each other. > - This pull request adds manual, CODEOWNER-controlled publishing to S3 and CloudFront. > - The action returns stable branch links and permanent build links in its summary and a Markdown artifact. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Storybook build and manual visual-review workflow. **Current behavior** The repository has no manual branch-preview publisher. A single GitHub Pages site cannot support independent publishers without combining their output. **Proposed behavior** A CODEOWNER selects a source branch and approves publication. Each branch has a stable CloudFront URL. A completed build becomes the branch target only after its upload succeeds. The action attaches `storybook-deployment.md` with the preview links and source commit. **Reason and benefit** Maintainers can share multiple branch previews at the same time. Branch builds have no repository token permissions or AWS credentials. Dependency caching and install hooks are disabled. The publisher cannot write runner dashboard files or delete objects. **Breaking changes** None. Normal visual checks keep their existing behavior. This does not change application code or GitHub Pages settings. **Additional context** Searched public issues and PRs for Storybook deployment work. No duplicate deployment proposal was found. This is maintainer infrastructure, not a roadmap-level core feature. ## What Changed - Add `Storybook Deploy` with a source-branch input and a manual entry through `Storybook Visual`. - Check the original actor and rerunner against default-branch CODEOWNERS. Require a protected deployment environment with CODEOWNER reviewers. - Separate public-source builds with no repository permissions from an OIDC publisher restricted to the Storybook S3 prefix. - Publish distinct branch URLs and retain build URLs. Preserve Storybook deep links across the branch redirect. - Add the run summary, a downloadable Markdown deployment report, focused tests, and operator setup docs and IAM policies. ## Verification - `node --test scripts/__tests__/storybook-deploy.test.mjs`: 19 tests pass. - `actionlint .github/workflows/storybook-deploy.yml .github/workflows/storybook-visual.yml`: passes. - [Feature branch live publication and deployment-only rerun](https://github.com/paperclipai/paperclip/actions/runs/34533202273): passed. - [Master branch live publication](https://github.com/paperclipai/paperclip/actions/runs/34533204743): passed. - Both public branch URLs render a component story without browser errors. A deployment-only rerun updates only the selected branch entry and preserves the previous build URL. - AWS policy simulation allows Storybook uploads and denies dashboard writes and object deletion. - Full local typechecking passes. Full local tests, build, and current-head PR checks are running. - [Revised build and Markdown artifact validation](https://github.com/paperclipai/paperclip/actions/runs/34605623088): passed. Downloaded the report and verified its branch URL, build URL, and source commit. - The public verifier also checks that the stable branch URL points to this build and rejects stale targets. ## Risks - Storybook previews are public. Maintainers must publish only public UI fixtures. - Retained builds accumulate until an operator prunes them. - Environment reviewers must stay synchronized with CODEOWNERS. The workflow fails closed if its environment loses required protection. - The existing CloudFront distribution is shared with runner reports. Separate S3 prefixes and a dedicated role prevent the publisher from overwriting those reports. ## Model Used OpenAI GPT-6 via Codex, with reasoning, shell tools, and browser verification. The exact runtime model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d10cbde815 |
fix(recovery): reject stale productive continuation wakes (#13173)
## Thinking Path > - Paperclip manages agent work through tasks and runs. > - Recovery continues assigned work when no live execution path remains. > - A recovery sweep can read an in-progress task before its run completes. > - The sweep can then observe the successful run after completion has changed the task status. > - This pull request checks current status and assignment under the existing enqueue lock. > - A stale continuation leaves a skipped wake receipt and creates no run. > - Task chat also omits an empty continuation cancelled before it started because its task had become terminal. ## Linked Issues or Issue Description Related public work: #10779 and #8419. Those older open changes address terminal disposition across other recovery paths. This change uses the existing scheduler guard for productive successful-run continuation and adds real database lock contention coverage. **What happened?** Recovery could combine an old in-progress task snapshot with a newer successful run. It queued an automatic continuation after the task was done. Dispatch cancelled that run before it started, but task chat displayed “Couldn't start” below the successful answer. This can happen after the native runner's finish result has already been accepted. It does not require a missing comment. **Expected behavior** Productive continuation must remain eligible when enqueueing acquires the task lock. Completion, cancellation, reassignment, or a move away from in-progress must prevent creation of the run. Actual execution stops must remain visible. **Steps to reproduce** 1. Let recovery select an assigned in-progress task whose latest run succeeded with productive progress. 2. Hold the task row lock in another transaction and change the task to done. 3. Let recovery attempt to enqueue while that transaction holds the lock. 4. Commit completion. Before this fix, recovery creates a redundant run from the stale snapshot. **Paperclip version or commit** Reproduced against master at `4042eb1c4` with deterministic integration tests. **Deployment mode** Built from source with PostgreSQL. The bug is in core recovery and is not adapter-specific. ## What Changed - Pass the existing status-and-assignee guard for productive terminal continuation recovery. - Preserve a skipped wake receipt with the expected and actual task state, without creating a run. - Test actual PostgreSQL lock contention for native and legacy completion, cancellation, backlog, review, blocked state, and reassignment. - Omit empty redundant pre-start cancellations from native and legacy task chat. Preserve stop markers for runs that started. - Document recovery eligibility at enqueue time. ## Verification - All seven new race cases failed before the guard was connected. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Task chat suite: 98 tests passed. - Recovery integration suites: 290 tests passed, including 31 stale-queue tests. - Full CI verification passed on `7ea71f04d`: all 31 active checks succeeded, including all test shards, browser tests, build, typecheck, and canary release dry run. Storybook visual regression was skipped by its path filter. - Greptile reviewed this commit at 5/5 with no review threads. - The local `pnpm test:run` aggregate reported a setup failure in the unchanged `tool-access-service.test.ts` suite. Its isolated rerun passed all 231 tests without edits. The duplicate aggregate was stopped after the complete CI matrix passed; it is not counted as a successful local full-suite run. ## Risks Low risk. The backend guard applies only to productive successful-run recovery. It requires the task to remain in-progress with the same agent. Other wake sources keep their current policy. The UI change only suppresses empty redundant cancellations; run records remain available. No schema change or migration is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code edits, and local test execution. The exact deployment snapshot and context window are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a05b828bcd |
Reduce run polling and workspace inspection amplification (#13174)
## Thinking Path > - Paperclip manages agent work and shows run progress to operators. > - Run lists, live events, transcripts, and workspace details must remain responsive as usage grows. > - Run-list redaction rereads the full context for every run. Hidden tabs can still trigger requests through live events and manual timers. > - Workspace detail reads repeat Git inspection even when concurrent callers request the same state. > - This pull request batches registry reads, pauses hidden-tab refreshes, and caches Git inspection for display. > - Cleanup keeps fresh Git checks, and redaction keeps company and run boundaries. ## Linked Issues **What happened?** Run-list responses perform one extra database read per run and parse full context JSON to obtain small secret registries. Hidden tabs continue transcript reads and event-triggered refetches. Workspace detail requests repeat Git scans. **Expected behavior** A run list reads registries once. Hidden tabs stop recurring run reads and reconcile when visible. Concurrent workspace detail reads share a short-lived Git result. **Steps to reproduce** 1. Open run lists and task transcripts in several tabs while agents run. 2. Hide some tabs and observe transcript and event-triggered requests. 3. Request a 200-run list and count redaction database queries. 4. Request the same workspace detail concurrently and count Git inspections. Related: #5255 adjusts polling cadence. This change addresses hidden-tab lifecycle, batched registry reads, and workspace inspection reuse. No duplicate with this scope was found. ## What Changed - Batch heartbeat and live-run redaction into one company-scoped registry query. Select only registry JSON for run and issue redaction. - Resolve duplicate secret values once per request. Preserve each run's registry and remove registry material from responses. - Suspend company event sockets and transcript reads while hidden. Refresh active queries and resume transcript offsets on return. - Prevent queued event invalidations and developer health polling from fetching in hidden tabs. Gate legacy run-log readers in both UI variants. - Exclude legacy plugin placeholder connections from remote health probes. Select only due connection IDs in SQL before the sweep limit. Preserve existing plugin records. - Cache concurrent Git display inspections for five seconds, with at most 256 entries. Leave close-readiness and cleanup checks uncached. - Add regression coverage and document the performance behavior. - Stabilize the existing Rust descendant-lineage fixture: allow a bounded 30 seconds for 300 durable notifications under concurrent test load, retaining every correctness assertion and adding timeout diagnostics. ## Verification - Regression coverage verifies one registry query for 200 runs, per-run isolation, request-local secret resolution, decryption failures, Git cache expiry/bounds, hidden-tab pause, and visibility recovery. - Real PostgreSQL redaction/run-route suites passed all 57 tests; workspace-service coverage passed. The health-sweep regression verifies plugin placeholders and chat connections remain untouched and do not consume the sweep limit. - Both legacy transcript viewers retain history and resume their byte offset after visibility changes. The related visibility/progress/chunk suites passed all 29 tests. Other focused UI suites and token gates passed. - Full `pnpm -r typecheck` and `pnpm build` passed. Affected-package typechecks/builds passed after review fixes. The concurrent Rust provider suite passed 84 tests (two ignored), and Rust formatting passed. - Full local `pnpm test:run` stopped after the general-server group: 10,538 passed, 65 skipped, four failed. Fresh chat-delivery and health-sweep reruns passed; building the debug runner fixture cleared the native-event test. One unchanged native-session recovery assertion still fails locally with a semantic-digest error instead of the expected settled-session message. The full local command is therefore not green. CI runs the later groups separately and skips the two native-session tests requiring a prebuilt runner binary (confirmed in its 37-test native-session suite). - All CI gates pass on final head `ee610e737`: typechecking, general and serialized tests, browser tests, runner verification, build, and canary dry run. One server shard passed on its single retry after exposure fixtures encountered port 42001 where they assumed 42000; that suite also passed locally (25 passed, three platform-specific skips). - Greptile reviewed the final head at 5/5 with no actionable findings. ## Risks - Workspace delivery display can lag local Git changes by five seconds. Destructive operations still inspect current state. - Hidden tabs do not receive company live-event notifications until visible. Active queries refresh on return. - This change preserves legacy plugin records and does not repair instance-specific workspace rows. There is no database migration. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser inspection. The exact model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted regressions; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.8 |
||
|
|
932c8bec56 |
fix(ci): bake the managed runtime identity into cloud images (#13210)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed deployments start from the image built by the Cloud workflow. > - The managed runtime requests user and group 1001. > - The image currently builds the node user as 1000. > - Startup must remap that user, which can walk a large mounted home directory. > - This pull request uses the existing Docker build arguments to bake user and group 1001 into Cloud images. > - Matching the runtime identity removes that startup work and helps avoid health-check retries. ## Linked Issues or Issue Description Refs #13208, #1923, and #7861. Searched open and closed PRs for the Cloud UID change. The older #7861 addresses build context and volume ownership repair. This change uses the existing identity arguments in the Cloud workflow and preserves ownership repair. **What happened?** A measured rollout had a container log `Updating node UID to 1001` after startup. The container stayed at this step for at least 2 minutes 55 seconds before rollback stopped it. The baked node identity was 1000, while the managed runtime requested 1001. A health check timed out and the target required a second deployment attempt. **Expected behavior** Cloud images should already have the managed runtime identity. A matching image should skip user and group remapping. Fresh or mismatched volumes must still receive ownership repair. **Steps to reproduce** 1. Build the current Cloud image with its default build arguments. 2. Start it with `USER_UID=1001`, `USER_GID=1001`, and a populated home volume. 3. Observe the startup user remap before the application starts. **Paperclip version or commit** `fc06f7f05f42c675be71ff0927b6334405d520ed` **Deployment mode** Docker on managed hosts. ## What Changed - Pass `USER_UID=1001` and `USER_GID=1001` to the Cloud image build. - Check the pushed digest's baked identity before the entrypoint can repair it. Then check the normal entrypoint's effective identity and writable home before publishing the verified full-SHA tag. - Add a workflow regression and two entrypoint cases for a matching Cloud identity, including a mismatched volume. - Document the runtime identity and the first-build cache cost. ## Verification - Focused workflow and artifact tests: 27 passed. - Entrypoint tests: 11 passed. Actionlint passed. Full local `pnpm -r typecheck` passed. Full local `pnpm build` passed. The manual [Cloud image build](https://github.com/paperclipai/paperclip/actions/runs/34575473213) passed on the exact PR head. It checked Sentry, baked and effective identity, writable home, orphan reaping, and full-SHA publication. The new identity check took one second. All 30 PR checks passed; the Storybook workflow was intentionally skipped. Greptile reviewed commit `114d408f637a0b53e2e2b1339c263779b1e4ae54` at 5/5 with no findings or open threads. - The full local suite for the same application source was already run in #13205. Its macOS general-server phase had 10,471 passes and 70 failures in seven unchanged files. Those failures included missing Runner fixtures, filesystem errors, timeouts, a port conflict, and a load-count mismatch. After configuring Cargo and rebuilding fixtures, 37 of 38 native tests passed; one unchanged native-resume assertion still failed. Linux PR CI passed. This change adds entrypoint tests and does not change application code. ## Risks - The first build must rebuild layers that depend on the base image identity. Later builds can reuse them. - A future managed runtime identity change must update these build arguments and checks together. - The Dockerfile's self-hosted defaults remain 1000. Runtime overrides and mounted-volume ownership repair remain supported. - The observed startup delay supports this change, but fleet timing also includes provider startup, image pull, canary order, and retries. No fixed end-to-end gain is claimed before a live rollout. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and tool use. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused workflow tests; full-suite limitations are listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.7 |
||
|
|
fc06f7f05f |
fix(ci): isolate chaos verification by caller workflow (#13208)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments require verified artifacts for the merged source commit. > - Cloud readiness and the npm release independently run the same source checks. > - Their shared chaos workflow used only the source ref as its concurrency key. > - One caller could cancel the other caller's required job for the same commit. > - This pull request scopes that key to the caller workflow and source ref. > - Both callers can finish their checks without blocking deployment readiness. ## Linked Issues or Issue Description Refs #13192 and #13205. Searched for related open issues and PRs; no duplicate fix was found. **What happened?** The master push for `398d304e15739d1ee6105633bd8a0e42c929d33f` started Cloud readiness and Release together. GitHub cancelled the Cloud readiness chaos job before it acquired a runner. Its annotation reported a higher-priority waiting request for the same concurrency group. The required readiness gate cannot pass after that cancellation. **Expected behavior** Cloud readiness and Release must each finish source verification for the same SHA. Standalone chaos evals must also have a separate group. **Steps to reproduce** Merge a commit to master while the npm release queue is empty. Both callers reach the reusable chaos workflow with the same source SHA. See [the cancelled job](https://github.com/paperclipai/paperclip/actions/runs/34569569760/job/103168603926). **Paperclip version or commit** `398d304e15739d1ee6105633bd8a0e42c929d33f`. **Deployment mode** GitHub Actions on master. ## What Changed - Add the caller workflow name to the chaos workflow concurrency group. Retain source isolation and cancellation of duplicate calls within the same workflow. - Add a regression test that evaluates the group for Cloud readiness, Release, and standalone evals at the same source SHA. - Document the concurrency boundary in the readiness runbook. ## Verification - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 26 tests. - The new regression test fails against the previous concurrency key and passes with this fix. - `actionlint -shellcheck= -pyflakes= .github/workflows/runner-chaos-evals.yml .github/workflows/release-verify.yml .github/workflows/cloud-readiness.yml` passed. - `git diff --check` passed. - The full local typecheck passed for the same application source in #13205. Its macOS general-server test phase had 10,471 passes and 70 failures in seven unchanged application test files: missing Cargo/Runner test binaries, filesystem permissions, timeouts, a port conflict, and a load-test count mismatch. Linux CI test checks passed. The full local build passed with Cargo on PATH. This PR changes workflow configuration, its test, and documentation only. - All CI checks pass on the final head, including typecheck, tests, browser suites, build, and canary dry run. Greptile is 5/5 with no open findings. After merge, verify both callers' chaos jobs complete for the same master SHA and record the resulting readiness time. ## Risks - Two callers may now run chaos tests at the same time. This uses two existing GitHub runners, which is the intended cost of independent verification. - Renaming a caller changes its concurrency group. The fixed prefix keeps this child group separate from caller-level concurrency groups. - The readiness gate continues to require every verification prerequisite. No gate is bypassed. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (26 focused workflow/artifact tests) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.6 nightly/v2026.911.0-nightly.0 |
||
|
|
398d304e15 |
docs: measure cloud deployment through target health (#13205)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Hosted deployments need a verified image and migrator for the same source commit. > - The cloud readiness workflow certifies those inputs before a deployment consumer acts. > - Its completion time does not show when a tenant runs the new commit. > - This pull request documents each milestone from merge through target health and fleet completion. > - Operators can use the evidence to find the slow stage and measure a complete deployment. ## Linked Issues or Issue Description Refs #13192, #13188, and #13189. Searched related issues and PRs; no duplicate timing documentation change was found. **Issue type** Missing documentation. **Where is the issue?** `doc/cloud-build-readiness.md`, Timing and rollout. **What's wrong?** The timing instructions stop at the readiness job. That omits consumer queues, artifact resolution, and target deployment. An image can be ready while the tenant still runs an older commit. **Suggested fix** Record separate merge, image, readiness, canary health, and fleet completion timestamps for the same full source SHA. Keep preparation-only runs out of deployment results. ## What Changed - Define the evidence needed for each merge-to-deployment milestone. - Explain how consumer queues can hide upstream build gains. - Require target source identity as well as health, and report exclusions, retries, cache state, and queue conditions. ## Verification - `git diff --check` passed. - `node --test scripts/preview-artifacts.test.mjs scripts/__tests__/release-verify-workflow.test.mjs` passed: 25 tests. - Cross-checked the readiness identity and artifact prerequisites against the current workflows and consumer contract. - Full local `pnpm -r typecheck` passed using the session's installed Rust toolchain. The full local test suite and subsequent build are still running. - All CI checks pass and Greptile is 5/5 on the exact head, with no unresolved findings. This changes one documentation file and adds no runtime behavior. ## Risks - Low risk: documentation only. Timing must still use trusted run evidence and the actual target commit. A single measured run is not a latency guarantee. ## Model Used - OpenAI GPT-6 / Codex, with reasoning, repository editing, and command/API tools. Exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (25 focused workflow/artifact tests; full checks pending) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.911.0-canary.5 |
||
|
|
d56be3f3fc |
fix(ci): verify deployable cloud artifacts independently (#13192)
Verify source, build the cloud image, and wait for exact-source migrator packages concurrently. Emit Cloud deployable v1 only when every prerequisite succeeds for the merged full SHA. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5cc51fad06 |
fix(release): publish exact-source cloud migrators on merge (#13188)
Publish exact-source shared and database migrator packages for each master merge through the existing trusted Release workflow, independently of the full release and image build. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6a7025ebe3 |
ci: skip cloud runner cleanup when disk headroom is ample (#13191)
Skip cloud runner disk cleanup when both the Docker and workspace filesystems have at least 64 GiB free. Preserve the existing cleanup for low, unavailable, or invalid measurements and verify the actual shell behavior across eight scenarios. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5c660a32f3 |
ci: build cloud images independently for each merge (#13189)
Build cloud images independently for each master commit through a reusable workflow. Preserve production release dependencies and image runtime checks, and write cloud registry caches per commit with bounded ancestor imports to prevent overlapping builds from replacing each other's cache. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
59d74b68b2 |
ci: cache the native Runner in a separate Docker stage (#13195)
Compile the native Runner from its complete Cargo and protocol inputs in a separate cached Docker stage. Preserve Cargo validation and generated-contract checks during the normal application build, normalize input timestamps across checkouts, and compile the isolated target in PR CI. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
42961b6ef1 |
fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9e970df4c5 |
ci: cache Rust dependencies in release Runner verification (#13194)
Cache external Rust dependencies in trusted master release verification after selecting the package-owned toolchain. Keep source compilation and all validation unconditional; restrict both restore and save to the matching master push. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3fd556b8f6 |
ci: preserve weekly Docker tool cache across commits (#13190)
Keep stable Docker tool installation layers independent of application build version and commit metadata. Preserve the existing weekly tool refresh and runtime build stamp. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
399daa1f25 |
ci: publish verified full-SHA cloud image tags (#13187)
Publish the full-source-SHA cloud image tag only after the pushed digest passes runtime, revision, and platform checks. This makes verified images directly resolvable by Cloud. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f70accd3a4 |
fix(onboarding): answer the Claude paste at once, and show the code as dots (#13193)
The Claude card's button only moved to Connecting once the login was stored - a submit, a status poll and a completion read after the paste - so for about a second the customer had done their part and the button still read Waiting for code. The panel now reports the submit as it starts (onCodeSubmitted) and the step shows Connecting from that moment. That could not simply move the phase earlier: the two-second hold started when Connecting did, so it would have hired whether or not a credential existed. The hire now waits for both the stored login and two seconds of Connecting counted from the paste. onSubmitFailed gives the button back when a submitted code does not become a stored login, the field locks while a code is out, and Cmd+Enter no longer hires mid-connect. Reports that land after Back are ignored. The panel stays mounted through Back's exit, so a late failure reopened the card being left, and a late success hired a customer who had backed away. The second predates this change; its test fails the same way against master. The authorization code shows as dots. The OpenAI card is untouched.canary/v2026.911.0-canary.2 |
||
|
|
9effe51b63 |
fix(server): let the cloud-harness sandbox environment self-heal past operator-drift protection (#13177)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every cloud-harness-managed stack gets one platform-owned "Paperclip Computer" sandbox environment, reconciled from `PAPERCLIP_MANAGED_CONFIG` on boot > - That reconciler deliberately refuses to overwrite a row it classifies as operator-modified, to protect a self-hosted operator's hand-edited environment (#10979) > - But for the cloud-harness-managed row specifically, no operator has any path to hand-edit it at all — so a hash mismatch there can only be drift between two platform-driven reconciliation passes, never a real customization > - Roughly 10 staging stacks got stuck on a broken sandbox image because of exactly this: a `sandbox_image` campaign correctly delivered a fixed snapshot, but the reconciler classified the row as operator-modified and silently skipped applying it > - This pull request adds an explicit `platformFullyManaged` flag so the cloud-harness caller can assert that guarantee and let its own drift self-heal, without weakening the protection for every other caller (self-hosted kubernetes-execution-mode, tests, admin routes) where an operator genuinely can edit the row > - The benefit is that a sandbox-image rollout can no longer get silently stuck fleet-wide, while self-hosted operator customization keeps exactly the protection #10979 built ## Linked Issues or Issue Description No public issue exists for this internal-instance-discovered bug; opening directly per CONTRIBUTING.md path B, following the bug report template fields. **What happened?** After a `sandbox_image` campaign delivered a fixed Daytona snapshot fleet-wide, ~10 of 79 active staging stacks kept booting agents against the old, broken snapshot. Their `PAPERCLIP_MANAGED_CONFIG` env var and the reconciler's own bookkeeping (`built_in_managed_resources`) both correctly showed the new snapshot — but `environments.config.snapshot`, the field the runtime actually reads to acquire a sandbox lease, was never updated on those rows. **Expected behavior** A `sandbox_image` campaign (or any `PAPERCLIP_MANAGED_CONFIG` delivery) to the cloud-harness-managed sandbox environment should always converge that row's `config` to the newly-desired value, since no operator can have a competing edit to protect. **Steps to reproduce** 1. Boot a cloud-harness-managed stack; let the reconciler create the managed sandbox row and record its stock hash in `built_in_managed_resources`. 2. Somehow cause the row's live content hash to no longer match the recorded binding hash without an operator ever touching it (in the field, this happened via drift between two platform-driven reconciliation passes carried out across a catalog-version bump — the exact trigger wasn't fully pinned down, but is irrelevant to the fix). 3. Deliver a new `PAPERCLIP_MANAGED_CONFIG` (e.g. via a `sandbox_image` campaign). 4. Observe `ensureManagedSandboxEnvironment` classify the row `operator_modified` and skip writing `config`, even though `updateAvailable: true` is reported and the binding itself already advanced to the new stock hash. **Paperclip version or commit** `master` as of this PR. **Deployment mode** Cloud-managed stacks with `enableManagedSandboxOnly` declared (any Paperclip Cloud–provisioned staging or production stack). Related PR for context (not a duplicate — this is additive to it, not a revert): #10979, which introduced the `operator_modified` classification this PR narrowly opts the cloud-harness path out of. ## What Changed - `server/src/services/environments.ts`: added `platformFullyManaged?: boolean` to `ManagedSandboxEnvironmentInput`. When set, a plain content-hash mismatch against a real prior binding (i.e. `operator_modified` that isn't an archive-reaffirmation) is reclassified as `stock_update_available` before the skip-vs-apply branch, so it flows through the normal update path instead of being frozen. - `server/src/services/managed-environments.ts`: pass `platformFullyManaged: true` from both `ensureManagedSandboxEnvironment` call sites — the main boot ensure and the provider-recovery reactivation path. These are the *only* two callers driven by `PAPERCLIP_MANAGED_CONFIG`; `ensureKubernetesEnvironment` (self-hosted `kubernetes-execution-mode` bootstrap) and every other caller are untouched and keep the original protective default. - `server/src/services/managed-environments.test.ts`: updated the two `toHaveBeenCalledWith` assertions that now include the flag. - `server/src/__tests__/environment-service.test.ts`: two new tests — one confirming the bypass applies drift under `platformFullyManaged`, one confirming archive-reaffirmation still wins even under the flag. Archive-reaffirmation is deliberately *not* bypassed even under `platformFullyManaged`: a `sandbox_image` update must never resurrect a row something else deliberately kept archived after Paperclip's own provider-unavailability archival. That's a distinct, still-real signal, orthogonal to config drift. ## Verification - `vitest run` on `managed-environments.test.ts` and `managed-resource-drift.test.ts`: 26/26 pass, including the two updated assertions. - `environment-service.test.ts` — the file both new tests live in, and the file holding the two pre-existing tests this change must not regress ("classifies operator drift, preserves the row, and exposes the pending stock update" and "preserves an existing unmanaged sandbox row holding the desired name") — requires a real embedded-Postgres instance (`describeEmbeddedPostgres`) not available in the sandbox this was developed in; `getEmbeddedPostgresTestSupport()` reports unsupported there, so the whole file is skipped locally. I traced the reconciliation logic by hand against all four relevant tests (the two new ones plus the two pre-existing ones) line by line to confirm the expected outcomes, but **CI running this suite for real is the actual gate here**, not this description — please don't merge on a green run of everything else alone if this suite doesn't show as executed. - `tsc --noEmit`: zero errors in any of the four touched files. The pre-existing ~229 errors elsewhere in `server` are unrelated missing-module issues from packages needing a build step first, confirmed unchanged by this diff. - Manually reproduced the underlying bug against real staging data (a `paperclip-cloud`-managed stack whose `environments` row was stuck exactly this way) before writing the fix, and confirmed via direct SQL inspection that the recorded `built_in_managed_resources` baseline already held the correct desired snapshot on every affected stack — i.e. the reconciler already *knew* the right answer, it was just refusing to apply it. That data point is what ruled out "the campaign didn't actually deliver the update" as the cause. ## Risks - Scope is intentionally narrow: only the two `PAPERCLIP_MANAGED_CONFIG`-driven call sites pass the new flag; every other caller of `ensureManagedSandboxEnvironment`/`ensureKubernetesEnvironment` is byte-for-byte unchanged. The two pre-existing regression tests that specifically cover self-hosted operator-edit protection don't pass this flag and are unmodified. - The main residual risk is the unresolved root cause of *why* the hash drifted in the first place (a race between two close-together reconciliation passes, or a catalog-version-dependent change to what gets hashed, most likely) — this PR makes that drift self-healing rather than fixing whatever produces it. If the drift is being caused by a genuine concurrency bug (rather than an expected, occasional side effect of a stock-field/catalog-version change), that bug still exists and could recur; it just no longer gets stuck when it does. - Low risk of behavior change for real self-hosted deployments: none of them can reach the new code path, since only the two now-flagged call sites exist inside `managed-environments.ts`, itself gated to `PAPERCLIP_MANAGED_CONFIG` (which self-hosted `kubernetes-execution-mode` explicitly refuses to run alongside — see the existing mutual-exclusivity check this PR does not touch). ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, with tool use (file edits, shell/git, `gh` CLI, direct Postgres inspection of live staging data via `psql`/`pg`, Railway SSH for on-host diagnosis). No extended-thinking mode. Standard Claude Code context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — see Verification: the file holding the four load-bearing tests can't run in this sandbox (no embedded-Postgres support); traced by hand instead, CI is the real gate - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — none applicable beyond the inline doc comments this PR adds - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — pending CI run on this PR - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending review - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c5c80e1feb |
ci(release-verify): split server tests five ways like pr-trusted (#13185)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every master push publishes a canary through release.yml, gated by release-verify.yml — the fleet's staging deploys and the nightly/beta/stable chain all start from those canaries > - release-verify splits the server test suite across three shards with a 20-minute job cap, while pr-trusted splits the same suite across five > - The server suite grew on 2026-09-10 and the three shards moved to 17-19 minutes; that evening every push-triggered canary run was cancelled by the 20-minute cap mid-verify, and no canary published after 18:50 UTC > - This pull request mirrors pr-trusted's five-way server split in release-verify, putting shards back at the 10-15 minute range with real headroom > - The benefit is a canary lane that reports test verdicts instead of dying on an infrastructure cap ## Linked Issues or Issue Description **What happened?** Push-triggered Release runs stopped publishing canaries on 2026-09-10. Runs at 19:34, 22:30, and 22:37 UTC were all cancelled by "The job has exceeded the maximum execution time of 20m0s" on a `verify_canary / General tests (server (N/3))` shard. No canary published after 18:50 UTC, which also starves the staging fleet's continuous deploys. **Expected behavior** release-verify's server shards finish well inside the 20-minute cap and runs conclude with a test verdict, as pr-trusted's five-way split of the same suite does (10-15 minutes per shard). **Steps to reproduce** 1. Compare server shard durations in the `verify_canary` job across 2026-09-10: 11-14 minutes in the morning, 17-19 minutes from 15:06 UTC, over 20 minutes by evening. 2. Observe runs 34521169020, 34537798488, and 34538332689 cancelled at the cap. **Paperclip version or commit** `master` at `d1ba17eec` (current tip; its canary run was one of the cancelled ones). ## What Changed - `release-verify.yml`: the `general-server` matrix goes from three shards to five, byte-for-byte the shape `pr-trusted.yml` already runs, with a comment recording why. ## Verification - The identical five-way split runs green on every pr-trusted run (10-15 minutes per shard today, including on PRs merged this evening). - The suite's own growth (slower chat-connector tests) is being addressed separately; this PR only removes the artificial cliff. ## Risks - Low risk: two more runners per verify run; no test content changes. If shard durations regress further, the cap fires again — which is the correct signal once shards have honest headroom. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), extended thinking, tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.911.0-canary.1 |
||
|
|
2585ed0550 |
test(server): settle three contention flakes that killed canary verifies (#13186)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every master push publishes a canary through release-verify; the staging fleet and the nightly/beta/stable chain start from those canaries > - The server suite grew substantially on 2026-09-10 and now runs under real contention in CI, where three tests assert timing properties that only hold on an idle machine > - Each of the three failed a release-verify canary run that day (runs 34497348802 and 34517515849), and together with the shard timeouts (#13185) they kept any canary from publishing after 18:50 UTC > - This pull request makes the three assertions contention-tolerant without weakening the invariants they prove > - The benefit is a canary lane whose verdicts reflect the code, not the load on the runner ## Linked Issues or Issue Description **What happened?** Three server tests failed release-verify canary runs on 2026-09-10 under CI load: 1. `chat-channels.integration.test.ts › returns a retryable webhook failure when the delivery insert fails before durable receipt` — the duplicate-redelivery request drew the retryable 503 instead of an immediate 200 (run 34517515849). 2. `chat-channels.integration.test.ts › returns ephemeral guidance for exact Slack controls in channels without creating tasks or actions` — a `provider_effect` row was read before its async settlement reached `processed` (run 34497348802). 3. `runner-connection-eval-fixtures.test.ts › resets paired attempts…` — the fixture's `TRUNCATE companies CASCADE` was chosen as a deadlock victim (40P01) against the helper app's own background sweeps (run 34497348802). **Expected behavior** Verify runs fail only for real regressions. A momentary-contention 503 on a duplicate redelivery, an in-flight settlement row, and a deadlock-victim reset are all recoverable states the code handles by design. **Steps to reproduce** Run the three tests under a loaded 3-shard release-verify split; the timing assertions flake. Under `pr-trusted`'s lighter shards they usually pass, which is why the PRs that introduced them were green. **Paperclip version or commit** `master` at `d1ba17eec`. ## What Changed - The duplicate-redelivery assertion retries on 503 the way Slack itself would (bounded, 250 ms apart), then asserts the 200 and the unchanged dedup invariants: duplicate count increments, still exactly one issue. - The channel-controls settlement read is wrapped in a bounded `vi.waitFor`, the same pattern the file's durable-receipt paths already use. - The runner eval fixture retries its TRUNCATE on Postgres error 40P01, bounded at five attempts, and rethrows anything else. ## Verification - All three run green locally: the two chat tests via `-t` filters, the eval fixtures file in full (6 tests). - Each change is assertion-shape only; no product code is touched. - Observation for a follow-up, not this PR: the chat integration file costs ~15 s transform + ~28 s import per vitest worker before any test executes — splitting it would give back real shard time. ## Risks - Low risk: the retries and waits are bounded, so a genuine regression (permanent 503, settlement that never lands, persistent deadlock) still fails within the same timeouts as before. ## Model Used - Claude (Anthropic), model ID `claude-fable-5` (Claude Fable 5), extended thinking, tool use via Claude Code CLI. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
d1ba17eeca |
fix(adapter-utils): fail fast when the sandbox control channel is lost mid-turn (#13158)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter utilities run agent turns and report their results to the control plane > - A lost sandbox control channel can leave an agent turn without a result > - The host then waits for the full adapter timeout instead of reporting the loss > - This pull request adds a push loss signal and a bounded host wait > - The benefit is a prompt failure terminal when the agent stops answering ## Linked Issues or Issue Description **What happened?** A sandbox control channel loss during an Agent Client Protocol turn left the host waiting for the four-hour adapter execution timeout. **Expected behavior** The host should detect the terminal channel loss, stop the turn, and report a safe failure without waiting for the agent. **Steps to reproduce** 1. Start an Agent Client Protocol turn through a sandbox adapter. 2. Close the duplex control channel while the turn remains active. 3. Observe the host response before the adapter timeout expires. **Paperclip version or commit** Test the pull request commit set at `10b6bbc5525a79fd575298607dd5a25ae448fc8a`. **Deployment mode** The change applies to sandbox-backed adapter execution. ## What Changed - Add `onLoss(listener)` to the duplex bridge handle. - Register the loss listener at turn start and read losses latched before turn start. - Cancel the turn on loss and arm a 30-second host deadline. - Close the stream locally when the deadline wins and create a host terminal. - Derive the public error from the closed `DuplexLossReason` enum. - Add tests for loss order, cancellation, timeout, and safe error output. ## Verification - Run `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Run `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts -t "run-disposition seam"`. - Confirm that the full pull request workflow passes. ## Risks The new deadline changes a lost-channel path from a long wait to a host-built failure after 30 seconds. Orderly completion keeps its existing behavior. The deadline race against a pending `turn.result` has no direct test. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The runtime does not expose a more specific deployment version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
60ee13a0f7 |
feat: allow operator UI snippets on Cloud instances (#13168)
Adds an optional Cloud-only HTML snippet so operators can load Plain’s standard chat bubble. **6 files, 16 implementation lines added; 102 additions including tests and docs.** ## Thinking Path > - Paperclip serves Cloud and self-hosted users. > - Closed beta users need a way to report problems. > - Plain provides a ready-made chat widget. > - Cloud operators can load it through a generic deployment setting. > - Self-hosted instances ignore that setting. ## Linked Issues or Issue Description **Subsystem affected** Server-served UI HTML. **Problem or motivation** Enable a chat bubble in Cloud without adding a support feature to the React app. **Proposed solution** Insert trusted `PAPERCLIP_CLOUD_UI_SNIPPET` HTML before `</body>` when the existing Cloud-managed predicate is true. The setting is off by default. Related Cloud-gated integration: #12190. ## What Changed Review the [final diff](https://github.com/paperclipai/paperclip/pull/13168/files) in this order: 1. `server/src/cloud-ui-snippet.ts`: the eight-line Cloud gate and HTML insertion. 2. `server/src/static-index-html.ts` and `server/src/app.ts`: apply it to static root/index, SPA routes, and Vite HTML. 3. Two test files and `doc/cloud-ui-snippet.md`: boundary checks and setup instructions. React UI, customer identity, and database behavior are unchanged. The existing feedback flag remains. Plain chat is anonymous; no Paperclip name, email, or organization is supplied. ## Verification - **Greptile: 5/5**, no actionable findings, reviewed commit `04bb44515`. - **[CI passed](https://github.com/paperclipai/paperclip/actions/runs/34535763243)**, including build, typecheck, server tests, and end-to-end tests. - Local: six focused tests, full typecheck, and build passed. The full local suite has not produced a final result; CI is the completed full verification. - Staging deployment and live chat testing remain to be done. ### Staging setup Set **one server environment variable**, `PAPERCLIP_CLOUD_UI_SNIPPET`, to: ```html <script> (function(d) { var script = d.createElement('script'); script.src = 'https://chat.cdn-plain.com/index.js'; script.onload = function() { Plain.init({ appId: 'liveChatApp_01M26J213F6RR53YRARZVAFCZZ' }); }; d.head.appendChild(script); })(document); </script> ``` This is the public staging app ID. **No API key or signing secret is needed.** Deploy to staging and restart the app with this setting. Test `/`, `/index.html`, and an organization dashboard; send a message and confirm a support reply returns. Production rollout is separate. [Plain embed docs](https://www.plain.com/docs/product/channels/chat) · [Configuration and rollback](https://github.com/paperclipai/paperclip/blob/04bb4451510b124f21b6771b7a8163faad61d8ec/doc/cloud-ui-snippet.md) ## Risks Only trusted operators should set this value. The HTML is public and scripts execute in the app origin; do not include secrets or user-provided HTML. Plain owns the anonymous browser session, with no Paperclip account-switch integration. To roll back, unset the variable, restart, and refresh open tabs. ## Model Used OpenAI Codex (GPT-6), with repository inspection and code execution. Exact runtime model identifier and context size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4042eb1c48 |
test(release-smoke): follow the connect-step source question and the first-task chat (#13166)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release pipeline promotes canary → nightly → beta → stable, and the nightly lane is gated by the Docker release smoke, a Playwright walk of first-run onboarding against the exact published artifact > - Onboarding changed twice since the smoke was last updated: the connect step now opens as a model-source question (#12796, #12801), and the seeded first task now opens as a chat with the lead that deliberately creates no run until the user answers (#13068) > - The smoke still waited for an immediate "Connect" button and then polled for an assignment-triggered heartbeat run, so it failed every scheduled nightly since 2026-09-03 and blocked all nightly and beta promotions > - This pull request updates the smoke to follow the current arc: pick the Claude source tile, press Connect, launch, then assert the seeded chat greeting, the opening question card, and the absence of heartbeat runs > - The benefit is a release pipeline that can promote current master again, with the smoke asserting the product's current contract instead of a removed one ## Linked Issues or Issue Description **What happened?** The scheduled nightly lane of `release.yml` has failed every night since 2026-09-03. The `smoke_nightly / smoke` job fails in `tests/release-smoke/docker-auth-onboarding.spec.ts` at `expect(connectButton).toBeVisible()`. No nightly has published since `2026.902.0-nightly.0`, so no beta can promote recent master. **Expected behavior** The release smoke follows the current onboarding arc and passes against a healthy published artifact. The nightly lane promotes the newest green canary each night. **Steps to reproduce** 1. Run `PAPERCLIPAI_VERSION=2026.910.0-canary.5 SMOKE_DETACH=true ./scripts/docker-onboard-smoke.sh`. 2. Run `pnpm run test:release-smoke` against the container with the previous spec. 3. The spec times out waiting for a "Connect" button. The step now shows a model-source tile row first, and after launch the seeded task is a chat with no heartbeat run. **Paperclip version or commit** Reproduced against published `paperclipai@2026.910.0-canary.5`; spec updated on current `master`. ## What Changed - The spec answers the connect step's model-source question: it asserts the "Connect a model" heading, picks the Claude tile from the "Model source" radiogroup, and only then waits for the "Connect" footer button (#12796, #12801 rebuilt the step around that question). - The spec replaces the assignment-run poll with the first-task chat contract from #13068: it asserts the deterministic greeting ("Welcome to Paperclip!"), the opening question card ("What would you like to do?"), and that the lead has zero heartbeat runs, because launch must not wake the assignee before the user answers. ## Verification - Launched the CI harness locally: `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.910.0-canary.5` (the newest canary, the one the next nightly would promote). - `pnpm run test:release-smoke` against that container: 1 passed. - The previous spec against the same container reproduces the CI failure mode first (Connect-button wait), and after the connect-step fix, the run-poll failure — both match the nightly logs. ## Risks - Low risk: the change touches only the release smoke spec. If onboarding's copy for the greeting or the question card changes, the smoke fails loudly at that assertion, which is this suite's job. ## Model Used - Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |