mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
codex/slack-managed-setup
4943
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
206f08351f |
fix(slack): finish avatar setup after manifest recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
34ffc905ea |
fix(slack): recover rejected grants and consume creation confirmations
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
61d199b086 |
test(slack): follow extracted setup stage contract
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b5322b4a9a |
test(slack): isolate second-agent identities across repeated runs
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a90d6c2384 |
fix(slack): persist reinstall intent across provider failures
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
483c466008 |
fix(slack): make manager refresh and installation recovery resumable
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
26622523c6 |
test(slack): cover grant revocation during OAuth and fixture consent return
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
20534f871b |
fix(slack): keep grant revocation enforced during install callbacks
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
785ae16699 |
feat(slack): add gated managed setup alongside customer-owned apps
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
835a022936 |
perf(server): wake delivery queues from transaction-aware work signals (#15625)
## Thinking Path > - Paperclip manages AI-agent work and its delivery to users and agents. > - Five delivery queues scan the database even when empty. > - Each queue already has durable rows; empty polling wastes idle hosting capacity. > - Producers can register intent inside their existing database transaction. > - Central transaction tracking can wake consumers after the outer commit without extra caller wrappers. > - A lost response needs reconciliation against PostgreSQL before an empty queue is safe to leave idle. > - This change uses one scheduler for outstanding work and lets settled, empty queues go quiet. ## Linked Issues or Issue Description Supersedes the closed #15614. Related idle-safety work: #15522 and #15599. **What existing behavior does this improve?** Delivery scheduling for feedback exports, chat completions, connection continuations, question answers, and tool-action receipts. **Current behavior** Feedback scans every five seconds. Four delivery queues scan on the heartbeat interval, normally 30 seconds. Startup and some event paths also scan them. **Proposed behavior** Each producer awaits one intent registration before its queue write. The existing `createDb` transaction boundary tracks nested savepoints and wakes the consumer after outer settlement. Startup scans restore durable work. Pending rows, failures, and unresolved transactions retain retry deadlines. Empty, settled queues have no timer or recurring scan. **Reason and benefit** Normal delivery starts after commit. Central tracking removes caller-specific post-commit plumbing. PostgreSQL transaction status resolves lost responses without new tables or a permanent uncertainty latch. **Breaking changes** No API or schema changes. The notification scope is one DB owner and its dedicated child pools. Direct SQL, separate roots, and other processes require an explicit wake/recovery integration. This path requires PostgreSQL 14+ transaction-information functions; embedded PostgreSQL uses 18. ## What Changed - Add one named-deadline scheduler and app-owned coordinator for five existing queues. - Instrument `createDb().transaction()` centrally, including nested savepoints. Dedicated child connections share their owner's signal scope. - Register work before the five queue insert paths. The first registration obtains the outer XID; unrelated transactions issue no extra SQL. Tool receipt insertion gains a small transaction around its existing insert. - Query `pg_xact_status` only for rejected, tracked transactions. Probe before scanning the queue. Retain the idle hold while PostgreSQL still reports an in-progress transaction or the probe fails. Release it after settlement and reconciliation. - Remove unconditional delivery recovery scans and caller-specific completion post-commit actions. Retain the existing task-scoped completion activity fast path and durable claims. - Coalesce wakes without postponing an earlier deadline. Retry outstanding rows and failures; suppress dispatch during idle drain while allowing transaction and queue reconciliation; suppress both during warm standby. Cancel deadlines and await active delivery sweeps on shutdown. - Serialize feedback flushes. Vote routes return after saving. Uploads have a 30-second deadline and shutdown cancellation; unfinished exports remain recoverable. - Document the writer contract and update the dated sleep inventory. ## Verification - Final head: `a4609956841107ad60c4cb50fe2acb05636b1363`. - Passed: all final-head hosted checks, with 54 successes and two intentional skips. The full test matrix, typecheck, build, canary, and aggregate verification passed in [run 37940948440](https://github.com/paperclipai/paperclip/actions/runs/37940948440). - Passed: Greptile 5/5 on the same head, with a successful check run and zero unresolved review threads. The branch is mergeable. - Passed locally after the final rebase: 50 coordinator, scheduler, and startup tests; server typecheck; server build. - Passed locally before the final import-only rebase: repo-wide typecheck and build; 237 PostgreSQL producer/delivery and database signal tests; feedback, native-question, and idle-safety regression tests. - Real PostgreSQL tests hold transactions open after an injected client response loss, then commit or abort. They verify that an empty scan cannot clear an in-progress transaction. Another test terminates a real backend and verifies recovery without replay. - Coordinator tests cover an hour with no empty-queue timers, pending/error retries, worker replacement, concurrent writes, earliest-deadline preservation, rollback during idle drain, committed work during drain, standby, and shutdown. - Local `pnpm test:run` was started and stopped after embedded PostgreSQL startup failures appeared. It did not finish and is not reported as a pass. Later local database reruns had skipped suites; those skips are not claimed as verification. The earlier PostgreSQL tests did execute and pass, and the final hosted full matrix passed. ## Risks - One XID query is added per transaction that registers delivery work. Nested transactions share that ID. Registration must be awaited before writing; new enqueue paths must follow this contract. - Work signals are local to a DB owner and its dedicated child pools. Independent roots, direct SQL, or other processes are not observed. Startup scans recover already committed rows, but do not fence late transactions from a previous process. Cross-process ownership and database failover require separate work. - A database outage or genuinely unresolved transaction keeps an idle hold until status can be reconciled. No timer clears uncertainty by assumption. - These changes remove five empty polling loops. They do not implement whole-instance sleep, database teardown, or an external waker. The earliest deadline covers only the registered queues. - Waiting tool reviews retain their existing retry cadence until their receipts can be delivered. - Feedback vote responses no longer wait for the remote upload. Sharing consent and the saved vote response are unchanged. - Real hosting-provider sleep and cost savings have not been measured. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository inspection, code editing, and command execution. The runtime did not expose the exact model revision or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ca4b7c4c0 |
refactor(heartbeat): extract run cancellation and cleanup (#15685)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service admits, executes, and settles agent runs. > - Earlier extractions separated workspace, preparation, state, retry, recovery, queue, and lifecycle handling. > - Cancellation and resource release still live inside the main service. > - These operations share execution ownership and saved-comment continuation rules. > - This pull request moves them into one run-control module and adds boundary tests. > - Reviewers can inspect cancellation and cleanup separately from adapter execution. ## Linked Issues or Issue Description **What existing behavior does this improve?** The structure and testability of heartbeat cancellation and cleanup. **Subsystem affected** server/ — orchestration services. **Current behavior** Stop, pause and budget cancellation, environment lease release, and saved-comment delivery live in `heartbeat.ts`. **Proposed behavior** Keep these operations in `server/src/services/heartbeat/run-control.ts`. Bind their database and existing callbacks with `createHeartbeatRunControl`. **Reason and benefit** This removes 1,115 lines from `heartbeat.ts`, leaving 10,443 lines. The new module has 1,360 lines. Related cancellation and cleanup rules stay together in one file. **Breaking changes** None. Function bodies, service methods, public helpers, status conditions, and side effects keep their behavior. **Additional context** Continues the merged lifecycle extraction in #15672. Related cancellation work: #15212. Searches of open and closed issues and PRs found no duplicate run-control extraction. This does not add a roadmap feature. ## What Changed - Move run cancellation, pause and budget cancellation, lease release, issue-lock release, and saved-comment resumption into `heartbeat/run-control.ts`. - Supply lifecycle, queue, recovery, and cleanup callbacks through an explicit dependency interface. - Keep executor and cancellation maps shared by all service instances. Preserve callback construction order with forwarding functions. - Keep the existing helpers available from `heartbeat.ts`. - Add 15 tests for inert construction, cancellation gates, shared Stop barriers, failed termination, warm-resource retention, pending wake batches, project scope, and stale budget enforcement. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Before extraction: **298 tests passed across six existing cancellation and cleanup suites**. - After extraction: **313 tests passed across seven files**, including the 15 new module tests and real PostgreSQL coverage. ```sh pnpm exec vitest run server/src/services/heartbeat/run-control.test.ts server/src/__tests__/heartbeat-native-runner-cancellation.test.ts server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts server/src/services/run-cancellation.test.ts server/src/services/adapter-execution-control.test.ts server/src/services/explicit-native-continuation.test.ts server/src/__tests__/native-cancellation-request.integration.test.ts ``` - Wider queue, recovery, reassignment, comment delivery, and native cancellation verification: **139 tests passed across nine files**, including the final module tests. The recovery suite first hit its 20-second database setup timeout while local builds were running. After removing verified orphaned PostgreSQL headers, its isolated run passed **380 of 381 cases**. One teardown assertion exceeded its one-second wait for a successor to settle; that exact case passed in a filtered rerun. Combined focused coverage is **818 distinct passing tests across 16 files**. The module suite appears in both runs. ```sh pnpm exec vitest run server/src/services/heartbeat/run-control.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-native-cleanup-admission.test.ts server/src/__tests__/heartbeat-lock-release-on-reassignment.test.ts server/src/services/heartbeat/queue.test.ts server/src/services/heartbeat/recovery.test.ts server/src/services/heartbeat/retries.test.ts server/src/services/heartbeat-stop-metadata.test.ts server/src/services/native-runtime/native-cancellation-request.test.ts ``` ```sh pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts -t "does not adopt unrelated queued comments for a non-coalescing recipient after Stop" ``` - Full `pnpm -r typecheck` and `pnpm build` passed on head `b26efd5f9de390ef8a675f84980f49f7fb7799c2`. - Full local `pnpm test:run` was started on this head and stopped after the complete CI suite passed. It did not finish and is not counted as a full local pass. - All **54 final-head checks passed**; two optional Storybook checks were skipped. This includes the full general and serialized test matrix, browser tests, typechecks, build, Runner verification, and canary dry run. The complete recovery suite passed in CI. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37945389261). - Greptile completed on head `b26efd5f9de390ef8a675f84980f49f7fb7799c2` with **5/5 and no actionable findings**. There are no review threads or merge conflicts. - Structural comparison confirms all 21 moved function bodies, the cancellation options type, remaining service logic, and all 140 existing public exports are preserved. ## Risks Moving closures can break callback construction order or shared Stop barriers. The factory starts no work and receives the same process-wide maps and existing callbacks. Tests cover concurrent Stop callers and failed termination. Native authority, status predicates, provider receipts, cleanup gates, and budget enforcement keep their original bodies. This PR has no schema or API changes. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
49bbfe3092 |
fix(server): register native runner transport in HTTP startup (#15682)
Register native Runner transport in the shared HTTP server constructor. Add startup regression tests and custom-launcher guidance. Verified with live OpenCode/OpenRouter tasks and complete CI coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1009.0-canary.4 |
||
|
|
0b80ea17ef |
refactor(heartbeat): extract run lifecycle handling (#15672)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service admits, executes, and settles agent runs. > - Earlier extractions separated workspaces, preparation, state, retries, recovery, and queue admission. > - Run status, progress, liveness, and completion handling still live in the main service. > - These operations share status predicates, event writes, and completion policies. > - This pull request moves them into one lifecycle module and adds boundary tests. > - Reviewers can inspect lifecycle policy separately from adapter execution. ## Linked Issues or Issue Description **What existing behavior does this improve?** The structure and testability of heartbeat run lifecycle handling. **Subsystem affected** server/ — orchestration services. **Current behavior** Run status writes, progress, events, liveness, completion handoffs, issue-comment finalization, and runtime settlement live inside `heartbeat.ts`. **Proposed behavior** Keep these operations in `server/src/services/heartbeat/run-lifecycle.ts`. Bind the database and reporting, recovery, and wakeup callbacks with `createHeartbeatLifecycle`. **Reason and benefit** This removes 2,682 lines from `heartbeat.ts`, leaving 11,558 lines. The lifecycle module has 2,982 lines and keeps related policy together in one file. **Breaking changes** None. Existing function bodies, service methods, public exports, status conditions, and side effects keep their behavior. **Additional context** Continues #15626 and the earlier merged heartbeat extractions. A search of open and closed issues and PRs found no duplicate lifecycle extraction. This does not add a roadmap feature. ## What Changed - Move status transitions, run events and progress, liveness classification, completion handoffs, plan-resume failure reporting, issue-comment finalization, and runtime/cost settlement into `heartbeat/run-lifecycle.ts`. - Supply reporting, recovery, and admission callbacks through an explicit dependency interface. Preserve construction order with forwarding callbacks. - Keep adapter execution, cancellation, and terminal telemetry reporting in the main service. - Add 16 tests for inert construction, continuation gates, late progress, status races, native ownership, usage metadata, revoked wakes, event sequencing, and redaction. - Update the existing publication-order source assertion to read the extracted event writer. Keep its persist-before-publish check. - Document the lifecycle module boundary in `doc/DEVELOPING.md`. ## Verification - Before extraction: 87 tests passed across seven existing suites. - After extraction: 103 tests passed across eight files, including the 16 new module tests and real PostgreSQL coverage. ```sh pnpm exec vitest run server/src/services/heartbeat/run-lifecycle.test.ts server/src/__tests__/heartbeat-run-status-payload.test.ts server/src/__tests__/heartbeat-run-event-sequencing.test.ts server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/__tests__/run-liveness.test.ts ``` - Additional recovery, cancellation, plan-resume, summary, and handoff coverage: **595 tests passed across 14 files**. The module suite appears in both runs. With the publication-order suite below, combined focused coverage is **694 distinct tests across 22 files**, with no skips. ```sh pnpm exec vitest run server/src/services/heartbeat/run-lifecycle.test.ts server/src/services/heartbeat/recovery.test.ts server/src/services/heartbeat/retries.test.ts server/src/services/heartbeat/queue.test.ts server/src/services/recovery/successful-run-handoff.test.ts server/src/services/recovery/review-path-recovery.test.ts server/src/services/heartbeat-run-runtime-status.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-native-runner-cancellation.test.ts server/src/__tests__/heartbeat-native-cleanup-admission.test.ts server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/heartbeat-run-summary.test.ts server/src/__tests__/heartbeat-context-summary.test.ts ``` - The publication-order and lifecycle suites passed together: **28 tests across two files**. The first CI run exposed an assertion that still read the old event-writer source path. That assertion now reads the lifecycle module. ```sh pnpm exec vitest run server/src/services/chat-publication-reconciliation.test.ts server/src/services/heartbeat/run-lifecycle.test.ts ``` - After resolving the import conflict with the Pi Runner integration and rebasing onto `57e977be7`: **44 tests passed across four files**, including the updated source assertion, lifecycle module, run-status payloads, and accepted-plan workspace refresh. ```sh pnpm exec vitest run server/src/services/chat-publication-reconciliation.test.ts server/src/services/heartbeat/run-lifecycle.test.ts server/src/__tests__/heartbeat-run-status-payload.test.ts server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts ``` - Full `pnpm -r typecheck` and `pnpm build` passed on final head `bca87df089c541450932f6b678ca776942d267b2`. Full local `pnpm test:run` was restarted on this head and stopped after the entire CI suite passed. It did not finish and is not counted as a full local pass. - The earlier full local run had one transient public MCP Cloud bootstrap failure. All **82 public MCP tests passed in isolation**. The first local run was stopped before the source-assertion fix and rebase. - All **54 latest-head checks passed**; two optional Storybook checks were skipped. One general-test shard passed all 1,314 tests but failed during runner temporary-directory cleanup with `EACCES`. One rerun passed. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37940997283). - Greptile completed on final head `bca87df089c541450932f6b678ca776942d267b2` with **5/5 and no actionable findings**. There are no review threads or merge conflicts. - Structural comparison confirms 63 module function bodies and ten declarations match the originals. Remaining service function bodies and all 140 public exports are preserved. ## Risks Moving closures can break callback construction order or status settlement. The factory receives explicit callbacks and starts no work during construction. Compare-and-set predicates, native ownership holds, revoked-wake guards, event sequencing, redaction, and completion policies keep their original bodies. This PR has no schema or API changes. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1009.0-canary.3 |
||
|
|
57e977be72 |
feat: integrate Pi 1.0 into the experimental Runner (#14921)
## Thinking Path > - Paperclip manages AI agents and their work. > - The experimental Runner owns provider processes and durable sessions. > - Pi needs working task execution and human controls. > - The five-PR stack must preserve changes already on master. > - Each layer now carries the complete integrated source for a safe sequential fallback. > - This PR belongs to native GitHub stack #15602, ending at #14956. ## Linked Issues or Issue Description Refs #14436, #14631, #14743 and #14956. Ship Pi 1.0 through the experimental Paperclip Runner. The five PRs are #14921, #14922, #14923, #14924 and #14956. The user authorized the complete merge after checks pass. Existing `pi_local` execution is unchanged. Accounting and wider provider/platform qualification remain deferred. ## What Changed - Recover missing final replies after workspace finalization changes owners, using accepted-turn evidence without rerunning work or granting external-chat publication. - Preserve the admitted Pi instruction root across warm runs, while retaining changed-root rejection. - Give Pi a bounded 15-second default shutdown grace so stop, drain acknowledgement and durable suspension can complete. Explicit deadlines and other providers retain their existing behavior. - Integrate the Pi 1.0 runtime and master contracts. - Use Pi profile 22. Preserve explicit caller-selected models and exact native thinking levels. Keep Pi's wrapper, helper, extension and question/control behavior unchanged from the qualified profile-19 runtime. - Preserve master's Dot lifecycle and consent fields, configured task environment, status guards and current Codex/Claude dependency versions. Cursor stays qualified. Copilot stays pending; profile 17 binds the changed shared protocol validation sources. - Exclude general AWS IAM credentials from Pi static/custom provider bindings and selected task projections; preserve the provider-scoped Bedrock bearer key. Profile 21 is retained as historical provenance. Rust and cloud install probes use the current declaration. - Patch bundled brace-expansion 5.0.9 to the exact official 5.0.12 payload. Pin the patch and complete runtime closures. Include the patch in normal installed setup tooling. Keep the upstream Pi shrinkwrap as provenance and permit only this exact security correction. - Include current attestation files in the Docker build context. Keep the repository lockfile unchanged from master. CI and private image builds resolve manifest changes before their frozen installation. ## Verification - Full local `pnpm -r typecheck` passes, including Runner Rust, server and UI. Focused integration checks pass: 194 Runner admission/environment tests, 63 profile/credential tests with one expected skip, 152 Dot/UI configuration tests, and Pi transcript/notice tests. - Full local `pnpm build` passes on the final source. - Fresh final-source checks pass: all 698 Rust workspace tests (32 binaries), 156 credential/profile/controller tests with one expected skip, Runner TypeScript typecheck, and 20 package/setup/sandbox tests. - The profile-21 Pi materializer passes on the native host with the official pinned Node 24.21.0 and its npm. It verifies all 150 locked packages, the patched dependency and the exact closure. Setup/package bundle tests and UI token gates pass. - The old hashes were reproduced for all three supported targets before calculating the patched graph. New closure hashes are darwin-arm64 `282022db10150c6632b3444df421342e7d534bdf5d5fb1097a2e79d0625a2bcf`, darwin-x64 `64e251e19009f755c0b04f73ce2138246faab71a961b0f13d75ebfcc34bef12e`, and linux-x64 `713b1fdff42fb56a1518bdc084f181d70bee8ebadc3e4b1d76321ed9108c8410`. Independent native platform execution is separate from graph identity reproduction. - Historical cloud qualification remains unchanged: all seven core cases pass on shipping source `10dc43c9ec65d88c2f782d62afb296d09494f215`, harness `1a4408a48cfb5a1f094a311141c257c92cd7a893`, image `sha256:5b3a775b383591bda1b0c1889e509acc70ce7f37c53f09733c81d59037f02280`, and accepted Sonnet 4.6/low fixture. All 215 canonical files and all seven cleanup checks pass independent verification. These are profile-19 results and are not relabeled as fresh profile-22 runs. - Current Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [The readiness plan](https://github.com/paperclipai/paperclip/blob/codex/pi-production-readiness/doc/plans/2026-10-02-pi-production-readiness.md) preserves campaign and failed-attempt provenance. - Merge only after every PR's current-head CI and fresh review pass. Linux CI covers the full suites, build and browser tests. The local embedded Postgres API-authority suite cannot start on this macOS/Node 26 host, so Linux CI must confirm that suite. ### Fresh profile-22 core qualification — 2026-10-08 All seven accepted core cases pass canonically on Pi profile 22, with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. This model is a fixture; production accepts the caller's explicit Pi provider/model. Runtime/install source: `3241a992f2a7703e59e97ed0fd3e5d6405de4401`. Frozen accepted harness: `1a4408a48cfb5a1f094a311141c257c92cd7a893`. Immutable cloud image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:506f22db7edd78f37c0c40bec1cc084af1850455026dbf467194bfbb8fcef141`. Pi digest: `sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`. [Hosted Linux image and clean-install verification](https://github.com/paperclipai/paperclip/actions/runs/37868328023) passes, including all 20 source-bound archives, normal CLI/Pi setup, companion import and the production pack reader. This exact installation source includes the latest master integration and the corrected Pi warm instruction-root fence. Full local typecheck/build and current-head hosted CI verify the final stack. All 13 focused real-root regressions pass. The full local executor suite passed 662 tests; 15 database tests could not start the Mac embedded PostgreSQL service. Hosted Linux CI passes the full required verification and E2E checks. These fresh results keep their own source identity; profile-19 results remain historical. | Core path | Canonical campaign | Retained archive SHA-256 | | --- | --- | --- | | File edit, validation, download and Done | `pi-core22-replyfix-0-1791511228` | 23 files; `a473e8603a3dd4737863291f8d3d1e392391f0b16d433c3e0e0e9d8baf7a97b0` | | Pending question and controller restart | `pi-core22-replyfix-1-1791511376` | 33 files; `6b829c4eb74e1f32a89c692a4ae7130dbfc1c6d3cf13915effe2103d9e242c8e` | | Three-turn session/process/workspace continuity | `pi-core22-replyfix-2-1791511587` | 23 files; `7a87021f8f9a3fdd3c58bb4467f8d82c635e3ea4795d6e75f144d9aa14818df8` | | Four typed questions and browser reconnects | `pi-core22-replyfix-3-1791511881` | 42 files; `9e31755252be1f4f9cb0626c984c142d4d1ae5f5bee3a7af08444db8d12c280a` | | Plan approval and completion | `pi-core22-replyfix-4-1791512031` | 22 files; `a0383ce1aab38e7b5a25ce0e9dd3bebea5c037ebd96ae6b29dae19015da2ae2c` | | Same-turn steering and permission denial | `pi-core22-replyfix-5-1791512261` | 39 files; `c929b8c7070f0b66aedc17e65ca46e6beab1e363926ac9f7e2a75fb250f05949` | | Stop during pending permission | `pi-core22-replyfix-6-1791512390` | 33 files; `7f0a58ae0f4d5bfc76149435f4e322537089c5bd16e7ffe9b5ad71f10a621a07` | All 215 canonical files (28714587 bytes) are independently hash-verified. All seven cleanup grades pass, with no owned runtime process or temporary root after each case. Automatic retries are zero. The owned cloud host stopped normally after retention. The prior profile-22 warm attempt remains failed and separately retained: archive SHA-256 `1e54eba5ec72b50cee1534b23d1d1d4f21a090006b8a64501ba70db972abfde5`. Its original canonical classification is preserved. Diagnosis reproduced a product bug comparing an agent-files root against an unset checkpoint-only field. The fix stores the admitted physical root separately from the adopted per-run collection capability. The real-root regression fails before the fix and passes afterward, including rejection of a changed physical root. Fixture, grader, model and all seven accepted case IDs are unchanged; this fresh campaign tests final-reply publication after file registration first. The intermediate restart attempt also remains failed and retained: archive SHA-256 `5dcaefdf1d17cf4cd54fd4cf810f45e736667392339b8ce7caf08bb4e225277f`. Its original canonical classification is preserved. Pi resumed, wrote the verified answer and completed its task; exact runner suspension was proven, but idle stop consumed about 5.2s and left under 3s for the drain acknowledgement. The Pi-only default shutdown grace is now 15s, preserving a full 5s drain round trip and a finite suspension reserve. Explicit caller deadlines, other provider defaults, literal drain receipts and exact suspension identity checks remain unchanged. The timing regression fails before this correction and passes afterward; all 18 focused settlement tests and Runner typecheck pass. The final-source file attempt is also preserved as failed (`candidate_failure`), archive SHA-256 `db6767b6773ea618997927ac77bdb005a5ac81492c7b9c0ffbc900449f829bc9`. Native edit, validation, exact downloadable artifact and Done/succeeded all passed, and the exact final reply was durably recorded. A workspace recovery owner completed before the live heartbeat reached presentation, leaving that reply absent from task chat. Recovery now materializes only a completed final reply from the accepted turn of an ordinary internal Done task, preserving issue/run/contract binding, suppression, external-chat authorization and same-run deduplication. The database regression covers the generated file-preparation receipt, suppression, unapproved external continuation and replay. Server typecheck and all 49 response-selection tests pass; hosted Linux verifies the database regression because embedded PostgreSQL cannot start on this Mac. The delayed-final-answer database regression passes on [the final root-source Linux server shard](https://github.com/paperclipai/paperclip/actions/runs/37868262553/job/113628594152), alongside 1,108 passing tests. The first root Runner shard had one unchanged durable-resume test exceed its 5-second timeout; the identical top-source shard and the isolated exact test passed. One rerun of that failed job and its required aggregate passed without source or test changes. The original failed job log and the single-rerun receipt remain retained. ### October 9 merge verification Current merge head: `5a8fe63512a7166aaef5cf50065a25008aa8b44b`. All current-head checks pass, including `ci / verify` and `ci / e2e`; exact-head Greptile review is 5/5 with no unresolved threads. Current master conflicts are resolved. The user authorized the maintainer override of the code-owner review gate after these checks. The seven retained live core cases remain bound to source `3241a992f2a7703e59e97ed0fd3e5d6405de4401` and its recorded cloud image. ## Risks - The security correction changes the dependency closure and profile identity. Old sessions must reopen on the new profile. Exact identities and credential bindings fail closed. - The runner remains experimental and requires explicit selection. Legacy Pi Local is unchanged. Caller model IDs pass through; the E2E model is a fixture. - Accounting and the broad platform/provider matrix remain deferred. This merge does not publish a release or deploy a service. ## Model Used OpenAI GPT-6 through Codex assisted with reasoning, repository inspection, editing and tool use. The exact serving ID and context window are not exposed in this session. Final live qualification uses Pi 1.0.0 with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low thinking. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4fb1ab53a8 |
chore(lockfile): refresh pnpm-lock.yaml (#15673)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.1009.0-canary.2 |
||
|
|
6f9d0a56ba |
fix: resolve installed Codex and preserve npm host dependencies (#15555)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Installed agents need their own runtime dependencies. > - Native Codex startup and browser login could depend on a global CLI. > - npm bundles do not inherit workspace dependency overrides or patches. > - Codex native packages must come from npm for the consumer host. > - This PR fixes executable resolution and the required npm dependency layout. > - Usable installed CLI versions can run without a numeric version gate. ## Linked Issues or Issue Description Refs: #15422. This is a prerequisite for the later Codex-default change. **What happened?** Packaged native Codex startup and browser login could require a global Codex CLI. The server's vendored runner lacked its own declared Codex bridge. A bundled wrapper could retain producer-host binaries or select the legacy adapter's separate platform version. **Expected behavior** Prefer the installed Codex executable. Accept usable older or newer versions. Use the selected host's PATH when the dependency is absent. Keep the patched JavaScript graph. Let npm install official native packages for the consumer platform. Grant sandbox reads only to the selected vendor resources. **Steps to reproduce** 1. Prepare a clean npm installation without a global Codex command. 2. Start native Codex or its browser login. 3. Inspect the selected executable, installed host package, and sandbox resource paths. **Paperclip version or commit** Frozen master base: `1881894973a2b25838d8abed9bd8aeebc3af4441`. **Deployment mode** Source and packaged self-hosted installation. ## What Changed - Share installed Codex command resolution across native startup, direct evals, and browser login. Accept usable version differences. Preserve explicit commands, recorded sessions, remote host boundaries, and the Linux ARM64 legacy login path. Return safe errors for missing executables and failed login terminals. - Declare the server's Codex bridge. Retain the patched JavaScript dependency graph in npm packages. Strip Codex native payloads. Declare official optional host packages on the published server manifest so native and legacy versions remain separate. Preserve consumer esbuild platform dependencies. - Adjust resource lookup for npm's separate platform packages. Retain package identity, path containment, resolver ownership, and narrow sandbox reads. Keep exact version and digest checks for explicit ACPX artifact qualification. - Extend existing packaging, installed-consumer, login, selection, recovery, and integrity tests. Document the retained behavior. The diff is now 23 files, 1,709 additions, and 53 deletions. The prior diff had 35 files and about 3,700 changed lines. Removed work is preserved on `codex/runner-packaging-full-snapshot` at `6f3060beaa842ffcee21af058370d3cab5f571e9`. Removed from this PR: release workflows and assembly, provider-pack changes, Docker materialization, Git installer changes, extra login HOME/working-directory isolation, and unrelated CI fixture repairs. This PR does not change agent defaults, stored runner choices, UI, schema, or provider qualification. ## Verification - Final candidate: `dde37d7ed0121c10b60b8801eda80f8dc17909ad`. [All 47 ordinary CI jobs passed](https://github.com/paperclipai/paperclip/actions/runs/37842288835) on attempt 1, including typecheck, tests, build, E2E, runner checks, and the installed-consumer canary. [Fresh Greptile review](https://github.com/paperclipai/paperclip/pull/15555#issuecomment-6059581546) is 5/5 on this head; no unresolved threads remain. Human approval is still outstanding. - Reused focused controls passed with Node 24: 71 initial narrowed checks, then 56 affected Codex/selection/eval checks after the relative PATH correction. Runner TypeScript no-emit, syntax, and diff checks passed. The existing fixture reproduces the original relative PATH failure and verifies working-directory selection, empty entries, ordering, and absolute launch. One CLI entrypoint test was initially blocked by sandbox IPC and passed with its existing local socket allowed. Login HOME, config-directory assignments, and working directory match master. - [The actual clean Linux npm consumer](https://github.com/paperclipai/paperclip/actions/runs/37842288835/job/113535523044) passed on the final candidate's CI integration. All 17 Paperclip tarballs omit native Codex payloads. Official npm host packages retain their own integrity and `inBundle=false`. Native Codex selects 0.160.0; the legacy closure retains 0.156.1. Package admission, command leases, narrow native sandbox resources, consumer hooks, preserved lock, and offline lifecycle controls passed. Provider calls were zero. No new workflow is added. - Actual official Codex 0.156.1 passed on the final committed source on macOS ARM64: installed dependency preference, absolute and relative PATH fallback, narrow resource lookup, and app-server initialize/initialized. Executed module hashes match the candidate. One scripts-disabled install, three version probes, and one handshake completed in 26 seconds; owned files and process group were removed. No login, account/model request, or provider task ran. - CI checked out `dfda8e708e87306d22ed735d78bdfcbe770e99b7` on base `65b558180533039a891ee0cd1ccab9988aa79adc`. Its 13 upstream paths do not overlap this PR's 23 paths or alter packaging/Codex inputs. The consumer report records producer `7b8e94c08657b5ddc265946459762340f125726c`, after the existing canary staged a generated-lock-only commit (one file, three insertions). These identities are kept separate; raw generated-lock bytes were not retained. - No local Docker or Rust build, paid provider turn, merge, or deployment. Later PRs must prove live onboarding and production cloud packaging before changing defaults. ## Risks - npm must install optional host dependencies. Missing dependencies still return errors. Paperclip tarballs do not pre-bundle Codex executables. - The repository requires CI-owned lockfile updates. PR CI resolves the changed manifest and stages its own producer lockfile. A raw source Docker build with `--frozen-lockfile` must wait for the existing master lockfile bot to merge its refresh, or use a disposable resolved checkout. No Docker build or deployment is qualified by this PR. - Ordinary Codex startup accepts version differences. Actual protocol or login failures remain errors. Explicit ACPX artifact checks retain their release pins. - The published server delegates platform installation outside its bundled JavaScript graph. Focused negative controls reject unsafe package metadata and paths. The hosted consumer test passed on this candidate’s CI integration. - Existing agents retain their stored runner choices. There is no data migration or automatic upgrade. Release pipeline and platform qualification work remain separate prerequisites for later defaults. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and parallel agents. The exact serving snapshot and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2b688a8fe3 |
feat(connections): add verified MCP providers and setup fixes (#15621)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents governed access to external tools. > - Several provider MCP servers lacked a supported catalog entry or failed during setup. > - Real browser tests identified specific registration, session, and form defects. > - This pull request adds seven catalog entries and fixes the shared paths used by eleven verified providers. > - Users can connect these providers through the existing Apps flow and control each action. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Apps catalog and remote MCP setup, discovery, and action tester. **Subsystem affected** Shared app definitions, the server tool services, and the Apps UI. **Current behavior** Seven providers lack a catalog entry. Airtable can select an advertised client-metadata flow that fails. Calendly rejects the registration name. Tavily needs an initialized session before a call. Firecrawl exposes invalid defaults for optional nested form objects. HTTPS setup can display a callback that differs from the server callback. **Proposed behavior** Add Calendly, Exa, Firecrawl, GSC Wizard, Parallel Search, Tavily, and Windsor.ai. Preserve Airtable, Linear, Make, and PostHog. Use reviewed provider options through the shared MCP and OAuth paths. Display the callback supplied by the server. Omit untouched optional object inputs. **Reason and benefit** Eleven providers passed bounded reads through the real Paperclip browser action tester. This PR includes that verified set and its required shared fixes. Unfinished providers remain outside this change. AgentMail keeps its existing integration. **Breaking changes** No database migration. Existing OAuth credentials and action policies keep their ownership and access rules. Airtable's reviewed method re-registers a retained client-metadata binding through DCR. Tavily's reviewed method initializes sessions before dispatch. The existing customer, managed, and Vercel OAuth gateway paths now refresh on upstream 401 and return `oauth_refreshed_retry_required` (409), requiring an explicit caller retry. They do not replay the rejected call automatically; the next invocation initializes a fresh credential-scoped session when required. Related catalog work: #15545. This branch preserves current master catalog entries and does not add another Google Workspace integration. Open and closed provider PRs and public MCP issues were searched. No duplicate for this verified set was found. The change extends the shipped Connected Apps roadmap item. ## What Changed - Add seven catalog entries with official provider branding and reviewed permissions. - Update Airtable, Linear, and Make metadata and show Make in Apps. - Add authoritative provider source overrides that use the existing generation and review checks. - Add the reviewed Airtable DCR option and compatible Calendly registration names. - Initialize Tavily sessions before discovery and calls, including retained connections. Keep credential-scoped session caching and single call dispatch. - Use the server's callback URL in the OAuth setup UI. - Omit untouched optional object defaults; keep supplied empty strings, false, zero, and object values intact in the action tester, with strict required-child validation. - Require an explicit retry after OAuth refresh instead of replaying a tools/call with missing session headers. - Record all eleven successful reads, write policies, auth-method limits, and qualification caveats. Omit credentials and private account data. - Find chat connector cards through catalog search in browser tests, so pagination does not hide Slack or other later entries. ## Verification - Real browser qualification: eleven bounded reads passed. Catalog refresh and reload proof are recorded in `doc/connections/verified-mcp-qualification-2026-10-08.md`. - Provider writes and actual agent-adapter sessions were not run. Alternative documented auth methods and expiry refresh remain unqualified. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Focused shared/catalog tests: 41 passed. Source override tests: 3 passed. Catalog regeneration fixture suite: 9 passed. Latest form regression suites: 31 passed. Gateway suite: 38 passed. Other focused UI suites passed before these last fixes. - UI token gates and both token sync checks: passed. - The catalog regeneration fixture includes the authoritative provider overrides; its full suite passes. The focused OAuth socket case passed on isolated rerun. - Complete local UI suite: 7,866 passed. CLI suite: 511 passed, 6 skipped. Complete shared suite: 892 passed. - The unchanged AgentMail/ClickUp discovery fallback suite passes all 45 cases. An OpenAI login test passed on isolated rerun. - Local broad database coverage is limited by macOS PostgreSQL shared-memory exhaustion (`shmget: No space left on device`, not disk space). The broad run was interrupted after diagnosis; no host settings or other running services were changed. - CI found that the Slack browser test assumed its catalog card was on the first page. The test now uses catalog search; the Slack case passes locally. Adjacent GitHub and iMessage cases could not start locally because embedded PostgreSQL initialization failed before browser assertions. - Final commit [`febe4520e`](https://github.com/paperclipai/paperclip/commit/febe4520e13d4b3a4121eafc3d22540c2b11379d): all 54 checks passed, with none pending or failed. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37847182065) passed typecheck, build, all general and serialized test suites, all eight browser shards, runner checks, canary dry run, and the aggregate verification gate. - Greptile reviewed that exact final commit at 2026-10-08 21:33 UTC and returned 5/5 with no outstanding findings or unresolved threads. - PR UI preview against the existing local test server: Calendly catalog/setup observed; Firecrawl read passed after expanding More options with nested optional inputs untouched. This retest proves frontend behavior, not the revised OAuth backend against a live provider. <details> <summary>Browser evidence</summary>    </details> ## Risks - Provider registration and consent behavior can change. Airtable's DCR option is explicit and keeps the existing issuer, resource, redirect, and PKCE checks. - Tavily uses initialized sessions without automatic call retries. After a successful OAuth refresh following 401, callers now receive a retry-required result. A later explicit retry can still require reauthorization if the provider rejects the refreshed token. Provider writes are not live-qualified. - Optional object cleanup affects the shared action tester. Regression tests cover absent objects, required children, defaults, and populated values. - GSC Wizard's underlying Google data scopes and Google flow completion were not independently verified. Its account reported paid/trial metadata of unknown origin. No purchase was performed. - No database migration, new AgentMail integration, or change to existing action grants is included. ## Model Used OpenAI GPT-6 through Codex assisted with implementation, research, tool use, and code execution. The root backend model ID and context window are not exposed in this session. The cheaper subagents used OpenAI `gpt-6-luna`. No model context size is inferred. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted checks and complete UI/CLI/shared suites passed; the broad local database run was blocked by the macOS PostgreSQL startup limitation documented above. The full database/workspace suite passed in CI. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f6d2a6b79e |
refactor(heartbeat): extract queue admission and wakeup dispatch (#15626)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service admits and dispatches agent runs. > - Earlier extractions separated workspaces, preparation, state, retries, and restart recovery. > - Queue admission and wakeup dispatch still occupy more than 4,200 lines in the main service. > - Their transaction, budget, ownership, and concurrency checks must stay together. > - This pull request moves those operations into one queue module and adds boundary tests. > - Reviewers can inspect queue policy separately from adapter execution. ## Linked Issues or Issue Description **What existing behavior does this improve?** The structure and testability of heartbeat queue admission and wakeup dispatch. **Subsystem affected** server/ — orchestration services. **Current behavior** Wake coalescing, queued claims, dispatch, daily limits, and native wake intents live inside `heartbeat.ts`. **Proposed behavior** Keep their behavior in `server/src/services/heartbeat/queue.ts`. Bind lifecycle callbacks and shared ownership with `createHeartbeatQueue`. **Reason and benefit** This removes 4,243 lines from `heartbeat.ts`, leaving 14,201 lines. The queue module has 4,576 lines and keeps related policy together in one file. **Breaking changes** None. The service methods, public exports, database transactions, locks, and execution gates keep their existing behavior. **Additional context** Continues #15622 and the earlier merged heartbeat extractions. A search of open and closed PRs found no duplicate queue extraction. Related #15600 adds a new experimental scheduler; this change preserves the current scheduler policy. This does not add a roadmap feature. ## What Changed - Move wake coalescing and batching, queue claims, daily caps, concurrency and priority checks, timer admission, queue resumption, and native status wake intents into `heartbeat/queue.ts`. - Pass lifecycle callbacks and the existing process-wide execution and wakeup sets into the factory. Preserve construction order with three forwarding callbacks. - Keep adapter execution, cancellation, and status writes in the service. - Add nine tests for inert construction, scheduling suppression, shared wake tracking, policy wiring, claim release, newer ownership, and budget gates. - Document the queue module boundary in `doc/DEVELOPING.md`. ## Verification - Before extraction: 133 tests passed across nine existing suites. - After extraction: **150 tests passed across 12 files**, including all nine new module tests and real PostgreSQL coverage. ```sh pnpm exec vitest run server/src/services/heartbeat/queue.test.ts server/src/__tests__/heartbeat-queued-run-claim-isolation.test.ts server/src/__tests__/heartbeat-start-lock.test.ts server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts server/src/__tests__/heartbeat-scheduling-suppression.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/heartbeat-running-followup.test.ts server/src/__tests__/heartbeat-dependency-scheduling.test.ts server/src/__tests__/heartbeat-auto-checkout.test.ts server/src/__tests__/heartbeat-archived-company-guard.test.ts server/src/__tests__/heartbeat-task-drain-admission-release.test.ts server/src/__tests__/heartbeat-worktree-suppression.test.ts ``` - Additional control and recovery coverage: **390 tests passed across five files**. Combined focused coverage is **540 passing tests across 17 files**, with no skips. ```sh pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-native-cleanup-admission.test.ts server/src/__tests__/heartbeat-native-status-context.test.ts server/src/__tests__/heartbeat-task-drain.test.ts server/src/services/chat-control-admission-retry.test.ts ``` - Full `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` was started and then stopped after the full CI suite passed. It did not finish and is not counted as a full local pass. - All **54 current-head checks passed**; two optional Storybook checks were skipped. This includes the full general and serialized test matrix, typechecks, build checks, Runner verification, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37847175452). - Greptile completed on head `cc0001961a052a716d81c136929f89312e0b5dee` with **5/5 and no actionable findings**. There are no review threads or merge conflicts. - Structural comparison confirms 31 moved function bodies and eight declarations match the originals. Remaining service function bodies and all 140 public exports are preserved. - An initial local run skipped database tests because the host had exhausted its PostgreSQL shared-memory slots. After verified stale segments were cleared, all 150 focused tests ran and passed with no skips. ## Risks Moving closures can break shared ownership or callback construction order. The factory receives the original sets and explicit callbacks. It starts no work during construction. Transaction boundaries, compare-and-set predicates, budget gates, cleanup holds, and lock ordering stay intact. This PR has no schema or API changes. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1009.0-canary.1 |
||
|
|
b9750b152f |
build(deps): bump rustls from 0.23.43 to 0.23.45 in /packages/paperclip-runner/runner (#15295)
Bumps [rustls](https://github.com/rustls/rustls) from 0.23.43 to 0.23.45. <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rustls/rustls/commit/2976d90fd1c2db6b518700dd101b714069cfcb17"><code>2976d90</code></a> Prepare 0.23.45</li> <li><a href="https://github.com/rustls/rustls/commit/f9d4ee92b20da535ef5982ffcaf6f62385f79641"><code>f9d4ee9</code></a> Handshake "alignment" check covers previously-received messages</li> <li><a href="https://github.com/rustls/rustls/commit/c4e92e8c9ee29ca3f552d4938efa73680c179179"><code>c4e92e8</code></a> Keep data needed for HRR processing together</li> <li><a href="https://github.com/rustls/rustls/commit/e553a7a70f99cad4722f06b32cc37857008a7d3f"><code>e553a7a</code></a> server: TLS1.2 is not available after a HRR</li> <li><a href="https://github.com/rustls/rustls/commit/5dff9da73798253d8d35909dd8548db522cd973c"><code>5dff9da</code></a> Test whether server negotiates TLS1.2 after HRR</li> <li><a href="https://github.com/rustls/rustls/commit/185a063b5fe1e9698c589d8be6a83d67224b95a1"><code>185a063</code></a> reject a second ClientHello that changes the cipher suite</li> <li><a href="https://github.com/rustls/rustls/commit/9fafe6dd0c15ea6e47399e94b2c6b647861f874b"><code>9fafe6d</code></a> reject a second ClientHello that drops pre_shared_key</li> <li><a href="https://github.com/rustls/rustls/commit/1b42c5f78e8e826d6afa72ce0ead3e3a4e624d79"><code>1b42c5f</code></a> providers: zeroize private key DER</li> <li><a href="https://github.com/rustls/rustls/commit/64ad386785c718c74262b79e6893813db381adfe"><code>64ad386</code></a> Bump version to 0.23.44</li> <li><a href="https://github.com/rustls/rustls/commit/1efbf662ed79d1392c542c0822805a818d232ac0"><code>1efbf66</code></a> bogo: remove PostQuantum setup</li> <li>Additional commits viewable in <a href="https://github.com/rustls/rustls/compare/v/0.23.43...v/0.23.45">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/paperclipai/paperclip/network/alerts). </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.1009.0-canary.0 nightly/v2026.1009.0-nightly.0 |
||
|
|
1c4ce44e13 |
fix: retain bounded workspace transfer failure evidence (#15637)
## Thinking Path > - Paperclip manages AI agents and the work they produce. > - Sandbox work must return to the host when an agent run ends. > - A failed transfer must keep its source available for recovery. > - Some transfer errors lose their useful fields before restore diagnostics reach Sentry. > - This change records fixed transfer stages, failure kinds, and numeric RPC codes. > - The next failure can identify the operation that failed without exposing workspace contents. ## Linked Issues or Issue Description **What happened?** A workspace transfer can fail after model execution succeeds. Several command, archive validation, and download errors become plain errors. The existing RPC envelope then reports `unknown`, with no transfer stage or command exit code. A host RPC timeout also loses its code from restore diagnostics. **Expected behavior** Keep bounded evidence through the provider, worker RPC, restore result, and Sentry context. Keep the same exception, error message, retry rules, and source retention policy. A transfer timeout must remain separate from the model execution timeout flag. **Steps to reproduce** Return a nonzero exit code from the sandbox archive command, return a per-file download error, exceed a tar listing limit, or time out the sync-out RPC. Observe the missing fields in the saved restore diagnostic. The added tests use local fixtures for these cases. Related public work: #15479 preserves the source after restore failure. #15481 carries bounded sync-out diagnostics through RPC. This change supplies missing producer evidence and extends that same envelope. I checked the roadmap and searched open PRs for duplicate transfer diagnostic work. ## What Changed - Add optional transfer stage and failure kind fields to the existing diagnostic envelope. Capture command exits and listing deadlines at their producers. - Record only known codes from typed RPC errors at the sync-out boundary. Revalidate every field before persistence and Sentry projection. - Keep annotations private to each outbound request and restore settlement. Preserve frozen error identity and prevent evidence from leaking across concurrent or later calls. - Document the fields. Cover producer failures, quota limits, old workers, concurrent error reuse, and redaction through the real Sentry SDK. ## Verification - Focused SDK, provider, host, persistence, and real Sentry tests: 449 passed, no skips. Independent review also ran the focused contracts and all provider tests. The optional Sentry SDK is required for this check, so the contract tests cannot skip. - A compiled Daytona transfer and compiled SDK worker pass a local RPC round trip. This uses the development TypeScript loader for workspace dependency exports. The transfer and capture modules are compiled JavaScript. - `pnpm -r typecheck` and `pnpm build` passed. The excluded Daytona package also passed `tsc --noEmit`. - The initial local `pnpm test:run` stopped in the general-server group: 60 suites could not initialize embedded PostgreSQL because the isolated install skipped its Darwin library-link setup. The existing package postinstall repairs this; `initdb --version` now passes. All 60 affected suites then passed: 1,342 tests, no skips. No local full aggregate pass is claimed. - Independent source review covers privacy, concurrency, packaging, and recovery behavior. ## Risks - These diagnostics help identify future failures. They do not establish or repair the cause of a past transfer failure. - New fields are optional. Older workers remain compatible. Unknown values are omitted. - The change adds a small request-local diagnostic scope. It keeps the original thrown errors, archive confinement, cleanup order, timeout values, retry count, and source retention policy. - No schema or deployment changes are required. No live provider actions are part of validation. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and independent agent review. The exact deployed model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b624a110e |
fix: checkpoint database backups before idle sleep (#15629)
## Thinking Path > - Paperclip coordinates agent work and preserves company state. > - Hosted instances can stop compute when admission and all work sources are quiet. > - The periodic database backup timer currently blocks every automatic sleep. > - Disabling backups would remove recovery points. > - This change offers an explicit final-backup checkpoint before sleep. > - A verified archive and restart marker preserve recovery while compute is stopped. ## Linked Issues or Issue Description **What happened?** An otherwise idle instance with automatic backups enabled always reports background work. Operators cannot reclaim idle compute without disabling those backups. The work inventory also queries the database before checking known local blockers. **Expected behavior** An operator can enable checkpoint mode. The server may authorize sleep only after all other work is quiet and a fresh backup is verified under the same owned hold. Backups resume on restart with the existing retention policy. **Steps to reproduce** Acquire a bounded idle drain on an instance with no application work and automatic backups enabled. Read its owned idle safety report. It reports present even when all other work is clear. With this change and `PAPERCLIP_DB_BACKUP_IDLE_CHECKPOINT_ENABLED=1`, a successful final backup can permit sleep; backup failures still refuse it. **Paperclip version or commit** Reproduced from master atcanary/v2026.1008.0-canary.22 |
||
|
|
4265cb3a2b |
fix(ci): cache native server integration builds in release verification (#15619)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Release verification must run the same server integration coverage as PR verification. > - Three server suites use real Rust Runner binaries. > - PR checks already run them with the shared Rust dependency cache, but release checks cold-build them inside ordinary server test shards. > - A cold Dot Runner build can consume its entire setup deadline before any test runs. > - This pull request gives release verification a required cached native integration lane and preserves every test. ## Linked Issues or Issue Description **What happened?** On master commit `1881894973a2b25838d8abed9bd8aeebc3af4441`, [Cloud readiness run 37837810707](https://github.com/paperclipai/paperclip/actions/runs/37837810707) failed in server shard 7. Dot Runner's `cargo build --release` exceeded its 300-second child-process deadline. The other 1,640 tests in that shard passed. The failed setup then tried to remove an undefined temporary path and emitted a second error. Cargo output was captured as an opaque buffer, which obscured build progress. **What did you expect to happen?** Build fixture binaries before tests in a lane with the existing Rust dependency cache. Run all native integration tests and make their result required for release readiness. A setup failure should keep its original diagnostic. **Steps to reproduce** Run release verification on a clean runner. The existing `general-server-without-chat` group retains the three Cargo-backed suites outside the PR workflow. The Dot suite can time out during a cold release build. Run the new workflow and partition tests against the previous source to reproduce five routing and coverage failures deterministically. **Version / commit** Observed at `1881894973a2b25838d8abed9bd8aeebc3af4441`; this change is based on `ed6abbf158b`. **Deployment mode** GitHub Actions Release and Cloud readiness verification. No application runtime or deployment action changes. Related work: #15581 also edits Dot integration tests for onboarding behavior. It does not repair release test routing. Searches found no open PR for this failure. ## What Changed - Add an explicit server test group that excludes the dedicated chat and native suites. Preserve the existing PR and local test groups. - Run all three native server suites in one required matrix lane with the existing trusted Rust dependency cache. - Build debug and release fixture binaries in a visible step with a 10-minute limit before tests start. Keep Cargo freshness checks and the existing test deadlines. - Stream Cargo diagnostics and clean up safely when Dot setup stops before creating its temporary directory. - Test complete, non-overlapping partitions for Release, Cloud readiness, other callers, and existing PR/local groups. Check cache restrictions, build order, and the source verification dependency. ## Verification - Passed 44 workflow and partition tests: `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - The same tests fail in five relevant cases against the previous workflow and selector. They pass after this correction. - An independent review repeated all 44 tests successfully. - Passed `actionlint .github/workflows/release-verify.yml`, `git diff --check`, and a secret scan. - Passed all 381 workflow policy tests: `node --test '.github/scripts/tests/*.test.mjs'`. - Passed full `pnpm build` in a clean worktree. - Built debug and release fixture binaries, then passed all 32 tests in the three real native integration suites: `pnpm test:run:general -- --group general-server-native-runner` (49.5 seconds, no skips). - Passed full `pnpm -r typecheck`. - Injected a synthetic Cargo setup failure. The suite reports that failure without the secondary undefined-path cleanup error. - Exact-head CI completed: 53 successful checks, including all server shards, both native Runner lanes, build, typecheck, browser tests, and the canary dry run. Two optional Storybook checks were skipped. - Started the duplicate local `pnpm test:run` aggregate and stopped it after exact-head CI passed. No completed local aggregate result is claimed. All test processes owned by that run exited. - Greptile scored the final head `ffe5ab5a28af2dfea9db1c1952c652f070bf52f8` at 5/5 with no review threads. - `Superagent Supply Chain Scan` is neutral, not passed: it only supports exact dependency pin replacements and cannot verify structural edits to `.github/workflows/release-verify.yml`. It reported no annotations. Independent source review, Greptile, actionlint, and the workflow policy tests cover this structural change. - GitHub still requires a code-owner review for the workflow files. There are no merge conflicts. ## Risks - Adds one CI matrix job, which increases concurrent runner demand. The existing cache writer and trust restrictions stay in place. - Missing dependencies still compile from the lockfile. A build that exceeds its step limit fails visibly; failed tests still block source verification. - Workflow execution uses the new lane after merge. PR tests validate the workflow contract and run the existing native test lane. - The supply chain scanner cannot analyze this workflow structure. Its neutral result is an explicit coverage limit; it is not counted as a successful scan. - No runtime, schema, deployment, publication, credential, or test-timeout changes. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and independent agent review. The exact deployed model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.21 |
||
|
|
2b4a1d7073 |
feat(dot): add guided invitations and persistent agent connections (#15581)
Add guided external-agent invitations with live Dot connection checks and ongoing access until revoked. Keep expired credentials invalid and enforce permissions for private avatar uploads. Preserve Runner recovery and remote launch boundaries. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
04de3132c2 |
refactor(heartbeat): extract restart recovery and lease cleanup (#15622)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service coordinates agent runs and their recovery. > - The earlier refactors separated workspaces, run preparation, run state, and retries. > - Restart recovery and lease cleanup still occupy more than 2,000 lines in the main service. > - These operations share process ownership and database claims that must stay intact. > - This pull request moves that group into one recovery module and adds boundary tests. > - Reviewers can now inspect recovery separately from queue admission and run execution. ## Linked Issues or Issue Description **What existing behavior does this improve?** The structure and testability of heartbeat restart recovery and lease cleanup. **Subsystem affected** server/ — orchestration services. **Current behavior** Hot-restart adoption, native restart recovery, shutdown draining, orphan reaping, and lease cleanup live inside `heartbeat.ts`. **Proposed behavior** Keep the same recovery behavior in `server/src/services/heartbeat/recovery.ts`. Bind the database and lifecycle callbacks with `createHeartbeatRecovery`. **Reason and benefit** This removes 2,181 lines from the main service. The extracted module is about 2,400 lines. It keeps related recovery operations together in one file. **Breaking changes** None. The public service methods, database claims, cleanup fences, and ownership checks stay the same. **Additional context** Continues the merged heartbeat extractions in #15568, #15573, #15578, #15591, #15601, and #15617. A search of open and closed PRs found no duplicate recovery extraction. This does not add a roadmap feature. ## What Changed - Move hot-restart snapshots and adoption, native restart recovery, shutdown draining, orphan reaping, and lease sweeps into `heartbeat/recovery.ts`. - Pass lifecycle callbacks and the existing shared execution sets into the factory. Preserve module-wide cleanup single-flight state. - Set the service's existing shutdown flag through a callback at the same point in shutdown preparation. - Add 14 focused tests for construction, shutdown ordering, empty selective drains, ownership races, shared state, retry timers, and background execution tracking. - Wait for the existing execution-drain barrier in the Stop recovery test instead of polling for one second. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - New module and additional cleanup coverage: **57 tests passed** across five files. ```sh pnpm exec vitest run server/src/services/heartbeat/recovery.test.ts server/src/__tests__/heartbeat-native-cleanup-admission.test.ts server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts server/src/services/hot-restart.test.ts ``` - Final recovery coverage on the rebased branch: **514 tests passed across 11 files**, including all 14 new module tests. ```sh pnpm exec vitest run server/src/services/heartbeat/recovery.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts server/src/__tests__/heartbeat-orphaned-active-lease-sweep.test.ts server/src/__tests__/native-session-resumption.test.ts server/src/__tests__/heartbeat-task-drain-admission-release.test.ts server/src/shutdown.test.ts server/src/__tests__/heartbeat-native-cleanup-admission.test.ts server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts server/src/services/hot-restart.test.ts ``` - Full `pnpm -r typecheck` and `pnpm build` passed, including after rebasing onto master. - Full local `pnpm test:run` was started again after the rebase. It was stopped after the full CI suite passed. It did not finish and is not counted as a full local pass. An earlier local AI-connection run hit an HTTP-test timeout; all corresponding CI coverage passed. - Structural comparison confirms 28 moved functions and 15 moved declarations. The only function-body adaptation is the shutdown flag callback. Remaining service functions and all 140 public exports match master. - All **54 current-head checks passed**; two unrelated Storybook checks were skipped. This includes the full general and serialized test matrix, typechecks, build checks, Runner verification, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37843023706). - Greptile completed on head `01fe39818d483e05f940b9fe842162a08e9b9e8e` with **5/5 and no actionable findings**. There are no review threads or merge conflicts. ## Risks The main risk is breaking shared execution ownership or changing recovery order while moving closures. The factory receives the original execution sets and lifecycle callbacks. It does no work during construction. Database transactions, compare-and-set predicates, cleanup deadlines, and native ownership gates stay intact. This PR has no schema or API changes. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
428b617fc9 |
test(evals): qualify native question resume paths (#15616)
## Thinking Path > - Paperclip lets agents pause tasks for human answers and continue the work. > - Native questions can yield a run or pause inside the provider's running turn. > - Those paths have different run identities and different forms. > - The documentation-placement test requires a semantic three-turn journey. > - A valid provider question therefore needs separate behavioral qualification. > - This pull request adds an opt-in two-answer journey with exact source and resume checks. > - Qualification exposed a Vite startup crash in idle-handler traversal, so the branch also distinguishes Connect mount paths from Express Route objects. ## Linked Issues or Issue Description **What existing behavior does this improve?** Product E2E evaluation of native task questions and answer delivery. **Current behavior** The semantic documentation case rejects the provider adapter's optional Other field before answering. Its three-run expectation also cannot prove a provider question that resumes within the same run. **Proposed behavior** Keep that semantic test intact. Add a separate opt-in case that classifies each recorded question by authoritative identities. Accept the optional Other field only for a verified provider question. Require both user answers, the correct continuation, and one saved final document. Related: #15564 added strict lifecycle checks. #15522 introduced idle-request tracking; this PR corrects its Connect/Vite compatibility after reproducing the startup crash. #14591 changes ACP question drafts and provider support; this PR has no overlapping files with that change. ## What Changed - Add the explicit-only `question-resume` suite for native Codex and ACPX Claude. Reuse the existing neutral choice-then-text task request. - Bind each question to its company, task, agent and run. Require an applied semantic tool receipt or the exact provider request identity. - Require a provider answer to continue the same run. Require a semantic answer to start one new run with the matching interaction and source-run wake fields. - Preserve question forms and both real board answers through the final checkpoint. Reject missing inputs, duplicate answers, extra runs and substituted identities. - Keep the existing lifecycle, output, task ownership and budget checks. Limit the new case to one attempt and one to three runs. - Add mutation calibrations and document the separate qualification claims. Keep production guidance, scheduling, provider code and the historical semantic case unchanged. - Fix startup with Vite middleware: treat Connect string mount paths separately from Express Route objects when traversing idle-request handlers. A real-Vite regression verifies startup and retained async work through an idle hold. ## Verification - 81 focused continuation, question and calibration tests passed before review. After the reference-preservation correction, all 50 question/resume tests pass, including two new negative cases that failed against the previous grader. - Final eval support suite passes: 1,921 Vitest tests and 128 Node tests, with one intentional skip. - Eval and workspace typechecks pass. Workspace build passes. - Catalog discovery finds only the two declared opt-in cells. Existing continuation remains 23 cells; default paid scope is unchanged. - Original [campaign 37831726391](https://github.com/paperclipai/paperclip/actions/runs/37831726391), source `1ac4b2873466896ff2892f49be82904927b6c57d`, failed during server startup before any test or provider run. Its original `transient_infrastructure` result and `cleanup: not_started` are preserved. Local reproduction identifies the Connect string route traversal crash; the question grader was never reached. - Setup-corrected [campaign 37833872898](https://github.com/paperclipai/paperclip/actions/runs/37833872898) measures source `9781a772389d41421df55759004676e1b905e512` with trusted workflow `8e59efc50b162126f33816841e06509634eba65c`. The workflow bytes are unchanged from the first dispatch. Same single selected native Claude cell, one attempt, at most three runs, ten-minute cell deadline, and 1,000-cent company/agent hard stops. Original result **PASS**, 24/24 checks and cleanup pass. Exactly three successful native Claude runs: two applied Paperclip `request_human_input` questions and one finishing run. Both real board answers persist, exact interaction/source-run response wakes match, one revision-1 task document contains Afternoon and the supplied reference, and final task state is done with no active lock, retry, recovery, monitor, pending interaction or child task. No model retry. - Current review head `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e` differs from the measured source only in the question/resume grader and its tests (11 insertions, 3 deletions). The saved document must contain the complete submitted reference, including its prefix and punctuation. Separate provider-free replay of the retained successful campaign passes all 24 continuation/lifecycle checks under this stricter grader. The original result and its measured source remain unchanged; no new provider run was made. - Observed paths were **semantic → semantic**. The provider built-in question/optional Other path was not exercised by this new live attempt. Its form acceptance has retained-original calibration; same-run answer/resume has deterministic positive/negative calibration. This neutral-task pass alone does not qualify the provider path; the subsequent explicit bridge trial below is separate evidence. Neither trial regrades the old Claude failure or establishes a causal behavior/performance comparison. - Subsequent explicit provider-path [campaign 37841107684](https://github.com/paperclipai/paperclip/actions/runs/37841107684) measures the final PR source `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`, using trusted workflow `2a5f65c9ae501b49b2a38ce7edd04209b806c9d0`. Exact cell: `continuation.runner-acpx-claude.local.provider-question-bridge`, Claude `claude-sonnet-5`, suite fingerprint `fc77203bf85fd99b06ebc53cd52982186c0ae0dc369fcc38a9ec3c737e9a31e5`. **Original PASS, 15/15 checks, cleanup passed, one actual provider run and zero eval rerolls.** The built-in question produced one choice plus its optional Other companion. The browser selected the requested reference, left Other blank, and submitted. Exact runtime request/response IDs and the selected option match the saved interaction; the same paused native run created one revision-1 document and finished Done with no pending interaction, lock, retry, recovery, monitor or child task. Independent evidence and runtime-receipt audits pass; marked initial/final screenshots were visually checked. [Original public report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37841107684-1/index.html). - This separate provider-directed trial qualifies one choice-answer bridge and traversal of an empty optional Other field. It does not establish natural tool selection, a typed Other answer, two consecutive built-in questions, free-text-only input, restart recovery or general reliability. No production, prompt or grader change was made for this dispatch. The old failure and the separate neutral semantic-path result retain their original grades. - Preserve within-run friction: one `write_task_document` input-schema rejection and two HTTP 409 “Document does not exist yet” responses preceded a successful `write_document` call. These are three unsuccessful tool calls inside the same run, not three extra provider runs; no flawless document-tool behavior is claimed. Billing reports token usage but remains unpriced/incomplete, so actual charges are unknown and local/hosted runtime is unmetered. Both company and agent hard stops were 1,000 cents, with one attempt, one selected cell, concurrency one and a ten-minute cell deadline. - Billing records all three model runs with token usage, but no priced dollar receipt (`unpriced`, `complete: false`). Actual charges are unknown; local/hosted runtime is unmetered. Numeric zero reported cost does not mean free. - Startup repair: real-Vite reproduction passes; 35 targeted idle-tracking tests pass, 34 database-dependent checks skip because embedded PostgreSQL is unavailable locally. Workspace typecheck and build pass again after the repair. - Full local repository database tests were not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Full Linux CI now passes on the final review head. - Final-head [CI run 37837804284](https://github.com/paperclipai/paperclip/actions/runs/37837804284) passes. Complete current-head audit: 53 successful check-runs, two intentional Storybook skips, and separate Snyk success. The ready transition’s contributor and security checks also pass (security scan: no findings). [Fresh review](https://github.com/paperclipai/paperclip/pull/15616#issuecomment-6068092601) is 5/5 on `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`; zero unresolved threads and no merge conflicts. ## Risks This is a new behavioral definition. The neutral two-question trial records semantic-tool use; the separate provider-directed trial records one built-in question. They do not establish consistent tool selection, documentation placement, default hiring, remote environments, crash recovery or general reliability. The separate explicit bridge trial covers one provider choice, an empty optional Other companion and same-run completion. Typed Other, two consecutive built-in questions and free-text-only provider input remain unqualified. The observed document-tool schema rejection and two 409 responses are retained as a separate follow-up, without attributing their cause. The historical case and its original grades remain intact. The reference question must still be text-only; the oracle does not turn an arbitrary extra question into an optional companion. Missing or unfamiliar identity evidence fails closed. The only production change distinguishes Connect string mount paths from Express route objects during idle-handler traversal. Unknown route objects still fail closed. No API, schema, migration or workflow change is included. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.20 |
||
|
|
65b5581805 |
Keep proven AI account selection blockers out of Sentry (#15618)
Preserve trusted evidence from explicit AI account selection rejections so known setup blockers stay out of Sentry. Keep task failure, blocked state, recovery actions, HTTP behavior, and unexpected error reporting unchanged. Validated with full typecheck/build, focused PostgreSQL lifecycle and reporting tests, independent reporting tests, and all PR CI checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2a5f65c9ae |
refactor(server): extract heartbeat retry scheduling (#15617)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat orchestration starts runs and recovers from supported failures. > - The heartbeat service still contains more than 22,000 lines after earlier extractions. > - Retry scheduling has a clear boundary, but it shares lifecycle callbacks with the service. > - This pull request moves retry scheduling into one module with explicit dependencies. > - It preserves the scheduling rules, database writes, and lock order. > - The benefit is a smaller service and focused tests for the retry boundary. ## Linked Issues or Issue Description **Current behavior** `server/src/services/heartbeat.ts` owns retry schedules, contention deferrals, due promotion, and retry-now requests. These operations sit among execution and recovery code. **Proposed behavior** Move these operations into `server/src/services/heartbeat/retries.ts`. Keep the public heartbeat API and existing helper exports. The service supplies its lifecycle callbacks and worktree cutoff. **Reason and benefit** Extract one complete area in a separate PR. This removes 1,641 lines from the main service without changing retry policy. Future retry changes can use a smaller module and direct tests. **Breaking changes** None. Existing imports, error-class identity, retry metadata, and API outcomes remain available. **Additional context** Refs #15601, the preceding run-state extraction. Related open retry changes: #14930, #13975, #15157, #12587, and #13276. This PR moves the existing implementation. Those policy changes remain separate. ## What Changed - Add `createHeartbeatRetries(db, dependencies)` for bounded retries, workspace and connection deferrals, shared-workspace holder checks, due promotion, and retry-now requests. - Keep status transitions, event publication, issue-lock release, plan-resume reporting, and worktree settings in the heartbeat service. - Preserve all 140 existing public exports. Re-export the same helpers, constants, and workspace-busy error class. - Add 14 tests for construction, callback order, cancellation races, concurrent scheduling, committed promotion, and retry-now scope. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Existing focused suite before the move: 84 tests passed. - Expanded focused suite after the move: 243 tests passed across 11 files. Command: `pnpm exec vitest run server/src/services/heartbeat/retries.test.ts server/src/__tests__/heartbeat-retry-scheduling.test.ts server/src/__tests__/heartbeat-workspace-busy.test.ts server/src/__tests__/heartbeat-model-fallback-warning.test.ts server/src/__tests__/heartbeat-running-followup.test.ts server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts server/src/__tests__/issue-scheduled-retry-routes.test.ts server/src/modules/run-dispatch` - After rebasing onto the latest master, 61 retry and plugin-idle tests passed across four files. - `pnpm -r typecheck` passed again on the rebased branch. - `pnpm build` passed again on the rebased branch. - `pnpm test:run` was started locally. The long serial run was stopped after full CI passed; it is not claimed as a complete local pass. - Full CI passed on commit `69e76f3eec75b0703570f1fce37ebfe9c63fa3ac`: 54 successful checks, two intentional Storybook skips, zero failures. This includes general tests, serialized server tests, browser tests, runner checks, typecheck, build, and canary verification. - The setup-timeout review fix passed all 14 module tests and the server typecheck. - Greptile completed a fresh review of commit `69e76f3eec75b0703570f1fce37ebfe9c63fa3ac`: 5/5, no open findings. - A structural comparison confirmed 22 unchanged moved functions and one unchanged three-line string reader, 16 unchanged declarations, one unchanged API body, and all 140 preserved exports. The remaining service body is unchanged after accounting for the new binding and removed declarations. ## Risks - A missing lifecycle callback or an eager call during construction could change retry behavior. The dependency contract and direct tests cover these boundaries. - Concurrent scheduling must keep the issue-first lock order and reuse one successor. The original transaction bodies are unchanged. PostgreSQL tests cover concurrent calls from separate factory instances. - Promotion must publish after commit and preserve worktree cutoffs. Direct tests cover both paths. - No schema, route, adapter, or retry-policy changes are included. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed6abbf158 |
build(deps): bump actions/cache/restore from 5.1.0 to 6.1.0 (#15163)
Bumps [actions/cache/restore](https://github.com/actions/cache) from 5.1.0 to 6.1.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/actions/cache/releases">actions/cache/restore's releases</a>.</em></p> <blockquote> <h2>v6.1.0</h2> <h2>What's Changed</h2> <ul> <li>Bump <code>@actions/cache</code> to v6.1.0 - handle read-only cache access by <a href="https://github.com/jasongin"><code>@jasongin</code></a> in <a href="https://redirect.github.com/actions/cache/pull/1768">actions/cache#1768</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/cache/compare/v6...v6.1.0">https://github.com/actions/cache/compare/v6...v6.1.0</a></p> <h2>v6.0.0</h2> <h2>What's Changed</h2> <ul> <li>Update packages, migrate to ESM by <a href="https://github.com/Samirat"><code>@Samirat</code></a> in <a href="https://redirect.github.com/actions/cache/pull/1760">actions/cache#1760</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/cache/compare/v5...v6.0.0">https://github.com/actions/cache/compare/v5...v6.0.0</a></p> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/actions/cache/blob/main/RELEASES.md">actions/cache/restore's changelog</a>.</em></p> <blockquote> <h1>Releases</h1> <h2>How to prepare a release</h2> <blockquote> <p>[!NOTE] Relevant for maintainers with write access only.</p> </blockquote> <ol> <li>Switch to a new branch from <code>main</code>.</li> <li>Run <code>npm test</code> to ensure all tests are passing.</li> <li>Update the version in <a href="https://github.com/actions/cache/blob/main/package.json"><code>https://github.com/actions/cache/blob/main/package.json</code></a>.</li> <li>Run <code>npm run build</code> to update the compiled files.</li> <li>Update this <a href="https://github.com/actions/cache/blob/main/RELEASES.md"><code>https://github.com/actions/cache/blob/main/RELEASES.md</code></a> with the new version and changes in the <code>## Changelog</code> section.</li> <li>Run <code>licensed cache</code> to update the license report.</li> <li>Run <code>licensed status</code> and resolve any warnings by updating the <a href="https://github.com/actions/cache/blob/main/.licensed.yml"><code>https://github.com/actions/cache/blob/main/.licensed.yml</code></a> file with the exceptions.</li> <li>Commit your changes and push your branch upstream.</li> <li>Open a pull request against <code>main</code> and get it reviewed and merged.</li> <li>Draft a new release <a href="https://github.com/actions/cache/releases">https://github.com/actions/cache/releases</a> use the same version number used in <code>package.json</code> <ol> <li>Create a new tag with the version number.</li> <li>Auto generate release notes and update them to match the changes you made in <code>RELEASES.md</code>.</li> <li>Toggle the set as the latest release option.</li> <li>Publish the release.</li> </ol> </li> <li>Navigate to <a href="https://github.com/actions/cache/actions/workflows/release-new-action-version.yml">https://github.com/actions/cache/actions/workflows/release-new-action-version.yml</a> <ol> <li>There should be a workflow run queued with the same version number.</li> <li>Approve the run to publish the new version and update the major tags for this action.</li> </ol> </li> </ol> <h2>Changelog</h2> <h3>6.1.0</h3> <ul> <li>Bump <code>@actions/cache</code> to v6.1.0 to pick up <a href="https://redirect.github.com/actions/toolkit/pull/2435">actions/toolkit#2435 Handle cache write error due to read-only token</a></li> <li>Switch redundant "Cache save failed" warning to debug log in save-only</li> </ul> <h3>6.0.0</h3> <ul> <li>Updated <code>@actions/cache</code> to ^6.0.1, <code>@actions/core</code> to ^3.0.1, <code>@actions/exec</code> to ^3.0.0, <code>@actions/io</code> to ^3.0.2</li> <li>Migrated to ESM module system</li> <li>Upgraded Jest to v30 and test infrastructure to be ESM compatible</li> </ul> <h3>5.0.4</h3> <ul> <li>Bump <code>minimatch</code> to v3.1.5 (fixes ReDoS via globstar patterns)</li> <li>Bump <code>undici</code> to v6.24.1 (WebSocket decompression bomb protection, header validation fixes)</li> <li>Bump <code>fast-xml-parser</code> to v5.5.6</li> </ul> <h3>5.0.3</h3> <ul> <li>Bump <code>@actions/cache</code> to v5.0.5 (Resolves: <a href="https://github.com/actions/cache/security/dependabot/33">https://github.com/actions/cache/security/dependabot/33</a>)</li> <li>Bump <code>@actions/core</code> to v2.0.3</li> </ul> <h3>5.0.2</h3> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/actions/cache/commit/55cc8345863c7cc4c66a329aec7e433d2d1c52a9"><code>55cc834</code></a> Merge pull request <a href="https://redirect.github.com/actions/cache/issues/1768">#1768</a> from jasongin/readonly-cache</li> <li><a href="https://github.com/actions/cache/commit/d8cd72f230726cdf4457ebb61ec1b593a8d12337"><code>d8cd72f</code></a> Bump <code>@actions/cache</code> to v6.1.0 - handle cache write error due to RO token</li> <li><a href="https://github.com/actions/cache/commit/2c8a9bd7457de244a408f35966fab2fb45fda9c8"><code>2c8a9bd</code></a> Merge pull request <a href="https://redirect.github.com/actions/cache/issues/1760">#1760</a> from actions/samirat/esm_migration_and_package_update</li> <li><a href="https://github.com/actions/cache/commit/e9b91fdc3fea7d79165fceb79042ef45c2d51023"><code>e9b91fd</code></a> Prettier fixes</li> <li><a href="https://github.com/actions/cache/commit/e4884b8ff7f92ef6b52c79eda480bbc86e685adb"><code>e4884b8</code></a> Rebuild dist</li> <li><a href="https://github.com/actions/cache/commit/10baf0191a3c426ea0fa4a3253a5c04233b6e18f"><code>10baf01</code></a> Fixed licenses</li> <li><a href="https://github.com/actions/cache/commit/e39b386c9004d72a15d864ade8c0b3a702d47a37"><code>e39b386</code></a> Fix test mock return order</li> <li><a href="https://github.com/actions/cache/commit/b6928203372a8571ff984c0c883ef3a1adfb0c06"><code>b692820</code></a> PR feedback</li> <li><a href="https://github.com/actions/cache/commit/60749128a44d25d3c520a489e576380cf00ff3f1"><code>6074912</code></a> Rebuild dist bundles as ESM to match type:module</li> <li><a href="https://github.com/actions/cache/commit/5a912e8b4af820fa082a0e75cfd2c782f8fbfe0e"><code>5a912e8</code></a> Fix lint and jest issues</li> <li>Additional commits viewable in <a href="https://github.com/actions/cache/compare/caa296126883cff596d87d8935842f9db880ef25...55cc8345863c7cc4c66a329aec7e433d2d1c52a9">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.1008.0-canary.19 |
||
|
|
1881894973 |
fix: prevent inbox archive deadlocks during completion (#15615)
Take a company-scoped parent lock before writing archive state and reuse the caller transaction. Preserve newer archive timestamps and attribution when a request waits, so completed tasks remain archived. Verified with real PostgreSQL race, rollback, company isolation, concurrency, timestamp and visibility tests, independent review, full CI and Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3d1d5294d9 |
Drain idle plugin workers before automatic sleep (#15599)
## Thinking Path > - Paperclip runs autonomous work through agents, schedules, and plugins. > - Automatic idle sleep must preserve accepted work and unfinished cleanup. > - Every enabled plugin currently blocks sleep, including unused providers. > - An enabled flag or package version cannot prove that a worker is idle. > - This change adds a live worker drain under the existing owned hold. > - Idle workers can allow sleep while unknown work stays protected. ## Linked Issues or Issue Description Refs #15522. Related: #15391 adds plugin readiness for agent admission; this change concerns instance sleep and does not replace that contract. **Subsystem affected** Plugin worker lifecycle and automatic idle sleep. **Problem or motivation** A workspace with no pending work cannot sleep when any plugin is enabled. Removing that check alone would lose accepted RPCs, background tasks, or cleanup after a caller timeout. **Proposed solution** Require a live `onIdleDrain` handshake from each worker. Close admission in both processes for the exact owner and expiry. Count accepted work until completion. Continue checking durable work separately. ## What Changed - Add bounded worker holds, exact-owner release, automatic expiry, and an abort signal for plugin-owned background work. - Count host and worker requests through their real completion receipts. Keep timed-out work counted. Check active notifications and terminal routes. - Accept enabled plugins only when their current workers provide matching runtime receipts. Missing workers, old SDKs, crashes, invalid replies, and unknown cleanup still prevent sleep. - Let unused Daytona workers opt in. Once a worker contacts the provider, it remains a blocker for that process lifetime. This restriction avoids treating its existing timeout and terminal-close behavior as a cleanup receipt. - Document the plugin author contract. No schema or user-facing API is added. ## Verification - All hosted CI checks passed on |
||
|
|
4fe45bf8fe |
refactor(heartbeat): extract run retrieval and session state (#15601)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat runs retain task sessions, usage totals, and bounded run data. > - After the preparation extraction, the main service still has over 24,000 lines. > - Run reads and session operations share one database and have no dispatch dependencies. > - This pull request moves that group into one module and preserves its callers. > - The benefit is a smaller run engine and one place to maintain saved run and session state. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat run retrieval and session state. **Current behavior** `server/src/services/heartbeat.ts` has 24,033 lines on the merged base. It mixes run projections, session queries, resume rules, compaction, and usage helpers with execution and recovery. **Proposed behavior** Move that group into `server/src/services/heartbeat/run-state.ts`. The main file loses 1,770 lines and ends at 22,263 lines. Keep current behavior, public exports, and service methods. This continues #15568, #15573, #15578, and #15591. **Reason and benefit** One domain extraction makes useful progress toward a few manageable files. Run reads and saved session state can be reviewed together without searching the run engine. A search found no duplicate open extraction. Related session changes include #15142, #14661, and #14670; this PR does not apply their proposed behavior changes. **Breaking changes** None. Existing public helpers retain their identity. Service method names and results stay the same. ## What Changed - Move bounded run projections, encoding checks, and run/event reads into `heartbeat/run-state.ts`. - Move task session persistence, explicit resumes, reset rules, compaction, and usage/billing helpers into the same module. - Bind database operations through `createHeartbeatRunState(db)`. Construction does no database work. The encoding cache stays local to each service instance. - Keep execution order, status transitions, session-goal recovery, cost accounting writes, dispatch, and cancellation in `heartbeat.ts`. - Preserve all 140 exports. A syntax-tree comparison confirms that 43 moved function bodies, 22 declarations, and nine service method bodies are unchanged. The copied private string helper is unchanged. Surrounding orchestration is unchanged apart from the factory binding. - Add eight tests for legacy export identity, database binding, encoding cache isolation, scoped resumes, cumulative usage, compaction, and stale or cancelled conversation session writes. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Before and after extraction, eight existing suites pass: 235 tests. They cover workspace/session helpers, runtime state, cost accounting, ledger attribution, task reset, timer reset, run lists, and run privacy. - The new module suite passes: eight tests, including five real PostgreSQL cases. Total focused coverage: 243 tests across nine suites. - Commands: `pnpm exec vitest run server/src/services/heartbeat/run-state.test.ts` and `pnpm exec vitest run server/src/__tests__/heartbeat-runtime-state.test.ts server/src/__tests__/heartbeat-cost-accounting.test.ts server/src/__tests__/heartbeat-ledger-billing-code.test.ts server/src/__tests__/heartbeat-task-session-reset.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts server/src/__tests__/heartbeat-list.test.ts server/src/__tests__/heartbeat-run-privacy-routes.test.ts`. - `pnpm --filter @paperclipai/server typecheck` passes. - `pnpm -r typecheck` and `pnpm build` pass. - The full local `pnpm test:run` did not complete and was stopped with SIGINT after it reported an OAuth scope health-check failure and a workspace-runtime timeout. All six OAuth variants and the workspace case pass in isolation. A single rerun of the 413-case tool-access suite passed the OAuth case but hit a different Notion callback timeout; that Notion case also passes in isolation. The full local suite is not claimed as passing. - Greptile is 5/5 with no findings or inline comments on head `2ca0908cfc112dd6ae26e2565cb3611b5bec52ba`. Its exact-head check completed successfully. - CI remains blocked on hosted-runner assignment after more than ten minutes. [`ci / Select trusted runner`](https://github.com/paperclipai/paperclip/actions/runs/37828454507/job/113487287178) is queued for `ubuntu-latest` with no runner assigned. All seven completed review/security checks pass; two optional Storybook checks are skipped. The full CI matrix has not started. It must complete before merge. ## Risks - Wiring errors could bind reads to the wrong database or share the encoding cache. Direct factory tests cover database and cache isolation. Real PostgreSQL suites cover query and transaction behavior. - Session writes must retain their issue row locks and generation fences. The original function bodies are unchanged. New database tests reject stale and cancelled conversation writes and clears. - Existing SQL_ASCII safeguards and bounded projections remain in place. No schema, API contract, migration, lockfile, or workflow changes are included. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8e59efc50b |
fix(ui): keep sidebar plugin outlets quiet when the server is unreachable (#15588)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The board UI sidebar shows plugin contributions: plugin nav slots,
plugin launchers, and the plugin panel at the bottom
> - These outlets read `GET /plugins/ui-contributions` through React
Query
> - When the server is unavailable, the refetch fails. Each outlet then
replaces its content with a red "Plugin extensions unavailable: Failed
to fetch" or "Plugin launchers unavailable" box
> - React Query still has the last good data, so the error box removes
useful content and adds noise in a persistent part of the UI
> - This pull request adds an option to hide the outlet error, and the
sidebar uses it
> - The benefit is a quiet sidebar during a server outage. The sidebar
keeps the last loaded plugin items. Other pages still show the error
inline
## Linked Issues or Issue Description
No public issue exists. I searched open and closed issues and PRs for
related work and found none.
**What happened?**
When the server is unavailable, the sidebar shows red error boxes in
place of plugin items. The text is "Plugin extensions unavailable:
Failed to fetch" or "Plugin launchers unavailable: Failed to fetch".
**Expected behavior**
The sidebar does not show errors while the server is unavailable. It
keeps the last loaded content.
**Steps to reproduce**
1. Start Paperclip with at least one plugin that contributes a `sidebar`
slot, a `sidebar` launcher, or a `sidebarPanel` slot.
2. Open the board UI and let the sidebar load.
3. Stop the server.
4. Wait for the next refetch of `/plugins/ui-contributions`, or focus
the window.
5. See the red error boxes in the sidebar.
**Paperclip version or commit**
`master` at
|
||
|
|
53105d5830 |
feat(apps): add Gauge connection (#15596)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Gauge (`withgauge.com`) measures how AI answer engines mention and cite a brand. Its hosted MCP server also gives SEO, analytics and ads reports, and runs a content pipeline that can publish to a connected CMS. > - Gauge is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no organization guidance, and no warning about live CMS publishing. > - The connector playbook supports this provider with existing definition fields. The server supports dynamic client registration with a reviewed `mcp` scope, and it accepts an organization API key as a bearer header. > - This pull request adds the Gauge definition, its artwork, the research and permission-review ledger rows, documentation, and deterministic tests. > - The benefit is a branded, governed Gauge connection with browser sign-in, an API-key option, and a clear publish warning. ## Linked Issues or Issue Description **Problem or motivation** Marketing and growth teams use Gauge to track their brand in AI answers and to run content workflows. They want their Paperclip agents to read visibility, keyword and traffic data, and to prepare content. Gauge is not in the Apps catalog. Operators must paste the MCP URL into the generic remote-MCP flow, which gives no branding, no method guidance, and no warning that content tools can publish live. **Proposed solution** Add a catalog-only Gauge connection built from the connector playbook. Browser sign-in uses Gauge's dynamic client registration and requests only the reviewed `mcp` scope. The user selects one Gauge organization on the consent screen. An optional method sends a customer-created organization API key as an `Authorization: Bearer` header. Both methods warn the operator to set publish actions to Ask first. Every discovered tool stays governed by the normal per-action policies. **Alternatives considered** A plugin was not needed because no custom UI, tables, workers, or webhooks are involved. The identity scopes that Gauge also advertises (`openid`, `profile`, `email`, `organizations`) are not requested, because Gauge selects the organization on its consent screen. A Gauge-specific `classifyRisk` rule was not added, because the tool names cannot be seen without an account. The generic rule already classifies `publish`, `create` and `update` tools as writes. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog and does not overlap planned core work. ## What Changed - Added the `gauge` provider to `scripts/ingest-app-definitions.mjs` (category `analytics`, API-key placement, guidance, warnings, description) and regenerated `packages/shared/src/app-definitions/gauge.json` and the generated registry. - Added the Gauge row to the self-serve MCP research ledger with `dcr_or_api_key` auth and risk tier S3. - Added permission reviews for `gauge/mcp-oauth` (explicit scope `mcp`, from Gauge's live authorization-server metadata) and `gauge/mcp-api-key` (provider key), with evidence links. - Added Gauge's mark (`ui/public/brands/apps/gauge.png`, the avatar of Gauge's official GitHub organization) and the brand manifest entry. - Added gallery copy for the Gauge card. - Added `doc/connections/GAUGE.md` (endpoints, scopes, administrator setup, capabilities and policy, manifest, brand provenance, validation hook) and linked it from the connections README and the permission audit. - Tests: Gauge definition shape, store visibility and artwork, URL recognition, reviewed scope, bearer-header placement, the connect form's sign-in default and API-key gating, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx server/src/__tests__/tool-access-service.test.ts` — 738 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui typecheck` — clean. - `node scripts/ingest-app-definitions.mjs --definitions-only` produces no Gauge drift. - Manual: on a local instance, open Apps → Browse and confirm the Gauge card and icon. Open `/apps/connect?source=gauge` and confirm that sign-in is the default and that the API-key method is under Advanced. - Live metadata probe on 2026-10-08: an unauthenticated `initialize` on `https://app.withgauge.com/mcp` returns 401 with `resource_metadata`. Both `.well-known` documents return the recorded endpoints and scopes. Dynamic client registration succeeds. The authorize endpoint accepts `scope=mcp` and rejects an unknown scope with HTTP 400. ## Risks - Low risk to existing providers: the change is additive catalog data plus tests. The generated registry only gains one import. - Gauge content tools can publish to a connected CMS, including live, and a bulk keyword update replaces each prompt's keyword list. Both methods warn the operator, and the API-key helper text tells operators to set publish actions to Ask first. - Gauge API keys are not scoped and reach the whole organization. Paperclip cannot narrow an issued key. - The permission-review ledger records live proof for both methods as not run. The lifecycle checklist in `doc/connections/GAUGE.md` still needs a documented pass with a Gauge organization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with tool use (shell, file editing, web research). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0111689d0a |
test(ui): cover monitor-only task policies (#15587)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators open tasks to read progress and edit reviewers and approvers. > - Stored execution policies can contain a monitor and omit the stages array. > - These policies previously crashed the task properties panel. > - The policy reader now applies schema defaults through the fix in #15593. > - This pull request adds regression coverage for monitor-only policies with external service metadata. > - The tests check that the panel renders and participant edits keep the existing monitor. ## Linked Issues or Issue Description **What happened?** A task with a monitor-only execution policy could crash with `Cannot read properties of undefined (reading 'find')`. The stored policy omitted `stages`. **Expected behavior** The task shows no reviewers or approvers. Adding both kinds of participant keeps the monitor and its external service metadata. **Steps to reproduce** 1. Load a task whose execution policy contains an external service monitor and omits `stages`. 2. Render its properties panel and inspect both participant controls. 3. Add reviewers and approvers. Check that the resulting policy keeps the monitor. Refs #15593, which provides the policy normalization used by these tests. ## What Changed - Test that both stage types return empty participant lists for a monitor-only policy. - Test adding reviewers and approvers while keeping the external service monitor. - Test retaining a monitor when no review or approval stages exist. - Render the properties panel with omitted stages and check its participant and monitor controls. ## Verification - The helper and properties panel suites pass on current master with these tests: 108 tests. - `pnpm check:token-gates` passes. - All CI gates pass on the updated PR commit, including build, types, Runner verification, full test suites, browser suites, security scans, and canary checks. - Greptile gives commit `2634bd1d` 5/5 with no actionable findings or unresolved review threads. - The full local workspace suite was not repeated for this test-only rebase. CI verifies the full suite, build, types, Runner checks, browser suites, security scans, and canary run. ## Risks - Low risk. This change adds tests against the existing policy behavior. - The fixture uses example external references and test participant IDs. - Existing documentation remains accurate for this regression coverage. ## Model Used - OpenAI Codex, based on GPT-6, with tool use and code execution. The exact serving model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0d57b9b98e |
refactor(heartbeat): extract run preparation (#15591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service prepares context and configuration before agent runs. > - The workspace extraction left the main file at about 27,000 lines. > - Preparation helpers form another large boundary with explicit inputs and one database dependency. > - This pull request moves those helpers into one module and preserves their callers. > - The benefit is a 24,000-line orchestration file and one place to maintain run preparation. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat run preparation. **Current behavior** `server/src/services/heartbeat.ts` mixes context and configuration preparation with execution, queueing, and recovery in 27,048 lines. **Proposed behavior** Move preparation into `server/src/services/heartbeat/run-preparation.ts`. The main file loses 3,048 lines and ends at 24,000 lines. Keep the existing imports and behavior. This continues #15568, #15573, and #15578. **Reason and benefit** A larger domain extraction makes useful progress toward a few manageable files. Context, identity, environment, and tool preparation can be reviewed together without searching the run engine. **Breaking changes** None. Existing public helpers and the configuration-incomplete error class retain their identity. ## What Changed - Move wake payloads, comments, attachments, skill mentions, adapter environment configuration, and MCP/tool setup into `heartbeat/run-preparation.ts`. - Bind issue context, pinned routine snapshots, organization rows, and responsible-user resolution through `createHeartbeatRunPreparation(db)`. - Keep factory construction free of database work. Keep queueing, dispatch, retries, cancellation, and execution order in `heartbeat.ts`. - Preserve all 140 existing exports. A syntax-tree comparison confirms that the 47 extracted function bodies and 25 moved declarations are unchanged. The private string normalizer is also unchanged. - Add six tests for legacy export identity, independent database binding, company scope, pinned routine configuration, wake author selection, company-owner fallback, and credential preflight. - Document the preparation boundary in `doc/DEVELOPING.md`. ## Verification - Passed: 183 focused tests across ten suites, including the six new tests. Command: `pnpm exec vitest run server/src/services/heartbeat/run-preparation.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/heartbeat-agent-session-message.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts server/src/__tests__/heartbeat-runtime-skills.test.ts server/src/__tests__/heartbeat-project-env.test.ts server/src/__tests__/heartbeat-responsible-user-invariant.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/run-secret-redaction.test.ts server/src/__tests__/low-trust-red-team-routes.test.ts`. - One runtime-skills case reported a Postgres deadlock in its after-test `TRUNCATE` cleanup. A focused rerun passed both runtime-skills cases. The other nine suites passed in the combined run. - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: all 54 GitHub checks on `8d0cb9a251b10be5735fa13035e8a31e50b132e7`. The two optional Storybook jobs were skipped. The complete CI test matrix is green. - The full local `pnpm test:run` was started and then stopped after the complete CI test matrix passed. It had no failures reported before stopping and did not complete locally. - Greptile: 5/5 on the same commit, with no inline comments or actionable findings. ## Risks The extraction crosses context, identity, credential, and tool-access module boundaries. Existing company filters, secret redaction, low-trust rules, and error classes stay intact. The new factory captures only the service database and does no database work during construction. Existing integration tests cover responsible-user authority, credential boundaries, MCP tokens, and quarantine. New tests cover the factory wiring and legacy identity. A syntax-tree comparison confirms that orchestration changed only to bind the extracted loaders. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f4acc8f14 |
build(deps): bump @modelcontextprotocol/sdk from 1.30.0 to 1.31.0 (#15361)
Bumps [@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk) from 1.30.0 to 1.31.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@modelcontextprotocol/sdk's releases</a>.</em></p> <blockquote> <h2>1.31.0</h2> <h2>Upgrade notes</h2> <ul> <li>Stored OAuth tokens and client information now include an <code>issuer</code> field. Storage that rejects unknown fields needs to allow it.</li> <li>Pass <code>expectedIssuer</code> when constructing <code>ClientCredentialsProvider</code>, <code>PrivateKeyJwtProvider</code> or <code>StaticPrivateKeyJwtProvider</code>. Constructing them without it is deprecated.</li> </ul> <h2>What's Changed</h2> <ul> <li>[v1.x] Bind stored OAuth credentials to the authorization server that issued them by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2888">modelcontextprotocol/typescript-sdk#2888</a></li> <li>chore: bump version to 1.31.0 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2890">modelcontextprotocol/typescript-sdk#2890</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0</a></p> <h2>1.30.1</h2> <h2>What's Changed</h2> <ul> <li>[v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON-RPC batch length by <a href="https://github.com/maxisbey"><code>@maxisbey</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2717">modelcontextprotocol/typescript-sdk#2717</a></li> <li>fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) by <a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> <li>chore: bump version to 1.30.1 by <a href="https://github.com/claude"><code>@claude</code></a>[bot] in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2848">modelcontextprotocol/typescript-sdk#2848</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/MukundaKatta"><code>@MukundaKatta</code></a> made their first contribution in <a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/4b0051f400219f8d8855f9a5433c6df35f15a639"><code>4b0051f</code></a> chore: bump version to 1.31.0 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2890">#2890</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/51ad4f03190dc5c84ab8b1a25f9b78b277be0dc7"><code>51ad4f0</code></a> [v1.x] Bind stored OAuth credentials to the authorization server that issued ...</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/289ac2c3af7e1536160e80414296b175171a1a87"><code>289ac2c</code></a> chore: bump version to 1.30.1 (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2848">#2848</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/12b425678a76cd54b0452a2ccf1e5dc7740f73ef"><code>12b4256</code></a> fix(auth): preserve resource URI without trailing slash (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>) (<a href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1972">#1972</a>)</li> <li><a href="https://github.com/modelcontextprotocol/typescript-sdk/commit/a9f6eb709b85459d01e8f2a9e881fef2621756c1"><code>a9f6eb7</code></a> [v1.x] fix(server): read HTTP request bodies with a size limit and bound JSON...</li> <li>See full diff in <a href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.31.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>canary/v2026.1008.0-canary.17 |
||
|
|
317ed367d4 |
Handle incomplete task policies and preserve review limits (#15593)
Apply the reviewed change for Handle incomplete task policies and preserve review limits. Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.16 |
||
|
|
3d0e74e743 |
Keep proven external connection failures out of Sentry (#15590)
Apply the reviewed change for Keep proven external connection failures out of Sentry. Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.15 |
||
|
|
0ac194450a |
fix: make Copilot provider-pack wrappers portable (#15586)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Container images include a movable provider pack for agent execution. > - pnpm executable wrappers can contain the temporary build directory. > - New optional Copilot packages add wrappers that the build does not replace. > - This pull request gives those installed wrappers relative executable paths. > - Image publication can finish while the existing path check remains enforced. ## Linked Issues or Issue Description Refs #15572 and #15560. The small portability loop comes from cryppadotta's larger Copilot runtime PR #15560. This separate fix repairs image publication without waiting for that feature's qualification and runtime changes. That PR can remove its duplicate loop after this lands. **What happened?** The standard Docker build stops with `Provider pack shim copilot-linux-x64 retains its temporary build path`. The failure occurred before and after #15522. See [the failed master build](https://github.com/paperclipai/paperclip/actions/runs/37802065316). **Expected behavior** Installed native Copilot wrappers resolve their pinned executable after the provider pack moves. Optional packages that are absent do not gain a command. The builder still rejects wrappers with temporary paths. **Steps to reproduce** 1. Run the provider-pack build from the affected master revision on Linux x64. 2. Let `pnpm deploy --prod` install the optional Copilot package. 3. The wrapper scan rejects its temporary `NODE_PATH`. **Paperclip version or commit** Master `3367b75ccce34d02f355cda1f1ed3fe0b34cf93d`. **Deployment mode** Docker and provider-pack builds. ## What Changed - Replace installed Copilot platform wrappers with relative native executable launchers. - Move the existing executable wrapper writer into an importable helper. Retain the same launch behavior for Node, Claude and OpenCode. - Test relocation, argument handling, exit status, absent optional packages and missing executable packages. Register the tests in the existing Runner test preparation command. - Hash the helper in Daytona image identity and prove that helper changes invalidate the image cache. - Document the packaging rule. Keep dependencies, provider qualification and the final temporary-path check unchanged. ## Verification - 17 Node packaging tests passed across the new wrapper tests, provider-pack release tests and candidate selection tests. - Six existing bundled remote-provider-pack tests passed using the Runner Vitest configuration. - Nine Daytona image identity tests and Runner E2E typecheck passed after the cache-input correction. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Real `pnpm deploy --prod` reproduction on macOS ARM64: the original Copilot wrapper contained the temporary path. After the rewrite and directory relocation, the actual executable returned Copilot CLI 1.0.88 with exit code zero using only `/usr/bin:/bin` in `PATH`. - Syntax checks, `git diff --check` and the pre-push secret scan passed. - The full local `pnpm test:run` reproduced the same five skill/connector fixture failures observed earlier in this workspace. It was stopped after current-head clean-checkout CI passed; later local phases were not run. This local run is not claimed as passing. Focused packaging tests, workspace typecheck/build and all hosted CI passed. - [Hosted Docker verification passed](https://github.com/paperclipai/paperclip/actions/runs/37804902678): Linux AMD64 and ARM64 image builds, multi-architecture publication and the process-reaping smoke check. This run tested `779d94989c54d9abbeba0838194c951186af66a6`; the only later changes are Daytona cache identity and its regression test. Provider-pack build code is identical. Local Docker did not respond within the bounded probe. - Current head `db2c19e270b8d5a7bab5db39d3de0e4761cf57c3`: 54 successful checks, two conditional skips, no failures and no merge conflicts. Apex is 5/5 with no unresolved comments. ## Risks The helper uses each installed package's exported executable. An installed wrapper with a missing package still fails the build. This change does not execute Copilot during image construction, alter dependency pins, change runtime admission or weaken the temporary-path check. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact serving model identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local-suite limitation is documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.14 |
||
|
|
75918997a8 |
refactor(heartbeat): extract workspace preparation and resolution (#15578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service prepares workspaces and dispatches agent runs. > - Its main file still has more than 30,000 lines after two small extractions. > - Workspace preparation forms a larger boundary with one database dependency. > - This pull request moves that code into one workspace module with focused tests. > - The benefit is a smaller orchestration file and one place to maintain workspace policy. ## Linked Issues or Issue Description **What existing behavior does this improve?** It improves the structure and test coverage of heartbeat workspace preparation. **Current behavior** `server/src/services/heartbeat.ts` mixes workspace preparation, run execution, and scheduling in 30,634 lines. **Proposed behavior** Move workspace preparation into `server/src/services/heartbeat/workspaces.ts`. Keep the public imports and run behavior stable. The main file loses about 3,540 lines in one extraction. Related extractions: #15568 and #15573. **Reason and benefit** A domain-level extraction makes useful progress toward a few manageable modules. Future workspace changes can be reviewed without searching the whole run engine. **Breaking changes** None. The existing public exports and workspace validation error class retain their identity. ## What Changed - Move managed checkout preparation, workspace validation and reuse, referenced project resolution, and session/workspace config freshness into `heartbeat/workspaces.ts`. - Bind run workspace resolution to the database through `createHeartbeatWorkspaceResolver(db)`. - Keep the checkout single-flight map at module scope so all callers share pending materialization. - Keep all 140 existing `heartbeat.ts` exports. The 84 moved function bodies and 53 other moved declarations are unchanged in a syntax-tree comparison. - Add seven tests for legacy export identity, database binding, issue selection, session fallback, concurrent checkouts, same-name repositories, and retry after failure. - Document the module boundary in `doc/DEVELOPING.md`. ## Verification - Passed: seven focused workspace suites, 264 tests. Command: `pnpm exec vitest run server/src/services/heartbeat/workspaces.test.ts server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-referenced-projects.test.ts server/src/__tests__/heartbeat-remote-referenced-projects.test.ts server/src/__tests__/heartbeat-workspace-branch-containment.test.ts server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts server/src/__tests__/heartbeat-project-env.test.ts`. - Passed: `pnpm -r typecheck`. - Passed: `pnpm build`. - Passed: all 54 GitHub CI checks on `14dec775c7fa83716432f73740bd697106fdef68`. The two optional Storybook jobs were skipped. The full CI test matrix is green. - The full local `pnpm test:run` hit a 15-second timeout in the first `public-mcp.test.ts` case. A focused rerun passed all 80 public MCP tests. The long local run was stopped after all CI test shards passed; it did not complete locally. - Greptile: 5/5 on the same commit, with no inline comments or actionable findings. ## Risks Module initialization and shared checkout state are the main extraction risks. Legacy exports retain the same function and class objects. The checkout map stays outside the database factory. Existing integration suites cover company boundaries, workspace containment, reuse, and referenced project authorization. New local Git tests cover concurrent materialization and failure cleanup. The orchestration body is unchanged except for the resolver factory binding. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.13 |
||
|
|
3367b75ccc |
fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A host can stop an idle instance to reduce unused compute. > - Zero active runs do not prove that requests, saved work, or cleanup are complete. > - A client can disconnect while its request still writes, and failed cleanup can remain only in memory or on disk. > - This pull request holds new admission and checks accepted work, cleanup, and persisted work under an owned drain. > - The host gets an empty report only while those checks remain valid. ## Linked Issues or Issue Description Refs #13413. That companion PR uses the same owned-hold vocabulary for runtime services. This PR covers HTTP requests, scheduler work, accounting, cleanup, and the instance work inventory. It does not include the preview gateway changes. **What existing behavior does this improve?** The instance task-drain API reports process counters. It does not prove that a host can safely stop the instance for idle sleep. **Current behavior** A quiet run set can coexist with an unfinished request, cleanup after a completed run, future work, or a failed accounting write. **Proposed behavior** Provide a bounded owned idle hold. Block new ingress, track accepted handler promises, inspect durable and local work, and return `none` only when the same hold stays quiet through the checks. Keep normal deployment drains compatible. **Reason and benefit** Hosts can identify eligible idle instances without treating a disconnect or failed cleanup write as completed work. ## What Changed - Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked release to task drain. Existing holds cannot be replaced by another API request. - Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new WebSocket upgrades. Track async handlers in nested Express routers and error middleware until they settle, even after the response or client disconnect. - Count accepted live-event WebSocket authentication through settlement, even after disconnects. Count detached built-in agent, managed-home and runtime-service startup reconciliation after readiness. - Keep health and control mutations tracked. Count control-request authentication separately from the read-only report, including concurrent user/company/membership writes. - Pause new scheduler admissions during idle holds. Count work already in flight, including database backups, and reject scans whose work generation changes. Periodic backups block sleep without a host wake schedule. - Inspect accounting and orphan-cleanup spool directories without skipping temporary or malformed entries. Retain orphan tokens through queue splices, flush failures, and buffer overflow. Keep failed usage capture counted until its database failure fence is written. - Check persisted work across companies in a bounded read-only transaction. Include deferred agent-file cleanup and saved watchdogs whose watched issues are complete. Enabled plugins and unsupported retained work remain blockers. - Require the exact idle owner and completed startup tracking before returning an empty report. Keep reports free of tenant details. - Document the hosting protocol, retry behavior, conservative blockers, and the remaining external provider-stop race. ## Verification - Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper extractions are preserved. GitHub confirms no merge conflicts. - Focused heartbeat renderer/run-log, drain, control-auth, admission, route and PostgreSQL inventory checks: 234 tests passed in 10 suites after the rebase. - The POST task-drain contract includes the expected `409` conflict response. The 84 OpenAPI and instance-settings route tests and server typecheck passed after that final documentation fix. - Final accepted-upgrade/startup regression run: 104 tests passed in five suites, including success and failure after readiness or disconnect. Final server typecheck also passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm install --frozen-lockfile` and the PR diff secret scan passed. No dependency or lockfile changes are added by this PR. - The prior verification covered HTTP disconnects and early responses, async error handlers, concurrent authentication writes, signed bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and failed accounting fences. Those tests passed. - The last full local `pnpm test:run` stopped in the general-server phase with 16,658 passing tests and five skill/connector fixture failures caused by an ancestor workspace skill directory. That full local run preceded this rebase and has not been repeated for the import conflict. Full current-head CI passed: 53 successful checks and two conditional skips, with no failures. - All eight review findings are fixed and their threads resolved, including accepted upgrade authentication, detached startup writes and the POST conflict contract. Current-head Apex review is 5/5 with no new findings; all eight review threads remain resolved. - No live provider stop or production change was performed. ## Risks - The hosting controller must use the owned protocol and hold external admission through its final validation and provider stop. A legacy drain cannot authorize idle sleep. During an idle hold, new requests receive 503 with `Retry-After: 1`; the host must handle queueing or retry before enabling this path. - This is a single-process protocol. An unexpected restart after the last validation can race an external provider stop. The host must serialize deploy/wake/sleep operations and bind the validation to the instance it stops. Multiple replicas need shared fencing. - Some retained state conservatively prevents sleep, including every enabled plugin. This PR does not promise that every inactive instance becomes eligible. - The HTTP adapter uses Express 5 router layers. Real Express tests cover nested routes, errors and disconnects. New routes must register before tracking is installed. Detached work must have durable state or explicit work tracking. - Unrecoverable in-memory cleanup debt keeps the instance awake until reconciliation. This change does not make such debt survive an unplanned process crash. No schema migration or provider configuration change is included. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use and code execution. The exact deployment variant/model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; the full local fixture limitation and clean full CI result are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
71af2fbc3b |
fix(slack): teach agents how people connect their accounts (#15576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections let people talk to an agent with their own Paperclip permissions. > - Each bot has a saved command that starts account linking. > - The Access page explains this flow, but Slack agent turn context omitted it. > - An agent could guess the command or confuse channel membership with Paperclip access. > - This pull request gives agents the saved command and account confirmation instructions. > - People can ask the agent how to join without changing access controls. ## Linked Issues or Issue Description **What happened?** Slack agent context explained messages, tools, and questions. It did not explain how a teammate can connect an account. The saved slash command can differ from the current agent name. **Expected behavior** Give the saved bot command, such as `/research_ops connect`. The teammate runs it themselves. They sign in through a private confirmation link. A new company member must receive admin approval before account confirmation. The connection manager can copy instructions from Access → Invite people. **Steps to reproduce** 1. Configure a Slack bot with a custom slash command. 2. Rename its assigned agent. 3. Inspect a fresh or resumed Slack task prompt. Before this change, it contains no account invitation instructions or saved connect command. **Paperclip version or commit** Base: `d6df12cef69fcaf2d2fe66a393168931d5b8b4e7`. **Deployment mode** Applies to local and hosted Slack chat connections. Deterministic tests used a local isolated database. The new model probe has not run against a live Slack bot. Related work: Refs #15413 and #13638. This fixes agent guidance for their existing account-linking flow. It adds no new membership system or invitation endpoint. ## What Changed - Read only the saved public slash command from the company-scoped conversation endpoint. Supply it only to the assigned agent for Slack turns with an active or verifying connection. - Validate the command with the shared Slack configuration schema. Refer to Access → Invite people when it is missing or invalid. Never guess from the current agent name. - Explain personal account confirmation, link expiry, and company membership approval on fresh and resumed turns. Distinguish channel invitations from Paperclip access. - Add deterministic guidance regressions, a manual invitation model probe, and setup documentation. ## Verification - Passed: 63 tests in `heartbeat-context-summary.test.ts` and `heartbeat-chat-task-link.test.ts`. - Passed: four real-heartbeat regressions in `heartbeat-slack-invitation.test.ts`. They check the saved command after an agent rename, a persisted resumed session, full and compact prompts, missing-command fallback, inactive endpoints, and a different assigned agent. They also reject caller-supplied command fields and exclude other setup metadata. - Passed: the isolated `chat-channels.integration.test.ts` case `discovers a Slack connect identity without starting work or granting access`. It covers the private link, duplicate connect requests, and nonmember access requests without a membership grant. - Passed: `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/server build`. - Passed: `pnpm test:slack-connector --list` and `git diff --check`. - Passed: repository `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. Typecheck required local IPC access for the migration check. - Passed: final-commit CI, with 54 successful checks and two optional Storybook checks skipped. CI includes all server, chat, workspace, serialized, Runner, and browser test shards, typecheck, build, and canary dry run. - The full local `pnpm test:run` was started. It was stopped after full CI passed; it had no final local summary. The focused invitation checks passed locally. Do not count the interrupted local run as a full-suite pass. - Greptile gave the exact final commit `f37e52bde621ef7f5b2bb345074d53345f8e4ee9` a 5/5 score. Both review findings were fixed and their threads resolved. - The new `invite-person` model probe is manual. These deterministic results do not establish a live model or Slack acceptance pass. ## Risks - Model guidance cannot prove that a person joined. The existing account-linking and membership checks remain authoritative. - Legacy rows without a saved command use the Access page fallback. Invalid command text is excluded from the prompt. - The query selects only the public command. It does not expose registration secrets, tokens, or personal confirmation links. - No schema changes, new provider requests, permission grants, or telemetry changes. ## Model Used - OpenAI GPT-6 through Codex. The exact backend variant and context window are not exposed in this session. Used reasoning, repository inspection, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.12 |
||
|
|
d6df12cef6 |
fix(ui): add icons to chat connection choices (#15574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connectors let people chat with agents and let agents use external tools. > - Slack setup asks the user to choose between those two uses. > - Both choices currently use text alone. > - This pull request adds an agent avatar and a hammer to the shared chooser. > - The icons help users identify each choice at a glance. ## Linked Issues or Issue Description **What existing behavior does this improve?** The chat-or-tools choice in connector setup. Related setup work: Refs #15413. **Subsystem affected** ui/ — React board UI and Storybook. **Current behavior** Both choice cards show a title and description with no identity icon. **Proposed behavior** Show the existing agent avatar beside chat. Show a hammer beside tools. Align both icons and keep the existing navigation. **Reason and benefit** Users can tell the two choices apart with less reading. **Breaking changes** None. The shared chooser applies the icons wherever that choice appears. Only Slack currently offers both paths in the catalog. Other chat-only providers keep their direct setup path. ## What Changed - Add the shared agent avatar to the chat choice. - Add a decorative hammer icon to the tools choice. - Add desktop and mobile Storybook stories that render the production page. ## Verification - Token gates pass. - Focused catalog, routing, and UI contract tests pass, including the new chooser accessibility and navigation test. - Checked the production chooser in Storybook at desktop and mobile widths. - Local repository typecheck, build, and Storybook build pass. - All 56 checks pass on commit `0f0673f1bb5cd340084f3a5229983da049f1c8db`, including the full CI test suite and browser shards. Greptile reports 5/5 with no findings. - Started the full local test command. Stopped that duplicate run after the full CI suite passed. The focused local tests completed successfully. - In Storybook, open Connections → Slack → Automatic setup → 00 · Choose connection. The mobile variant is in the same group. ## Risks Low risk. This changes card layout and decorative images. Existing button text, click handlers, credentials, and permissions stay the same. The avatar is generic because the user has not selected an agent yet. ## Model Used OpenAI GPT-6 through Codex. Used code inspection, edits, shell tools, and browser tools. The session does not expose the exact runtime model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
941a3fa991 |
Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass. Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com> Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
87a7312cfd |
refactor(server): extract heartbeat run-log formatting (#15573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat orchestration records agent output and run events. > - The heartbeat service still contains more than 30,000 lines after the first extraction. > - Log formatting has fixed limits and does not write run state. > - These helpers form a small next step toward a more manageable heartbeat service. > - This PR moves them into the heartbeat folder without changing their function bodies. > - Direct boundary tests and existing caller tests check the output. ## Linked Issues or Issue Description Refs: #15568. **What existing behavior does this improve?** This improves the structure and test coverage of heartbeat run-log formatting. **Subsystem affected** server/ — orchestration services. **Current behavior** `heartbeat.ts` contains excerpt handling, payload size limits, and log chunk formatting beside run orchestration. **Proposed behavior** Move these helpers and their constants into `server/src/services/heartbeat/run-log.ts`. Keep the same function bodies, output, limits, and public exports. **Reason and benefit** Run-log formatting has its own small module and focused tests. This keeps the second extraction small enough to review on its own. **Breaking changes** None. The existing public import path and persisted output stay the same. Related run-log PRs: Refs: #6373, Refs: #8841. Those PRs change redaction behavior. This PR only moves existing formatting code. ## What Changed - Move 109 lines of helper functions and six constants into `heartbeat/run-log.ts`. - Move `appendExcerpt` and retain the existing public exports for `boundHeartbeatRunEventPayloadForStorage` and `compactRunLogChunk` in `heartbeat.ts`. - Add 17 direct regression cases for size and depth limits, cycles, shared references, immutable input, image omission, redaction order, excerpt tails, and UTF-8 boundaries. - Document both heartbeat extractions in `doc/DEVELOPING.md`. ## Verification - Before extraction, four focused files passed with 55 tests. - After extraction and the two new excerpt tests, the same four files passed with 57 tests. - Run `pnpm exec vitest run server/src/services/heartbeat/run-log.test.ts server/src/__tests__/heartbeat-run-log.test.ts server/src/__tests__/heartbeat-list.test.ts server/src/__tests__/redaction.test.ts`. - The original and extracted function blocks match byte for byte after adding the export keyword to `appendExcerpt`. - Existing tests still import the public helpers from `heartbeat.ts`. - `node scripts/check-module-boundaries.mjs` and `git diff --check` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - The full suite passed in CI. The duplicate local `pnpm test:run` was stopped after CI passed; it did not complete locally. - All CI gates passed on commit `4e65c666e3b5b46162b3a528b37fba8902a1d200`: 54 passing checks, two skipped Storybook checks. - Fresh Greptile review of the same commit: 5/5, no actionable or inline findings. The PR has no merge conflicts. ## Risks - Low risk. Moving code can cause an import or build error. - The same redaction and adapter utility modules remain in use. - Database writes, current-user redaction, live event delivery, and run state remain in `heartbeat.ts`. - The event schema and payload output do not change. - No database, API, Telemetry, Observability, or UI contract changes are required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51653becc2 |
refactor(server): extract heartbeat task markdown rendering (#15568)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Heartbeat orchestration supplies each run with task context. > - The task Markdown renderer is inside a service with more than 31,000 lines. > - The renderer formats task data and does not write run state. > - It is a small first step toward a more manageable heartbeat service. > - This PR moves the renderer without changing its function body or public export. > - Direct tests and existing caller tests check the output before and after the move. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the structure and test coverage of heartbeat task context rendering. **Subsystem affected** server/ — orchestration services. **Current behavior** `buildPaperclipTaskMarkdown` occupies 352 lines inside `heartbeat.ts`. The service also owns run execution, scheduling, recovery, and cancellation. **Proposed behavior** Move the renderer to `heartbeat/task-markdown.ts`. Keep its function body and the existing `heartbeat.ts` export unchanged. Leave orchestration in place for this first PR. **Reason and benefit** Prompt formatting has its own small source file and direct regression tests. Reviewers can check this one extraction before a later refactor. **Breaking changes** None. The input type, rendered output, and existing import path stay the same. Related renderer changes: Refs: #4732, Refs: #14030. These PRs change prompt behavior. This PR only moves the current renderer. ## What Changed - Move the 352-line renderer to `server/src/services/heartbeat/task-markdown.ts`. - Import and re-export it from `heartbeat.ts`. Remove imports used only by the renderer. - Add seven direct regression cases. Cover empty context, ordered comment-only wakes, input preservation, nested code fences, ancestor limits, attachment-only wakes, and rejected plans. - Keep the renderer and its direct tests in `server/src/services/heartbeat/`. Document this folder as the home for relevant later extractions in `doc/DEVELOPING.md`. ## Verification - The focused suite passed before and after extraction, and after moving to the heartbeat folder: four files, 85 tests. - Run `pnpm exec vitest run server/src/services/heartbeat/task-markdown.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/heartbeat-chat-task-link.test.ts server/src/__tests__/codex-local-execute.test.ts`. - The extracted function matches the original function byte for byte. The folder move only adjusts imports. - `node scripts/check-module-boundaries.mjs` passed. - `git diff --check` passed. - The initial extraction passed local `pnpm -r typecheck` and `pnpm build`. - The folder update passed `pnpm --filter @paperclipai/server exec tsc --noEmit` and `pnpm --filter @paperclipai/server build`. - The initial extraction passed the full CI suite. Its duplicate local `pnpm test:run` was stopped after CI passed; it did not complete locally. - The folder update passed all CI gates on commit `e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 54 passing checks, two skipped Storybook checks. Three test shards passed after one retry following simultaneous runner shutdowns; those failures had no failed test assertions. - Fresh Greptile review of commit `e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 5/5, no actionable or inline findings. ## Risks - Low risk. A moved module can change import resolution or expose an import cycle. - Existing caller tests still load the compatibility export from `heartbeat.ts`. - The same guidance constants and public task URL resolver remain in use. - No run-state writes, transactions, locks, shared process state, or cleanup paths moved. - No database, API, telemetry, or UI contracts changed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. The exact serving model ID and context window were not exposed in this session. Used reasoning, repository inspection, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.11 |
||
|
|
24a15c5afa |
test(evals): verify durable waiting across continuation checkpoints (#15564)
## Thinking Path > - Paperclip manages agents and durable tasks across provider runs. > - Human questions must preserve task ownership and stop work until a real answer arrives. > - The continuation eval checks the final work, but its lifecycle oracle misses several broken waiting states. > - A correct final answer can hide a stale execution lock or a lost intermediate answer receipt. > - This pull request checks each wait and retains every question and run identity. > - A source audit records which waiting operations belong to native runtimes and which still require legacy API calls. ## Linked Issues or Issue Description **What existing behavior does this improve?** The existing Product E2E continuation oracle for human question and approval journeys. **Current behavior** A pending interaction can pass the waiting check even when its task has the wrong status, a stale lock, or a retry. Only the first pending question and first run receipts are checked at the final checkpoint. **Proposed behavior** Require a healthy wait on the same task and assignee. Preserve all intermediate run receipts and every pending question's answered identity. Allow a paused native provider question only when its pending runtime request identifies the running native run and execution lock. Related: #15548 and #15554 cover earlier bookkeeping slices. #15544 changes production continuation summaries; this PR changes the eval oracle and does not overlap that fix. ## What Changed - Check waiting state, execution ownership, and all answer/run identities in the existing lifecycle oracle. - Add negative calibrations for broken records and positive coverage for both semantic waits and paused provider questions. - Retain task activity at each checkpoint for inspection of successful persisted mutations. - Verify 1,000-cent company and agent budget limits before continuation work. Disable automatic cell rerolls. - Document runtime ownership, existing coverage, and the remaining instruction decision. Keep production instructions unchanged. ## Verification - `pnpm test:e2e:runner:unit`: 1,867 Vitest tests pass, one skips; 128 Node tests pass. - `pnpm test:e2e:runner:typecheck`: passes. - `pnpm -r typecheck`: passes. - `pnpm build`: passes. - Negative calibration: 27 added cases fail against the prior oracle and pass with these checks. - Four existing local continuation cells are selected for a separate bounded live canary. Live results are pending; no behavioral pass is claimed here. - The full local `pnpm test:run` suite was not repeated because embedded PostgreSQL was unavailable in the preceding workspace verification. Required Linux PR CI must pass before readiness. ## Risks The stronger oracle can expose existing product or fixture defects. A paused provider question and a terminal semantic wait have distinct valid states. The four-cell canary does not qualify approval/review, dependency unblock, crash races, remote execution, or general task quality. Task activity records successful persisted writes, not failed API attempts; repeated progress comments are not automatically defects. Historical eval grades remain unchanged. No production scheduling, prompt, tool, schema or migration changes. ## Model Used OpenAI Codex, GPT-6-based, with tool use and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f47614046d |
fix(slack): upload agent avatars directly during Cloud setup (#15566)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Slack setup creates a dedicated bot for an agent. > - The setup should upload that agent's avatar. > - The current upload asks Slack to fetch an image from the board origin. > - Cloud requires a tenant session at that origin, so Slack receives HTTP 401. > - This change renders the PNG on the server and uploads the file directly. ## Linked Issues or Issue Description Refs #15413. **What happened?** Automatic Slack setup left the app and bot with default icons on Cloud staging. An unauthenticated request to the exact avatar URL returned HTTP 401 with `tenant_session_required`. Local setup did not expose this Cloud ingress requirement. **Expected behavior** New Slack bots receive the assigned agent's 512-pixel avatar with the Paperclip dark background. **Steps to reproduce** 1. Create a Slack app through automatic setup on a Cloud tenant. 2. Complete installation. 3. Inspect the bot avatar in Slack. The previous URL-based upload cannot fetch the image without a tenant session. ## What Changed - Render the assigned agent's preset PNG with the existing bounded worker pool. - Send PNG bytes as multipart `file` data to `apps.icon.set` instead of passing a board URL. - Keep the temporary token in the Authorization header. Let fetch set the multipart boundary. - Recheck management permission and credential-lease ownership after rendering. Close the worker pool during chat service shutdown. - Cover actual PNG dimensions, uploaded bytes, failure recovery, and secret-safe responses. Update deployment documentation. ## Verification - Passed: 84 focused tests across automatic Slack registration and on-demand agent avatars. - Passed: full repository build. - Passed: full repository typecheck. All 54 current-head GitHub checks passed, including the complete test matrix, all browser shards, build, typecheck, canary, and security checks. Greptile completed on `79e79f6fc` with 5/5 and no actionable findings or open review threads. - The Cloud fetch failure was reproduced without browser credentials. No Cloud access rule was changed. - A real Slack upload with this new path still requires deployment and a fresh automatic setup. Existing apps retain the manual avatar-upload fallback. ## Risks - The renderer can time out or Slack can reject the upload. Both failures preserve the saved app and leave installation usable. - The renderer adds a bounded, lazy worker pool to Slack registration. Shutdown closes it. - No migration, bot permissions, credential retention, or Cloud authentication behavior changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser inspection. The runtime does not expose a more specific authoring model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.10 |
||
|
|
2f0c485dec |
fix(skills): ship the completion helper with the installed skill (#15554)
## Thinking Path > - Paperclip manages work for AI agents. > - Legacy agents use the Paperclip skill to save task status and comments. > - The skill names a script relative to the task workspace. > - That script exists only in the Paperclip source repository. > - Agents in other workspaces can hit a missing command or search for it. > - This PR ships the helper inside the skill and uses the installed skill path. > - The repository command remains available through a forwarding wrapper. ## Linked Issues or Issue Description Fixes #9527. Refs #15548 for the preceding runtime checkout guidance. Related: #6052 addresses LF line endings for the repository helper; this change addresses helper delivery and path resolution. ## What Changed - Bundle the existing issue update helper with the Paperclip skill. Preserve its HTTP checks, echoed-status check and two-attempt limit. - Resolve the command from the installed skill directory. Use a verified PATCH when that path is unavailable, without searching the filesystem. - Keep the repository command as a wrapper that works from any directory. - Test shell execution and exact status/comment payloads through both provider skill-home layouts, including paths with spaces. - Add helper sources and existing verification tests to stock-harness admission. Record an absent historical helper explicitly. Add the missing declaration for the admission fingerprint export. ## Verification - Complete directly affected source suites: 30 tests pass. They cover skill delivery, preserved multiline comments and links, authentication headers, empty responses, mismatched status, transient retries and definitive rejections. - Product E2E typecheck passes. Support suites: 1,835 Vitest tests pass, one is skipped; 128 Node tests pass. - Full local build and workspace typecheck pass. - Full local repository tests are not claimed as passed. Embedded PostgreSQL was unavailable in this worktree during the preceding task; Linux CI will run the repository gates. - The authorized matched Codex/Claude comparison is pending. It uses the existing assigned-skill case and original oracle, one initial attempt per profile and variant. - CI and a fresh Greptile review are pending. Keep this PR in draft until readiness gates complete. ## Risks - Correct path resolution depends on the harness supplying the installed skill path. The instructions use verified PATCH when that path is unavailable. - The helper still requires Bash, curl and jq. Its existing retry and response-verification behavior is unchanged. - Tests use the shared skill-directory symlink mechanism and an HTTP fixture. Real provider completion behavior still requires the bounded live comparison. - This fix does not redesign native completion, legacy recovery or ambiguous transport handling. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.9 |
||
|
|
71cd0a2621 |
fix(skills): honor the current run harness checkout (#15548)
## Thinking Path > - Paperclip manages work for AI agents. > - The runtime claims eligible assigned tasks before it starts an agent. > - The wake tells the agent when the runtime already holds that claim. > - The legacy skill still requires another checkout in every case. > - This PR makes the skill honor the current task and run claim. > - Manual checkout and server ownership checks remain in place for other cases. ## Linked Issues or Issue Description **Where is the issue?** `skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5. **What's wrong?** The wake can say that the harness already checked out the issue. The skill still tells the agent that it must call checkout. These instructions conflict. **Suggested fix** Skip the second checkout only when the runtime wake explicitly confirms the claim for this issue and run. Retain manual checkout when that statement is absent or the agent selects another task. Refs #14948 for the existing shared prompt reduction. ## What Changed - Honor the explicit runtime claim in the scoped wake procedure and Step 5. - Keep context reads, status writes, deliverable handling and conflict rules. - Add checks for normal and resumed wake text and excluded automatic claims. - Retain successful checkout HTTP activity for legacy stock-task evals. Bind each receipt to the exact company, task, agent and run. Keep this observation separate from the original task grades. ## Verification - Checkout observation calibration: nine tests pass. - Focused skill, wake and database ownership tests: in progress. - Full repository build, typecheck and tests: in progress. - Planned live comparison: the existing assigned-skill document case on legacy Codex and Claude. One attempt per variant and profile. No automatic retries. The baseline and candidate share the observation code and task oracle. - Live results are pending. This draft does not claim behavioral qualification. ## Risks - Agents may misread prompt guidance. The API still enforces ownership; the text grants no new authority. - The exception is specific to the current issue and run. It does not remove ordinary legacy completion writes or authorize another task. - Activity measures successful checkout HTTP calls. Failed attempts require separate run-log inspection. Missing or mismatched observations cannot count as zero calls. - One trial per profile cannot establish general reliability, speed or cost trends. ## Model Used OpenAI Codex, GPT-6 family. The exact model build and context window are not exposed in this session. Used code editing, shell tools and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.1008.0-canary.8 |