mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
1483bb8bcf5571bbec68db22119f7c685290d970
27
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
84fe89906d |
fix: complete native agent review handoffs (#13581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native execution uses durable runs, issue locks, wake requests, and typed tool authority > - A child can finish with a native agent review request while its original assignee stays responsible for the work > - The reviewer then needs a bounded execution path that can inspect the child, record one decision, and finish safely > - Before this change, assignee-only gates rejected the reviewer or left the parent waiting after the child review ended > - This pull request adds typed reviewer admission, scoped reviewer tools, durable wake and recovery handling, and parent continuation evidence > - The benefit is that native review handoffs complete without changing child ownership or granting broad mutation access ## Linked Issues or Issue Description Refs: #13314 Refs: #13574 **What happened?** A native child run could report `needs_review` for an agent reviewer. The reviewer wake then failed assignee and execution-lock checks. The child remained in review and the parent remained waiting. **Expected behavior** The named reviewer should receive one durable wake. The reviewer should inspect the child and resolve the exact review card. The child assignee should stay unchanged. The parent should receive the recorded review outcome after the child reaches its terminal state. **Steps to reproduce** 1. Run a native task with a different named agent reviewer. 2. Keep the child assigned to its original worker. 3. Let the worker finish with a native completion review request. 4. Start the durable reviewer wake. 5. Resolve the review and finish the reviewer run. 6. Observe the child and parent state. **Paperclip version or commit** Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification source: `eea171aae`. **Deployment mode** Built from source. **Installation method** Built from source (pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Access context** Both. **Database mode** Embedded PostgreSQL in the isolated live test fixtures. ## What Changed - Add server-validated native review assignment facts. - Admit only the exact company, issue, source run, decision, revision, addressee, and resolver policy. - Give reviewer runs a narrow set of Paperclip read and resolve tools. File and shell access follow the configured agent and environment policy, so reviewers can run tests. - Separate server-owned reviewer instructions from untrusted persisted review data. Escape the data boundary; retain server-enforced authorization. - Keep the child assignee unchanged. Atomically claim the reviewer run, wake request, and issue execution lock. A competing lock prevents provider startup. - Require the exact running reviewer session and current issue lock to resolve its assigned card. Reject missing, unrelated, or terminal reviewer runs. - Add durable reviewer wake, lock, stale-card, and abandoned-run recovery handling. - Prevent duplicate native wake dispatches during deferred admission and recovery. - Carry accepted or rejected child review outcomes into parent task context and continuation evidence. - Add focused server, runner, and native protocol coverage. - Preserve upstream continuation rules. Add child review decisions as separate evidence, while keeping real human answers in their own field. - Return actionable completion validation feedback to both providers. Permit a corrected completion after rejection. Keep strict terminal acknowledgment validation. - Apply exclusive shared-workspace locks to sandbox environments. Local and SSH folders can run concurrently, including when old settings request serialization. - Repair test timing, native event parsing, and the review artifact assertion. Allow a valid reject, correct, and accept review sequence. Check the accepted card against its reviewer run and decision. Keep polling within the existing deadline when review acceptance precedes the parent wake projection; report a specific missing-continuation error at timeout. - Apply the ACPX pending-call limit to reserved finish/block calls, with capacity-release and cancellation tests. ## Verification - `pnpm build`: passed on `eea171aae`. - `pnpm -r typecheck`: passed on `eea171aae`. - `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on `b31ad9ab8`; runner E2E typecheck also passed. - `pnpm check:token-gates`: passed. - Focused DB review, reviewer authority, and prompt-boundary checks: 31 tests passed. They cover invalid reviewer runs, competing locks, atomic admission, duplicate claims, and valid resolution. - Heartbeat, workspace, and recovery checks: 30 tests passed. - ACPX sidecar suite: 27 tests passed. Moving the capacity guard back below reserved handling makes both new regression cases fail. - Four focused live continuation checks passed on their first attempt at `f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and question tool guidance (12/12 each). These cases do not use the reviewer prompt path changed afterward. - Fresh Codex and Claude review-handoff checks passed all 29 native checks each on their first attempt at `eea171aae`. Both runs received the expected fixed prompt and completed cleanup. Only the six selected live flows were tested; no full paid provider catalog run. - The final commit only extracts the existing test-harness timeout diagnostic into a shared helper and adds positive and negative coverage. Removing the accepted-review guard makes two regression assertions fail; restoring it passes all six timeout tests. Production runtime code, prompts, deadlines, and grading criteria are unchanged by this final commit. - Deadline regressions: a valid continuation delayed 20 seconds succeeds within its 30-second unit-test deadline; an absent wake returns a specific candidate-failure diagnostic at that same deadline. Both assertions failed before the fix. Production E2E deadlines remain unchanged. - Historical native failures remain recorded: Docker availability failures; a valid reject/correct/accept sequence that the first-card grader misread; and a test that rejected the gap between accepted child review and parent wake projection. No failed result was regraded. The latest tests use a protected reference to the pinned Docker image and the unchanged artifact oracle and time limits. - Full repository verification runs in GitHub CI. Local verification uses the focused suites above, full build, and full typecheck. An unchanged Codex shutdown timing test failed once in CI, passed in isolation, and its full shard passed on the final commit without changes to that test or its causal code path. The original failure is retained in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no outstanding actionable findings. All review threads are resolved. All current-head CI gates passed, including the isolated native runner Docker build (55 successful checks; two skipped by the workflow). ## Risks - Reviewer admission depends on exact persisted decision and interaction bindings. A stale or changed card is rejected. - Paperclip control-plane tools are limited to inspection and review resolution. This is not a filesystem permission boundary; provider file and shell access retain the configured policy. - Deferred wake recovery changes dispatch receipt coalescing. A scheduler regression could delay a continuation if the receipt state is wrong. - Parent review outcomes are evidence for the model. They do not grant tool authority or change issue ownership. - This change does not address legacy lease-hold handoff behavior. > Roadmap review: native execution, review gates, and durable recovery are existing roadmap capabilities. This PR completes a narrow reliability path for those capabilities. ## Model Used OpenAI `gpt-6-astra` with reasoning, tool use, and code execution. OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and journal work. Context window size is not exposed by this session. Live test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the PR authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0b67bfe71 |
feat: queue approvals and answers during active runs (#13539)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users guide running agents through messages, questions, and approval cards. > - Messages already wait in a queue when an agent is running. > - Card responses did not appear in that queue. Some question answers also steered a later run without a user click. > - A fast approval could invalidate the agent's review handoff and cause it to stop its own run. > - This pull request gives card responses the same queue controls and preserves the exact response during delivery. > - Users can wait for completion or explicitly send the response with Interrupt or Steer. ## Linked Issues or Issue Description Refs #13517, which is merged. This PR targets master and adds queued interaction responses on top of the onboarding changes. Related continuation work: #10519 and #12866. **What happened?** Accepting a proposal while its source run was active left a saved response outside the message queue. The agent could then lose its review path, reassign the task, and cancel itself. Answers to older questions could also steer another active turn without a click. **Expected behavior** Save the response immediately. Queue its continuation behind the active run. Deliver it after completion, or when the user explicitly chooses Interrupt or Steer. Preserve approval revisions and answer choices. **Steps to reproduce** 1. Let an agent publish a confirmation card while its run is still active. 2. Accept the card before the agent finishes its review handoff. 3. Inspect the message queue and the task's next run. **Paperclip version or commit** Reproduced on da8a3876c with the onboarding changes from #13517. **Deployment mode** Local development from source. The fix covers legacy adapters and native Runner turns. ## What Changed - Project resolved cards into the existing queue as immutable responses. Keep answers and exact approval revisions. - Require an explicit click to steer a response into a compatible native turn. Use Interrupt when a fresh session is required. - Preserve typed response context through interruption, cleanup waits, and normal queue promotion. Keep the direct answer channel for a provider blocked on its original question request. - Accept the source run's review handoff after its card resolves. Reject stale agent reassignment that would orphan a queued response. - Add deterministic regression tests and an `accept-while-running` case to the first-task suite. Require recorded timestamp overlap before that case can pass. - Keep the first-task skill name out of user-facing messages. ## Verification - Red-green: the original route failed the queue regression; the changed route passes it. - Focused server/UI tests: 139 passed, including 64 queue-route tests. - Runner harness unit tests: 314 passed. - Server, UI, and Runner E2E typechecks passed. UI token gates passed. - Full repository typecheck and build passed. Server typecheck passed again after review fixes. - Review regressions: 165 queue/reopen route tests, 53 wake admission tests, and 18 run identity tests passed. Approval acknowledgement recovery and both message/approval arrival orders are covered. - Full local test run: 12,401 passed; three new admission regressions ran against a cached pre-fix module. A fresh run of that entire suite passed (53 tests). The complete CI suite passed on the final commit. - Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional Storybook checks skipped. Every server/workspace/browser shard, Runner verification, build, typecheck/release registry, canary, policy, and security check passed. Greptile: 5/5, no unresolved threads. Earlier interrupted CI workers were replaced by this fresh complete run. - After integrating the updated parent: 314 harness tests, 119 queue/admission tests, 44 onboarding/question-delivery tests, and 13 native recovery tests passed locally. Full repository typecheck and build passed. - Clarified the skill wording preference: routine replies describe the action without announcing the internal skill; direct questions and permission/security/execution disclosures remain truthful. - The paid `accept-while-running` scenario is registered for all four local first-task profiles. It has not been run against a model in this change. - Rebased onto the merged parent at `11921075a`; the resulting tree exactly matches the locally verified integration tree. Final-head CI on `b53054807` passed: 54 successful checks, 2 optional Storybook checks skipped, no failed checks. Every new server/browser shard, aggregate verify/e2e gate, Runner, typecheck, build, canary, and security check passed on the first attempt. Greptile reviewed this exact head at 5/5 with no unresolved threads. ## Risks - Responses now wait instead of implicitly steering another active turn. A provider blocked on the original question still receives its answer directly. - Approval receipts cannot be edited, discarded, or reordered as comments. This preserves the recorded decision. - Interruption must still prove that the prior execution stopped. The tests cover cleanup waits and duplicate delivery. - The new paid overlap case can be unexercised if the model finishes before the click lands. It cannot pass without evidence of overlap. - No database migration is required. This repairs the existing approvals and execution controls; it does not implement the roadmap's work-stream queues. ## Model Used OpenAI GPT-6 through Codex. The exact deployed model ID and context-window size were not exposed in this session. Capabilities used: agentic reasoning, repository inspection, code editing, terminal commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cceeb0aa66 |
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e9b24c821 |
fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Concurrent task runs can share a managed AI subscription with one credential lease. > - #13438 added durable retries for busy subscriptions and task-lock checks. > - A comment wake can run for an agent who is not the task assignee. Such a run never owns the task lock, so the new checks suppress its retry. > - This pull request preserves those comment wakes while keeping the lock checks for assignee runs. > - It also shows the subscription wait in task status and preserves the failure count through the database projection. ## Linked Issues or Issue Description Refs #13438. **What happened?** When a subscription is busy, a non-assignee comment wake is cancelled without a successor. Repeated subscription waits also display an inflated attempt count because the execution query omits the preserved failure count. **Expected behavior** An eligible comment wake waits and resumes without claiming the assignee's task lock. The task shows “Waiting for AI subscription”. Waiting does not consume provider-failure retries. Assignee retries still stop when the task lock is cleared or transferred. **Steps to reproduce** 1. Configure an agent with a managed subscription and concurrent runs. 2. Hold its credential lease in one run. 3. Mention the agent on a task assigned to another actor. 4. Release the lease and inspect whether the comment wake has a scheduled retry. The test uses a real embedded Postgres database and a real credential lease. Provider execution uses a fixture. The reassignment test applies a database mutation after the real checkout. No live provider account is needed. **Paperclip version or commit** Based on `f912ecaac` from #13438. Before the reconciliation fix, the rebased regression at `2063cfd13` failed because the comment wake had no scheduled retry. The original configuration failure was reproduced on `5282cabde` before #13438 merged. **Deployment mode** Built from source with embedded Postgres for local verification. ## What Changed - Record non-assignee comment-wake authority at admission while holding the task and run locks. A later reassignment cannot grant this exception. Preserve it through repeated subscription waits. - Keep master’s lock checks for assignee retries, including the check inside the scheduling transaction. - Show the subscription wait and preserve its failure count in the execution projection query. - Limit pre-provider wait receipts to fresh executions. A persisted native execution input must retain its recovery path. - Add lease, retry, projection, cancellation, reassignment, pause, revocation, service-recreation, and ownership regression coverage. - Use master’s 60–120 second retry interval and document the resulting behavior. ## Verification - Red: after rebase, the comment-wake test failed with no scheduled retry; the other 50 contention and retry-scheduling tests passed. - Red/green: reassigning the task immediately after real checkout reproduced an unwanted successor for both assignment and comment wakes. Both tests pass after recording authority at admission. - Green: all 285 targeted tests pass across 12 suites, including #13438’s four cancellation-race phases, its hiring/connection suite, run-dispatch integration tests, and both reassignment regressions. - `pnpm -r typecheck` passes after rebase. - `pnpm build` passes on the final revision. - The full general and serialized test suites pass in [CI run 34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681) on `99f1e4c1d`. Full-suite verification ran in CI; local verification used the 285 targeted tests. - All 32 active PR checks pass, including build, typecheck, native runner verification, canary dry run, and all browser shards. Two Storybook checks are skipped by path filters. - Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All review threads are resolved. - `git diff --check` and `node scripts/check-module-boundaries.mjs` pass. ## Risks - The non-assignee exception requires authority recorded under admission locks or its server-created subscription retry. Tests cover caller-supplied flags, assignment and comment reassignment races, repeated comment waits, and assignee lock protections. - A wait uses master’s 60–120 second interval. A long-held lease can produce multiple wait records. - Existing native sessions can have provider effects. They must not receive a fresh-execution receipt that permits replay. - No schema migration or credential permission change is required. ## Model Used OpenAI Codex, model `gpt-6-astra`, with repository search, code editing, shell execution, and test tools. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f2c5e54dca |
fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - A task keeps its instructions, progress, and files when its assigned agent changes. > - The replacement runner lost the interrupted run's context and could overwrite an existing draft. > - A saved message also stayed attached to the former agent and could reopen the task after the replacement finished. > - File tasks could report Done with only a local path that the user could not open. > - This pull request transfers handoff context and saved messages, and makes requested files accessible through the existing attachment contract. > - Users can change agents and collect completed work without repeating instructions or confirming bookkeeping. ## Linked Issues or Issue Description Refs #13338. Builds on merged #13354 for queue admission and #13353 for remote workspace retry. #10123 concerns restricted recovery-model escalation; this change instead covers ordinary native handoff and file completion. **What happened?** Codex wrote a draft before a user assigned the task to Claude. The replacement lacked continuation context and replaced the draft. A queued user message could later restart the former agent and reopen the completed task. Separately, a runner could finish a requested file but return only a machine-local path. Remote native runs had no bound file publication tool. **Expected behavior** The replacement reads and preserves existing work, receives saved messages once, and keeps each message's author. The former agent stays stopped. A requested file has a working attachment or accessible work product before Done. Text-only tasks do not require attachments. **Steps to reproduce** 1. Ask Codex to save three newsletter names and then wait. 2. Queue an instruction to keep those names and expand the draft. 3. Use Interrupt and assign to select Claude. 4. Verify the original names survive, the result has a working download, and only the source and replacement runs exist. 5. Ask either provider for a Markdown checklist and open the file from its completed response. ## What Changed - Carry the exact same-task interrupted run's summary, semantic receipts, and history into handoff context. Tell the replacement to inspect existing files before editing. - Adopt saved ordinary task comments into the successor's receipt under the task lock. Preserve authors and separate mention, chat, and interaction contracts. - Prevent a former-assignee comment wake from reopening a completed task or starting a stale execution. - Reject workspace-only, fabricated, and cross-task file completion references with actionable runner feedback. New file output also needs a matching current-run publication receipt and asset filename/size/hash, or an accessible work product registered by the current run. Prior output can remain context alongside a current file, or be verified and re-registered internally. Authorized chat attachment reuse retains its verified current-run clone receipt; older receipt shapes require an intact matching source. - Bind remote file reads to the active environment runner and reuse the existing attachment and work-product publication path. - Enforce workspace confinement, regular single-link files, stable identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the native session fingerprint for the updated tool contract. - Contain rejected remote signals and protocol-failure cleanup, including logging failures. Preserve the original cleanup rejection for its owner; a rejected operation never supplies stop acknowledgement or cleanup proof. - Allow exactly one maximum-size base64 file through the native SSH command adapter, preserving a finite output cap. - Document handoff and accessible file completion rules. ## Verification - Each observed bug has a failing regression before its fix. Final post-rebase integration passed 732 tests across 13 files before the final receipt and signal guards; final affected results are below. - Publication provenance and compatibility: 8 provenance regressions and 2 compatibility regressions failed before their fixes; the final four affected suites pass 53 tests, including mixed old/new references and real authorized chat reuse. Controls cover old attachments and work products, filename/size/hash/origin mismatch, missing/wrong receipts, current-run publication, same-run durable proof, internally re-registering preserved bytes, and no-new-file follow-ups. - Remote signal rejection: the real Node subprocess previously exited 1 when the production launcher signalled a deleted sandbox. It now stays alive for both a failed signal and failed logging; all 349 executor tests pass. The failed signal still provides no termination proof. - Remote file reader and SSH command boundary: 33 tests passed, including real Linux descriptor reads and the actual SSH adapter subprocess output cap (network executable replaced by a deterministic fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap fails. - Live local Claude and Codex Stop journeys preserve the saved file, deliver queued instructions once, and reach Done with two total runs. The handoff journey preserves the original names and download with exactly two runs. Both local providers deliver exact checklist files without a completion confirmation. - Live combined Daytona verification passed: the original failed task's Retry reused its sandbox; a selected Git subfolder produced an exact downloadable file; warm and deliberately resumed Claude runs took about 33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134 seconds of contention/backoff. Both cloud downloads retained exact bytes after the two owned sandboxes were deleted. - The sandbox-deletion retest identified a separate ignored promise in protocol-failure cleanup. Two real Node subprocess regressions failed under fatal unhandled-rejection policy before the fix; all 36 protocol, lifecycle, and integrity tests now pass. The original close promise still rejects to its owning runtime. The final live retest passed: a normal Claude Daytona task completed in 132.352 seconds, then its sandbox was deleted. Thirteen samples over 361 seconds confirmed the same controller stayed healthy, the task stayed Done with unchanged run IDs, and its attachment retained exact bytes. The post-deletion browser download passed with zero page errors; all five owned sandboxes are confirmed absent. - Full repository typecheck and build passed on final commit `de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758) passed: 32 successful checks and two conditional skips. The earlier mixed-source full local test invocation was deliberately stopped before rebase, so no pristine green full local aggregate is claimed. Its known failures passed in later affected suites. ## Risks - Handoff may adopt only ordinary comments from its validated former owner. Other delivery contracts must remain independent. - File verification fails closed if a remote file changes during reading. The runner must retry publication or explain a blocker. - The updated session fingerprint starts a fresh provider process where needed to install the new tool contract. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI `gpt-6-astra` through Codex, with reasoning, code execution, browser testing, and tool use. The context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24007a9804 |
fix: promote review tasks when continuations start (#13318)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service starts and tracks agent work on issues. > - An issue can be in review before a user comment starts more agent work. > - The previous checkout rule did not claim an issue from review for this continuation. > - The issue could stay in review while an agent actively worked on it. > - This pull request lets a resolved interaction continuation claim an issue from review. > - The existing checkout update then sets the issue to in progress. > - The benefit is that issue status now shows active agent work correctly. ## Linked Issues or Issue Description **What happened?** An issue stayed in review when a new agent continuation started work on it. **Expected behavior** The issue must move to in progress when an agent starts work. An idle issue must stay in review. **Steps to reproduce** 1. Put an assigned issue in review. 2. Resolve an interaction that starts a continuation. 3. Start the heartbeat run. 4. Observe that the issue remains in review while the run works. **Paperclip version or commit** The problem reproduces on the master branch before this change. ## What Changed - Allow a resolved interaction continuation to claim an assigned issue from review. - Keep idle review issues unchanged. - Add regression tests for both behaviors. ## Verification - Run `node_modules/.bin/vitest run server/src/__tests__/heartbeat-auto-checkout.test.ts`. - Confirm that all four tests pass. ## Risks - Low risk. The change only expands checkout eligibility for resolved interaction continuations. - The existing guarded checkout update still controls ownership and the status update. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5.5. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f12b647ae8 |
fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can collect more messages while its agent works. > - Legacy runners must stop the active process before they can receive those messages. > - The old Interrupt action cancelled the run but could leave the queue idle and hidden. > - Codex could also classify a cancelled run as successful or start a fresh process after cancellation. > - This pull request joins cancellation, preserves the provider session, and dispatches the current queue after cleanup. > - The benefit is reliable interruption with the saved message order, edits, and deletions. ## Linked Issues or Issue Description **What happened?** Interrupt could strand a legacy message queue. The UI could hide pending messages after the run stopped. A Codex signal exit could race the cancellation write. A stale session warning could also trigger a fresh process after an interrupted resume. **Expected behavior** Interrupt stops the active turn and sends the remaining messages once, in their saved order. Deleted messages stay deleted. An interrupted Codex turn keeps its session and does not restart itself. **Steps to reproduce** 1. Assign a task to a legacy Codex agent that runs a long command. 2. Queue three messages. Edit one, discard another, and move the last message first. 3. Click Interrupt in the queue. 4. Repeat the interruption while the resumed session runs another command. Related work: Refs #13160, which moves native queue steering into the wake-queue module. This change fixes legacy interruption and keeps native steering unchanged. ## What Changed - Add a revision-checked, company-scoped endpoint for legacy queue interruption. - Promote only the requested queue after the provider stops and releases its lease. Retry its persisted interrupt intent from the scheduler after a promotion error or server restart. - Keep pending legacy queues visible after a run stops. Use server state for the interrupt result. - Serialize owned process cancellation before classifying the adapter result. Preserve late session and log metadata. Acknowledge cancellation only when an actual process or process group was owned; scheduler placeholders retain their normal release policy. - Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the session has started. - Add cancellation race, multi-actor queue order, durable retry, resume fallback, and stale request regression tests. Document the behavior. ## Verification - Real browser tests passed with legacy Codex CLI and ACP engines, using Codex 0.153.4 and gpt-5.6-sol. - All three automated ACP browser scenarios passed locally: immediate Interrupt delivery, no replay of an unfinished write, and pause requiring Resume. Updated the old test expectation that required a separate “go” after Interrupt. - Browser tests covered queued edits, deletion, reordering, deleting the final message, and repeated interruption. - Two consecutive CLI interrupts kept one provider session. Both stopped processes exited. The final message arrived once. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed. - All 346 post-review scheduling, recovery, queue-route, archived-company, worktree-suppression, and stale-queue regression tests passed. - All 318 process-recovery and durable-chat tests passed after the final cancellation guard. - Codex adapter, queue UI, issue-page, and OpenAPI contract tests passed. - `pnpm build` passed. - Full local suite coverage completed with `PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI shards: 618 general server suites, all 145 serialized server suites, and all workspace groups. Every failing suite passed a targeted rerun after the fixes, rebuilding the native test fixture, correcting macOS temporary-path setup, or retrying setup/timing failures. Existing skips remain. - The original monolithic run reported failures before the final fixes; its failed suites were rerun rather than rerunning all 618 suites again. The final process-recovery/durable-chat regression run passed all 318 tests. - All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: [run 34654820774, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2), including typecheck, build, all test shards, E2E, and canary. The signoff and Cursor sandbox tests each hit a timeout in the initial attempt; both suites passed locally, and both failed shards passed their single CI rerun. All three corrected ACP browser scenarios passed in CI. - Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no open review threads. ## Risks Cancellation order affects local adapters. The tests cover signal exits, graceful exits, adapter exceptions, termination errors, and cancellation write errors. Embedded adapters keep their cancellation controls. Ordinary run cancellation and task pause keep their distinct queue policies. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code execution. The exact serving model ID and context-window size are not exposed in this session. The live test runner used OpenAI gpt-5.6-sol. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1c4bcff2b1 |
fix(test): await issue lock before retry race assertions (#13273)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Retry decisions must respect changes to an issue owner. > - Database tests verify this with two concurrent transactions. > - The test started the competing operation before its fixture held the lock. > - That race can fail a correct source verification run and block deployment. > - This PR waits for lock acquisition before starting the competing operation. ## Linked Issues or Issue Description Refs #13257 for the deployment verification work that exposed this test race. No duplicate fix was found. **What happened?** [Cloud verification job 103424158921](https://github.com/paperclipai/paperclip/actions/runs/34648268409/job/103424158921) failed because the test did not observe a concurrent issue-row lock waiter. The fixture and retry operation both started without an ordering guarantee. **Expected behavior** The fixture must hold the issue lock before the competing retry operation starts. The test must still prove that the retry waits for the lock and observes the reassignment. **Steps to reproduce** 1. Add a temporary 50 ms delay before the fixture acquires the issue lock. 2. Run the promoteOrCancelDueRetry issue-lock test. 3. The original fixture fails with the same missing-waiter error as CI. 4. The synchronized fixture passes with that delay. The delay is not part of this PR. ## What Changed - Separate fixture readiness from its transaction completion promise. - Wait for readiness at both callers before starting the retry decision. - Propagate transaction failure during setup through Promise.race. ## Verification - The original fixture fails under the temporary delayed-lock probe. The fixed fixture passes the same probe. - All 31 tests in server/src/modules/run-dispatch/adapters/postgres.test.ts pass after removing the probe. - The real concurrent waiter, lock-order, and reassignment assertions remain intact. No timeout was increased. - Full local typecheck and build pass (221s and 50s). The full local test command stopped in its server phase after 10,610 passes, 65 skips, and 14 failures: 13 existing macOS skill-cache rename/permission failures and one unchanged Telegram test assertion. The Telegram test passes in a focused rerun. Later local phases did not run after that failure. Current-head Greptile is 5/5 with no findings. All 32 current-head checks pass, including full Linux typecheck, build, native verification, server suites, browser suites, and canary packaging. The unchanged GitHub browser mock assertion passed its single failed-shard retry. - git diff --check passes. ## Risks - Only test synchronization changes. Production database behavior is unchanged. - Awaiting the transaction itself during setup would deadlock the test. Returning the completion promise inside an object avoids that problem. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — all 31 database adapter tests pass; full local verification limits are disclosed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb640ec129 |
fix(execution): keep blocked wakes waiting without repeated runs (#13236)
## Thinking Path > - Paperclip manages AI agents and their work. > - Wake admission decides when a task can create an execution run. > - Recovery can prohibit replay while the previous execution needs review. > - Dependency reconciliation kept creating runs before dispatch rejected that same hold. > - Each rejected run added another startup notice without doing useful work. > - This change checks the hold during admission and records repeated automatic waits once. > - Tasks keep their messages and can resume when the current gates permit execution. ## Linked Issues or Issue Description Related changes: Refs #13173 (stale completed-task continuations). Refs #12651 (dependency waits during recovery). **What happened?** A blocked task with completed dependencies can remain under a durable execution reconciliation hold. Each scheduler pass created a queued run. Dispatch then cancelled it before the adapter started. The skipped wake did not satisfy dependency wake deduplication, so this repeated and filled the conversation with “Couldn't start” notices. **Expected behavior** A known execution hold creates a waiting diagnostic without a run. Repeated automatic observations share that diagnostic. Clearing the hold permits a new wake only after the other gates pass. New comments remain available for the next eligible execution. **Steps to reproduce** 1. Assign a blocked task with a completed blocker. 2. Give the task an active reconciliation action, or a resolved action whose automatic recovery evidence still prohibits replay. 3. Run dependency reconciliation repeatedly. 4. Observe repeated cancelled pre-start runs on the base branch. This branch creates no runs while held and admits work after the effective hold clears. ## What Changed - Check effective execution holds under the issue admission lock before inserting runs. Keep the final dispatch check for races. - Share automatic wait diagnostics across producers, wake keys, and service restarts. Apply the helper to reconciliation, dependencies, pause holds, availability, budgets, and disabled heartbeats. - Preserve ordinary comment and interaction receipts during execution holds. Prevent release from draining them while replay is blocked. Keep external-chat receipt authorization intact. - Group empty pre-start reconciliation cancellations into a neutral waiting notice. Keep started runs and the full run history. - Document the waiting contract and add database-backed, UI, and browser regressions. - Stabilize two existing verification tests: allow the asynchronous chat lease transition a bounded five-second wait, and accept either legitimate damaged-session refusal while retaining exact archive-evidence assertions. ## Verification Passed targeted tests: - `pnpm exec vitest run server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts server/src/modules/wake-queue/adapters/postgres.test.ts` — 28 tests. - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx` — covered in the initial combined test run; UI suite passed. - Run-dispatch adapter tests passed in the combined gate regression run. - `pnpm exec vitest run server/src/__tests__/durable-chat-wakeup.test.ts` — 41 tests, including held receipt replay, promotion, and revoked access. - `PAPERCLIP_E2E_PORT=3294 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts` — all 3 browser scenarios pass. Repeated held messages create no additional runs or provider prompts and do not replay writes. - `pnpm check:token-gates` - `pnpm check:module-boundaries` - `git diff --check` `pnpm -r typecheck` and `pnpm build` pass. The full local `pnpm test:run` invocation did not finish green: it encountered exhausted local PostgreSQL shared-memory slots, a missing fresh-worktree runner test binary, and tests loaded across in-flight edits. The affected chat/database suites passed on rerun (81 tests), and the targeted lifecycle/recovery verification passed (3 tests). After building the runner test binary, the full native session suite also passed (37 tests). Final-head [CI run 34621288475](https://github.com/paperclipai/paperclip/actions/runs/34621288475) passed on `8659618b0ed2b98df002a28f4c1bd97321b0db04`, including all server/workspace test shards, all three browser shards, runner verification, typecheck, build, release dry run, and the aggregate verification gates. All 31 reported checks passed; the two conditional Storybook checks were skipped as intended. Greptile reviewed that exact commit at 5/5 with no unresolved review threads. ## Risks The wait record is diagnostic only. It must never count as a delivered wake or bypass a current gate. Tests cover repeated and concurrent admission, resolved no-replay evidence, a remaining dependency after hold clearance, deferred comments, and release gating. Explicit user requests and authorized chat receipts do not share automatic diagnostics. No migration or historical data deletion is required. Existing provider retry budgets remain unchanged. ## Model Used OpenAI GPT-6 through Codex. The runtime does not expose the exact hosted snapshot ID or context-window size. Used repository inspection, code editing, command execution, tests, and review tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
86c2e0ac4a |
feat(server): log an activity row for each queued-comment queue mutation (#13159)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server records actions that change issues and their queued comments > - The queued-comment edit, reorder, and discard routes changed queue state without activity rows > - Operators could not inspect these queue mutations in the activity feed > - This pull request adds one identifier-only activity row for each successful queue mutation > - The benefit is a durable audit trail with no comment text in the activity log ## Linked Issues or Issue Description **What existing behavior does this improve?** The queued-comment edit, reorder, and discard routes now record their successful mutations in the activity feed. **Subsystem affected** Cross-cutting (server and ui). **Current behavior** The three queue mutation routes change queued comments but do not write an activity row. The activity feed has no label for these actions. **Proposed behavior** Each successful route writes one activity row with the actor fields, entity fields, queue identifiers, and queue revision. The discard row also includes the cancelled run identifier. The activity feed shows a label for each action. **Reason and benefit** Operators need a durable record of queue changes. Identifier-only details support audit and troubleshooting without storing comment text. **Breaking changes** None. The routes keep their existing response and authorization behavior. **Additional context** Each mutation writes its activity row on the same locked transaction that applies the mutation, so the two commit or roll back together. The route publishes the live activity event only after that transaction commits. The separate comment-cancel route opts out of this write and keeps its existing single activity row. ## What Changed - Add activity rows for queued-comment edit, reorder, and discard mutations. - Include queue identifiers, revisions, ordered comment identifiers, and cancelled run identifiers as applicable. - Add activity-feed labels for the three new actions. - Add route and activity-format tests for the new behavior. - Write each activity row on the same transaction as the mutation it records, through a new port method that the adapter implements. - Keep the comment-delete route opted out of that write, so a cancellation does not log two rows. ## Verification - [x] `npx vitest run server/src/__tests__/issue-queued-comments-routes.test.ts` passes. - [x] `npx vitest run ui/src/lib/activity-format.test.ts` passes. - [x] `pnpm --filter @paperclipai/server typecheck` exits 0. - [x] `pnpm --filter @paperclipai/ui typecheck` exits 0. - [x] `node scripts/check-module-boundaries.mjs` passes. - [x] The full CI suite is green. ## Risks Low risk. The change adds activity rows after successful mutations and does not change route responses, authorization, or stored comment text. ## Model Used OpenAI Codex, GPT-5 Codex. The model used repository inspection, Git operations, and command execution. The context window and reasoning mode are not exposed by this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0d8bbf7cf4 |
refactor(server): move the queued-comment queue mutations into the wake-queue module (#13145)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server coordinates issue execution and agent wake events > - Queued comment mutations belong to the wake queue that owns their state > - Route-local database writes split queue rules across two layers > - This pull request moves those mutations into the wake-queue module and keeps route authorization and response mapping > - The benefit is one transaction boundary with company-scoped writes and a shared checked response contract ## Linked Issues or Issue Description **What existing behavior does this improve?** The queued-comment edit, reorder, and discard endpoints write queue state directly from the route layer. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The route layer owns database transactions, locks, queue writes, and wake-row writes for queued comments. **Proposed behavior** The wake-queue module owns these operations. The routes keep authorization, input checks, error mapping, and response mapping. **Reason and benefit** The module gives all queued-comment callers one transaction boundary and applies company predicates to every adapter read and write. **Breaking changes** None. The endpoints keep their existing paths and response behavior. ## What Changed - Move queued-comment edit, reorder, and discard operations into the wake-queue module. - Add company predicates to seven queue writes. - Use the shared queue contract type for mutation responses. - Add module tests and route tests for the moved operations. ## Verification - `server/src/modules/wake-queue`: 128 tests pass across 6 files. - `server/src/__tests__/issue-queued-comments-routes.test.ts`: 19 tests pass. - The server TypeScript check reports the same 141 pre-existing errors before and after this change. - GitHub Actions must pass the required pull-request checks. ## Risks The change moves transaction and lock ownership across module boundaries. The new adapter, use-case, and route tests cover the moved behavior. No database schema changes occur. ## Model Used OpenAI Codex, GPT-5, current agent runtime, tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae0c1fbbd5 |
refactor(server): move the admission half of the deferred wake state machine into the wake-queue module (#13136)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The heartbeat service admits wake requests while an issue has an active execution run. > - That admission branch mixes wake policy, database reads, and database writes in one service. > - This structure makes the wake-queue boundary hard to test and extend. > - This pull request moves the admission policy and its database adapter into the wake-queue module. > - The result keeps heartbeat orchestration small and makes the admission behavior testable in isolation. ## Linked Issues or Issue Description **What existing behavior does this improve?** The heartbeat service now delegates deferred wake admission to the wake-queue module. The module keeps the existing merge, defer, and ordinary-wake outcomes. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The heartbeat service contains a 146-line branch that reads wake state, chooses an outcome, and writes the result. **Proposed behavior** The wake-queue module owns the pure admission decision and the adapter reads and writes. The heartbeat service calls one module method. **Reason and benefit** This boundary reduces service coupling and lets module tests cover the admission policy. The change keeps the existing reason strings and outcomes. **Breaking changes** None. The change preserves the current behavior and public API. **Additional context** This pull request follows [PR #13132](https://github.com/paperclipai/paperclip/pull/13132), which merged the first slice of this refactor. I searched GitHub for duplicate and related pull requests before opening this pull request. ## What Changed - Move deferred wake admission policy into `server/src/modules/wake-queue`. - Add module ports and a PostgreSQL adapter for the admission reads and writes. - Keep the existing wake outcomes and stored reason strings. - Extend the module boundary check to reject service imports from the application layer. - Add unit and adapter tests for the moved behavior. ## Verification - `node --test scripts/check-module-boundaries.test.mjs` passes. - The `server/src/modules/wake-queue` suite passes 58 tests. - The eight pinned heartbeat and queued-comment tests remain unchanged and require CI verification. - Every continuous-integration check must reach a terminal green state before merge. ## Risks The refactor changes the location of wake admission logic. A missed adapter condition could change deferred wake behavior. The tests cover the policy outcomes and the adapter writes. The residual tenant-scope risk remains documented in the review record. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. The implementation author ran the tests and prepared the commit set. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0bff1d5cb6 |
refactor(server): simplify the wake-queue module (#13139)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server uses a deferred wake queue to release work at the correct time > - The wake queue had repeated decisions, helpers, queries, and recovery data > - This repetition made the module harder to read and left a drain invariant implicit > - This pull request moves pure decisions into the policy layer and simplifies the queue flow > - The benefit is a smaller, clearer module with the same behavior and direct test coverage ## Linked Issues or Issue Description Refs #13136 **Problem:** The deferred wake queue carried repeated logic across the application and database layers. **Expected behavior:** The queue keeps the same wake, release, recovery, and escalation behavior after the refactor. **Solution:** Move pure decisions into the policy layer, share repeated data and predicates, and state the drain invariant in the application layer. ## What Changed - Remove the unused database port method and carry the blocked release notice kind as a typed field. - Move four pre-drain decisions into pure policy functions with table-driven tests. - Resolve the responsible user once in the application layer. - Use the canonical helper for agent invokability checks. - Share string helpers and run predicates across the module. - Split the deferred-wake decision flow and share recovery facts with a discriminator. - Bound the drain loop and throw when it processes a wake identifier twice. - Rename queue ports to describe their behavior. - Share row-loading code between stranded-issue escalation adapters. - Restore the interaction-continuation integration test case. ## Verification - Run the three wake-queue module test files. They pass 61 cases locally. - Run `pnpm check:module-boundaries`. It passes locally. - Run the `server/` type-check and confirm that no error names the changed wake-queue module or `heartbeat.ts`. - Run continuous integration and confirm that `promotes an interaction continuation after removing a coalesced self-authored comment` passes. ## Risks The change refactors queue control flow and database adapter boundaries. The main risk is a behavior change in deferred wake release or recovery. The new policy tests and the restored integration case cover these paths. The local environment cannot run the integration test because the same dependency failure occurs on the base branch. ## Model Used OpenAI Codex, GPT-5, context window not exposed by the runtime, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6dd48cad43 |
refactor(server): move the release half of the deferred wake state machine into a wake-queue module (#13132)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The heartbeat service releases deferred issue wakes and promotes the next run > - The release logic sat in a large service function, which made its decisions and database effects hard to test > - A promotion race could reopen an issue without creating the promoted run > - Company checks did not protect every read and write, and some readers used a second transaction connection > - This pull request moves the release logic into a layered wake-queue module and closes these race and company-scope defects > - The benefit is clearer decisions, safer writes, and focused tests while existing callers keep the same entry point ## Linked Issues or Issue Description Refs: #10195 ## What Changed - Move the release half of deferred issue execution from `heartbeat.ts` into `server/src/modules/wake-queue/`. - Add pure policy decisions with table-driven tests. - Claim a wake before the reopen write and advance to the next wake when the claim fails. - Add company predicates to guarded reads and writes. - Pass the transaction-scoped issue snapshot to reader ports. - Keep `releaseIssueExecutionAndPromote` as the public wrapper. ## Verification - Run `pnpm check:module-boundaries`. - Run `pnpm exec tsc --noEmit` inside `server/` and compare the result with the known baseline. - Run the wake-queue unit and adapter tests in continuous integration. - Check the promotion-claim ordering, guarded writes, transaction-scoped reads, and cross-company outcomes. - Search the pull request diff for internal issue identifiers. ## Risks - The refactor changes the transaction path for deferred wake release. - A stale or lost wake claim now skips that wake and continues with the next queued wake. - The public wrapper keeps its name and signature, which limits caller risk. - Continuous integration must confirm the full server test suite and build. ## Model Used OpenAI Codex, GPT-5, exact runtime model version not exposed, context window not exposed, with shell, Git, and GitHub tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018ca5daaf |
fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task controls coordinate provider execution and queued user messages. > - Stop could finish before an embedded ACP provider stopped its tools. > - A later request could be held for reconciliation without a clear task response. > - A restored provider could also retain the stopped run's API credential. > - This pull request verifies provider termination and preserves safe session continuation. > - Operators can continue known-safe work and see why uncertain work cannot start. ## Linked Issues or Issue Description **What happened?** Stop could leave an embedded ACP provider running. A queued follow-up followed by “go” could fail before it reached the provider. Task chat could show a generic missing-response message. Even a restored session could use the previous run's credential and fail its task update. **Expected behavior** Stop waits for confirmed provider termination. A later explicit wake continues the same compatible session only when recorded actions have known outcomes. It carries pending comments and the current run's environment. Uncertain actions retain a visible reconciliation hold. Composer Stop preserves the existing pause rule: conversation can continue while paused, but task work requires Resume. **Steps to reproduce** 1. Start an embedded ACP task. 2. Send a second request while the provider is running. 3. Interrupt the run, then send “go”. Also test composer Stop followed by Resume work. 4. Check that the request is delivered once and that the provider can complete the task through the current run's API credential. 5. Repeat with an unfinished write. Confirm that the write stops and that further execution stays blocked with a visible reason. **Paperclip version or commit** Built from source on master at `3bc60dd8b` plus this branch. **Deployment mode** Local source build with an isolated embedded PostgreSQL instance. Refs #11183. Refs #12552. Those changes address recovery after operator cancellation. This change also covers embedded ACP termination, session proof, pending-comment delivery, and task feedback. ## What Changed - Propagate Stop into embedded ACP and wait for bounded adapter cleanup and provider exit. Retain the actual ChildProcess object for forced termination on all platforms; never signal a recycled numeric PID. - Preserve interrupted checkpoints only for acknowledged, local, persistent sessions with settled reads or no tools. Keep writes, incomplete actions, and forced termination blocked. - Restore the same compatible provider session with the current run's environment. Reject fresh-session fallback for an interrupted checkpoint. - Adopt pending comments on the next explicit wake. Stop alone does not dispatch them. - Share the execution-blocker rule across dispatch, Resume, and task detail. Show Stopped or Couldn't start with the recorded reason. Resolve the stopped agent for the run link, including reviewer runs. - Keep execution reconciliation holds intact when generic recovery sees queued comments or healthy child tasks. - Add process, service, component, and browser regression coverage. Fix disposable database cleanup and React test settling exposed by the full suite. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed all three `acp-stop-continuation.spec.ts` browser journeys. They use an actual ACP child process and require task completion through the agent API. - Passed 165 adapter execution, operator-stop, and child-process control tests, 17 queued-comment route tests, and 65 tests in the two adjusted UI suites. Earlier focused recovery, heartbeat, and task-control tests also passed. - Manually used the browser to queue a request, Stop, send “go” while paused, and Resume. The same session answered once and moved the task to Done with the current run's credential. - Manually interrupted an unfinished write. Its file size stayed fixed for five seconds. “Go” showed the reconciliation reason and did not start another provider prompt. - Separate live Claude ACP smoke checks confirmed that Stop ended a disposable local write and that a no-tool interruption could resume the exact provider session. The browser fixture does not call Drive or another external app. - Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and general server groups pass with targeted retries: two transient server failures passed together on retry, and two embedded-database startup failures passed after removing abandoned shared-memory segments from this task's completed browser fixtures. All 144 serialized server suites completed, with 2,189 tests passing after two transient HTTP socket failures passed on retry. - Passed all 135 heartbeat process/recovery tests, including a deterministic regression that failed before the recovery-sweep fix. - Passed 18 dispatch integration tests, including stopped-reviewer links, company boundaries, and malformed run IDs. - Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads. The security scan and all required CI gates pass for the same commit. ## Risks - Safe continuation depends on complete tool reporting and a restorable local provider session. Unknown outcomes remain blocked and require reconciliation. - Provider cleanup can take time. A timeout does not grant replay permission. - The change adds optional adapter context fields and an optional issue projection. It does not change the database schema or require a migration. - Test cleanup truncates company data only in a disposable test database. ## Model Used OpenAI GPT-6, running as Codex with repository tools, code execution, and browser interaction. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
35fdc0c66b |
fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e095b84dab |
feat(connections): connect services from native task feeds (#13058)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents use connections to reach external services. > - A fresh native task can have no service tools installed. > - The agent needs a way to discover services and ask the responsible person for access. > - This pull request brings the existing connection-intent flow into native task execution. > - The person can connect from the task, and the agent can continue with updated tools. ## Linked Issues or Issue Description **Subsystem affected** Native runner tool authority, connection intents, task interactions, and shared connection setup. **Problem or motivation** A task that needs an unconnected service cannot finish its work. Leaving the task to configure access also loses context. A resolved request must survive a restart and resume the correct agent once. **Proposed solution** Expose connection discovery and access requests as server-owned native tools. Render a durable task card and use the shared setup dialog. Persist outcome delivery and start a fresh provider session after access is ready. **Alternatives considered** Sending the person to the Connections page adds navigation and does not solve continuation. Polling for authorization consumes runs and can create duplicate requests. **Roadmap alignment** This extends the existing connection-intent runtime and setup experience. It reuses the shared access model and the native runner. Related: #12345, #12347. The service-slug fix in #12906 is related but separate. Companion evaluation PR: https://github.com/paperclipai/paperclip-evals/pull/21. ## What Changed - Expose `connections_search` and `connection_request` with server-bound company, task, agent, and responsible user. Preserve the legacy entry points. - Discover catalog services and authorized custom connections. Check installation, identity, health, and executable permissions before reporting ready. - Keep pending cards through ordinary messages. Reuse requests and retire stale ownership. Put Connect at the right of Not now. - Reuse the shared setup flow in a task dialog. Keep access additive and default to the requesting agent. Recover from cancelled or blocked OAuth windows with a new-tab fallback. - Persist outcome delivery with an idempotent wake key. Resume in a fresh session and recheck ownership before dispatch. - Add native browser fixtures, offline Storybook states, server contracts, and evaluation fixtures. Update guidance and documentation. ## Verification - `pnpm build`: passed after replaying the change on current master. - `pnpm -r typecheck`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed. - New continuation-policy regression cases: 16 passed. - Docker-backed PostgreSQL regressions passed for requester-only OAuth access, assignment-only expiry, terminal expiry, and credential-free setup metadata. - Shared setup and task-card UI tests: 121 passed, including configured MCP reconnect URL recovery and preserving user edits across refetch. - Storybook browser checks: all 119 passed on the latest reconnect fix. - `pnpm test:run`: 4,734 tests passed in the first server group, but embedded PostgreSQL startup failures and resulting cleanup errors prevented a complete local pass. All Linux CI lanes passed on the latest reviewed commit. One external-object route test returned an unexplained 500 on the first run; it passed twice locally and the failed shard passed on retry without code changes. - Earlier feature-checkout evidence: three deterministic native browser journeys passed, including restart delivery and an actual fixture tool result. Legacy scripted coverage also passed. All 59 added stories were inspected in light and dark themes. - Live Notion testing recorded successful provider reads. The manual test used a local-trusted instance. It does not prove authenticated/cloud deployment or every provider journey. - Native browser rerun reached the embedded PostgreSQL startup limit before bootstrap, so the latest checkout’s full native browser journey remains unverified. Both OAuth page/task regression cases passed against isolated Docker-backed PostgreSQL 17. They verify no premature task access, requester-only completion, additive retries, and reconnect preservation. - Applied both new migrations twice to isolated PostgreSQL 17. Foreign keys remained intact, duplicate active delivery keys were rejected, and failed delivery records did not block retries. Reviewer path: start a fresh test drive, enable the native runner, use an agent that can perform work directly, and ask it to summarize a Notion page. Connect from the card, then verify the resumed provider call and source-linked answer. The default test-drive CEO is instructed to delegate, so it can introduce an unrelated hiring step. ## Risks - Two additive migrations create durable deliveries and a partial unique wake index. They are idempotent. The wake index can require a maintenance window on large tables because migrations run in a transaction. - OAuth and continuation cross asynchronous boundaries. Tests cover ownership changes, retries, additive access, and restart delivery; live provider behavior still varies. - The latest requester-scope fix has not yet been exercised through live OAuth. GitHub, API-key, authenticated-user, and all recovery journeys are not claimed as verified. ## Model Used OpenAI GPT-6-based Codex assisted with implementation, tests, and review using tools and code execution. The runtime does not expose the exact model version, context window, or reasoning setting. Live evaluation used `gpt-5.6-luna`; manual native testing used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used and disclosed unavailable runtime details - [x] I have checked ROADMAP.md and confirmed this extends existing connection work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run all required tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f65991a5f1 |
refactor(server): move scheduled-retry and queued-run dispatch into a run-dispatch module (#12920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The server heartbeat service dispatches scheduled retries and queued runs. > - The service kept policy decisions and database writes in one large file. > - This layout made policy branches harder to test and transaction boundaries harder to inspect. > - This pull request moves the policy rules and database transactions into a run-dispatch module. > - The benefit is a smaller service, pure policy tests, and clear transaction ownership. ## Linked Issues or Issue Description **What existing behavior does this improve?** The server heartbeat service promotes scheduled retries and cancels stale queued runs. **Subsystem affected** server/ — REST API and orchestration services. **Current behavior** The heartbeat service contains the policy rules and the database writes for these dispatch paths. **Proposed behavior** A run-dispatch module owns pure policy functions and semantic database transactions. The public service contracts stay unchanged. **Reason and benefit** The new layout separates branch rules from database effects. It makes each policy branch easier to test and keeps each operation’s row writes in one transaction. **Breaking changes** None. The public service contracts stay unchanged. ## What Changed - Move scheduled-retry promotion and queued-run staleness rules into pure functions. - Add table-driven unit tests for each policy branch. - Move promotion and cancellation writes into semantic transactions. - Keep row locking, company isolation, and post-commit effects unchanged. ## Verification - `node scripts/check-module-boundaries.mjs` passes. - `tsc --noEmit` from `server/` reports no errors. - The focused server test command passes 228 tests in six files. - Full pull request CI passes. - Greptile reports 5/5, and all review threads are resolved. ## Risks The main risk concerns changed transaction boundaries in scheduled-retry promotion and queued-run cancellation. The focused tests retain coverage for locking, transactionality, company isolation, and transport contracts. The public service contracts do not change. ## Model Used Codex, GPT-5, with code execution and tool use. The context window is not provided by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ed5b7c5c8 |
refactor(server): extract the active-run output watchdog into a feature module (#12853)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server recovery service monitors active runs and applies watchdog decisions > - The watchdog rules and database operations lived in one large recovery service > - This structure made the rules harder to test and made company scoping harder to inspect > - This pull request moves the watchdog into domain, application, and adapter layers > - The benefit is a smaller recovery service, pure policy tests, and clear company-scoped ports ## Linked Issues or Issue Description **What existing behavior does this improve?** The active-run output watchdog that detects silence, suppression, and terminal evidence. **Current behavior** The recovery service contains the watchdog policy, use cases, database operations, and process control in one file. **Proposed behavior** A feature module separates pure policy, use cases and ports, and Postgres and process adapters. The recovery service delegates its public watchdog methods to this module. **Reason and benefit** The separation makes policy decisions easy to test. Company identifiers on every reader and writer port make tenant scope clear. Smaller service methods reduce change risk. **Breaking changes** None. The recovery service keeps its public methods and call sites. Related public watchdog work includes [#7043](https://github.com/paperclipai/paperclip/pull/7043) and [#7770](https://github.com/paperclipai/paperclip/pull/7770). ## What Changed - Add the `server/src/modules/active-run-watchdog/` feature module with domain, application, and adapter layers. - Move watchdog policy, use cases, Postgres access, and local process control into the module. - Keep the recovery service public methods and delegate them to the module. - Add 53 pure module test cases and retain 8 Postgres integration cases. - Add company scoping and transaction rollback coverage. ## Verification - Run `vitest run --config vitest.config.ts src/modules` and confirm 3 files and 53 cases pass. - Run `vitest run --config vitest.config.ts src/__tests__/heartbeat-active-run-output-watchdog.test.ts` and confirm 1 file and 8 cases pass. - Run the full server suite in pull request CI. - Compare the type-check result with a fresh baseline on the same checkout. ## Risks The main risk is a behavior change in recovery decisions during the move across layers. The pure policy tests cover the moved rules. The integration tests cover database behavior, company scope, and transaction rollback. Pull request CI runs the full server suite. ## Model Used OpenAI Codex, GPT-5, runtime-managed context window, tool use, code execution, and repository review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |