mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
codex/native-completion-current-master-candidate
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1778075155 |
fix(server): continue unfinished tasks after status replies (#14626)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their task outcome through a structured finish result. > - A Board comment can permit a passive wait for the next response. > - That exception accepted reports that also admitted blocking unfinished work. > - The task then stayed In Progress without a runner, and recovery treated the wait as healthy. > - This pull request rejects that contradiction and uses the existing bounded continuation path. > - Real wait conditions and protection against obsolete requests remain in force. ## Linked Issues or Issue Description **What happened?** An agent answered a status inquiry with `yielded` and `response_wake`. The same result listed blocking remaining work. The server accepted an indefinite wait without a question, approval, dependency, or pause. No further run was queued. **Expected behavior** Unfinished ordinary tasks must continue or have a recorded reason to wait. A status reply alone must not suspend the work. **Steps to reproduce** 1. Add a Board status inquiry to an unfinished assigned task. 2. Submit a successful native result with `yielded`, `response_wake`, and `remainingWork[].blocksCompletion: true`. 3. Leave the task without any real wait condition. 4. Observe that the old policy preserves In Progress with no continuation and suppresses recovery. **Paperclip version or commit** Reproduced in the native status policy at `da887ea3e`. The branch is based on current master. **Deployment mode** Authenticated server with the native Paperclip Runner. Related work: Refs #13338 (native response waits and recovery). Refs #12071 (separate legacy recovery and retry-state work). This change fixes the native unfinished-response-wait exception. ## What Changed - Reject contradictory finish reports while the provider can still correct them. - Route accepted unfinished response waits through the existing one-follow-up continuation budget. Repeated incomplete results create a visible recovery action. - Preserve questions, approvals, dependencies, pauses, conversation lifecycles, and superseded Board requests. - Recheck the current Board source in the decision transaction before queuing repair. - Let normal recovery reconsider old committed waits that report blocking work and still have a current source. - Add policy and database regression tests. Document the rule. ## Verification - `pnpm exec vitest run server/src/services/native-runtime/status-arbiter.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts --no-file-parallelism`: 346 tests passed, no skips. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `git diff --check origin/master...HEAD`: passed. - Final head `8f0824d9bb4d631f2347f8afffe5fff6d574cb69`: all CI checks green (54 passed, 2 intentionally skipped), including all general tests, serialized server suites, runner tests, eight E2E shards, and the canary dry run. - Greptile: 5/5 on the final head, with no unresolved review threads. - The broad local `pnpm test:run` encountered an unrelated timing failure in `workspace-runtime.test.ts` (waiting for managed process-tree listeners). That test passed on an isolated rerun. The duplicate broad run was stopped; the complete CI suite passed on the final commit. - The transaction-race regression reproduced an obsolete continuation before the fix. All 12 source-change cases now pass, covering passive waits, corrective continuations, and exhausted-repair decisions. ## Risks - Agents that previously parked unfinished ordinary tasks must now continue or record an actual wait condition. - Old contradictory waits become eligible for normal recovery. Existing ownership, budget, pause, and supersession checks still apply. - The rule uses the structured blocking-work flag. It does not infer omitted work from prose. - No schema, dependency, or UI change. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The session does not expose a more specific deployment identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f55759942b |
fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b193077582 |
fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections turn messages into governed agent runs and return replies to Slack. > - Remote runs must replace an incompatible sandbox runner with the controller's packaged binary. > - The fallback used a package-relative path that does not match the vendored server layout. > - Slack callback health also compared the internal proxy address with the public callback address. > - Master now contains the runner fallback fix; this pull request fixes callback health and extends missing-runner regression coverage. ## Linked Issues or Issue Description Refs #13677, #13680, #13691, and #13686. **What happened?** A packaged controller stopped a Slack-triggered remote run with `runner_remote_artifact_unavailable` when the sandbox runner needed replacement. Working Slack callbacks also showed a stale URL warning behind the Cloud gateway. **Expected behavior** The controller stages its packaged runner when needed. Callback health uses the observed public address and still detects real address changes. **Steps to reproduce** 1. Run a packaged server with a sandbox that has an older runner or no runner. 2. Send a Slack mention to an agent that uses that sandbox. 3. Route signed Slack callbacks through a claimed Cloud gateway that rewrites the upstream host. 4. Check the run and the Slack callback health panel. **Paperclip version or commit** Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The fix branch also includes #13691. **Deployment mode** Packaged server with a Cloud gateway and a Daytona sandbox. ## What Changed - Extend the controller-owned runner fallback tests with a missing sandbox binary case; retain the fix now merged in #13686. - Prefer dedicated gateway diagnostic headers that survive provider rewrites of standard forwarded headers. Use validated host hints only for callback-health evidence on claimed Cloud instances after provider acceptance. Preserve request bodies, routing, authentication, and configured callback URLs. - Cover current, stale, and missing sandbox runners, all Slack callback surfaces, rejected callbacks, malformed proxy hints, real host and port changes, and the existing self-hosted behavior. - Document the artifact lookup and callback-health boundaries. ## Verification - Native session executor and binary resolver suites: 379 tests passed after merging current master (`45c99a0d0`). - Targeted callback integration suite with disposable PostgreSQL: 4 tests passed before rebase. - Full `pnpm -r typecheck` and `pnpm build` passed after merging current master. The full local suite passed 12,668 tests; one suite failed to start its disposable PostgreSQL. Rerunning that suite alone passed all 31 tests. - The built server resolver selected the executable under `server/dist/vendor/paperclip-runner/bin/`. - Final callback regression: all 4 targeted integration tests pass, covering provider header rewrites, default ports, uppercase/trailing-dot hosts, and ignored self-hosted hints. - Live staging proof of the runner fix: a previously failed Slack thread recovered, a new mention received its requested response, and the account-connect command succeeded. Unsigned callbacks returned 401. - Live browser and Slack acceptance passed: generated callback URLs, account linking, a new mention, an interactive question and answer, and all three callback-health indicators. The connector was activated through the onboarding UI. - All 54 latest-head checks pass, two optional checks are skipped, and Greptile is 5/5 with no unresolved threads. ## Risks - Runner fallback must select a binary for the remote platform. This preserves explicit remote artifact overrides and the existing capability checks. - Proxy headers are not identity proof. They are used only for diagnostics on claimed Cloud instances after the provider accepts the request. Self-hosted instances ignore them. Wrong public hosts and ports still warn. - No database migration or new public API contract. ## Model Used OpenAI GPT-6 via Codex. Used reasoning, repository tools, code execution, and browser/native-app testing. Exact model variant and context window were not exposed by the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9335b7db10 |
fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance. Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale. Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
84fe89906d |
fix: complete native agent review handoffs (#13581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native execution uses durable runs, issue locks, wake requests, and typed tool authority > - A child can finish with a native agent review request while its original assignee stays responsible for the work > - The reviewer then needs a bounded execution path that can inspect the child, record one decision, and finish safely > - Before this change, assignee-only gates rejected the reviewer or left the parent waiting after the child review ended > - This pull request adds typed reviewer admission, scoped reviewer tools, durable wake and recovery handling, and parent continuation evidence > - The benefit is that native review handoffs complete without changing child ownership or granting broad mutation access ## Linked Issues or Issue Description Refs: #13314 Refs: #13574 **What happened?** A native child run could report `needs_review` for an agent reviewer. The reviewer wake then failed assignee and execution-lock checks. The child remained in review and the parent remained waiting. **Expected behavior** The named reviewer should receive one durable wake. The reviewer should inspect the child and resolve the exact review card. The child assignee should stay unchanged. The parent should receive the recorded review outcome after the child reaches its terminal state. **Steps to reproduce** 1. Run a native task with a different named agent reviewer. 2. Keep the child assigned to its original worker. 3. Let the worker finish with a native completion review request. 4. Start the durable reviewer wake. 5. Resolve the review and finish the reviewer run. 6. Observe the child and parent state. **Paperclip version or commit** Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification source: `eea171aae`. **Deployment mode** Built from source. **Installation method** Built from source (pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Access context** Both. **Database mode** Embedded PostgreSQL in the isolated live test fixtures. ## What Changed - Add server-validated native review assignment facts. - Admit only the exact company, issue, source run, decision, revision, addressee, and resolver policy. - Give reviewer runs a narrow set of Paperclip read and resolve tools. File and shell access follow the configured agent and environment policy, so reviewers can run tests. - Separate server-owned reviewer instructions from untrusted persisted review data. Escape the data boundary; retain server-enforced authorization. - Keep the child assignee unchanged. Atomically claim the reviewer run, wake request, and issue execution lock. A competing lock prevents provider startup. - Require the exact running reviewer session and current issue lock to resolve its assigned card. Reject missing, unrelated, or terminal reviewer runs. - Add durable reviewer wake, lock, stale-card, and abandoned-run recovery handling. - Prevent duplicate native wake dispatches during deferred admission and recovery. - Carry accepted or rejected child review outcomes into parent task context and continuation evidence. - Add focused server, runner, and native protocol coverage. - Preserve upstream continuation rules. Add child review decisions as separate evidence, while keeping real human answers in their own field. - Return actionable completion validation feedback to both providers. Permit a corrected completion after rejection. Keep strict terminal acknowledgment validation. - Apply exclusive shared-workspace locks to sandbox environments. Local and SSH folders can run concurrently, including when old settings request serialization. - Repair test timing, native event parsing, and the review artifact assertion. Allow a valid reject, correct, and accept review sequence. Check the accepted card against its reviewer run and decision. Keep polling within the existing deadline when review acceptance precedes the parent wake projection; report a specific missing-continuation error at timeout. - Apply the ACPX pending-call limit to reserved finish/block calls, with capacity-release and cancellation tests. ## Verification - `pnpm build`: passed on `eea171aae`. - `pnpm -r typecheck`: passed on `eea171aae`. - `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on `b31ad9ab8`; runner E2E typecheck also passed. - `pnpm check:token-gates`: passed. - Focused DB review, reviewer authority, and prompt-boundary checks: 31 tests passed. They cover invalid reviewer runs, competing locks, atomic admission, duplicate claims, and valid resolution. - Heartbeat, workspace, and recovery checks: 30 tests passed. - ACPX sidecar suite: 27 tests passed. Moving the capacity guard back below reserved handling makes both new regression cases fail. - Four focused live continuation checks passed on their first attempt at `f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and question tool guidance (12/12 each). These cases do not use the reviewer prompt path changed afterward. - Fresh Codex and Claude review-handoff checks passed all 29 native checks each on their first attempt at `eea171aae`. Both runs received the expected fixed prompt and completed cleanup. Only the six selected live flows were tested; no full paid provider catalog run. - The final commit only extracts the existing test-harness timeout diagnostic into a shared helper and adds positive and negative coverage. Removing the accepted-review guard makes two regression assertions fail; restoring it passes all six timeout tests. Production runtime code, prompts, deadlines, and grading criteria are unchanged by this final commit. - Deadline regressions: a valid continuation delayed 20 seconds succeeds within its 30-second unit-test deadline; an absent wake returns a specific candidate-failure diagnostic at that same deadline. Both assertions failed before the fix. Production E2E deadlines remain unchanged. - Historical native failures remain recorded: Docker availability failures; a valid reject/correct/accept sequence that the first-card grader misread; and a test that rejected the gap between accepted child review and parent wake projection. No failed result was regraded. The latest tests use a protected reference to the pinned Docker image and the unchanged artifact oracle and time limits. - Full repository verification runs in GitHub CI. Local verification uses the focused suites above, full build, and full typecheck. An unchanged Codex shutdown timing test failed once in CI, passed in isolation, and its full shard passed on the final commit without changes to that test or its causal code path. The original failure is retained in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no outstanding actionable findings. All review threads are resolved. All current-head CI gates passed, including the isolated native runner Docker build (55 successful checks; two skipped by the workflow). ## Risks - Reviewer admission depends on exact persisted decision and interaction bindings. A stale or changed card is rejected. - Paperclip control-plane tools are limited to inspection and review resolution. This is not a filesystem permission boundary; provider file and shell access retain the configured policy. - Deferred wake recovery changes dispatch receipt coalescing. A scheduler regression could delay a continuation if the receipt state is wrong. - Parent review outcomes are evidence for the model. They do not grant tool authority or change issue ownership. - This change does not address legacy lease-hold handoff behavior. > Roadmap review: native execution, review gates, and durable recovery are existing roadmap capabilities. This PR completes a narrow reliability path for those capabilities. ## Model Used OpenAI `gpt-6-astra` with reasoning, tool use, and code execution. OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and journal work. Context window size is not exposed by this session. Live test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the PR authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9021c6721 |
fix: require explicit native completion reviews (#13314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their outcome through paperclip_finish. > - The server previously turned incomplete reports into human approval requests. > - Those requests could block a later successful run, even when no person had requested review. > - This pull request creates review cards only for explicit attention requests and withdraws proven old fallback cards. > - Agents receive useful completion feedback, while explicit approval gates and task state protections remain in force. ## Linked Issues or Issue Description Related: #13266 removed reviews caused by policy upgrades. This change removes the separate completion fallback. **What happened?** An agent reported needs_review while waiting for checks without requesting a human decision. Paperclip created a generic Native completion review. A later successful report could not complete the task because that old card remained pending. **Expected behavior** Ordinary low-risk work completes after a valid done report, a successful run, and workspace finalization. Incomplete work stays with the agent. Explicit approval requests remain visible and must be resolved. **Steps to reproduce** 1. Complete a native run with needs_review and no attention requests. 2. Continue the task and submit a successful done report. 3. Observe that the old implementation leaves the task in review behind a generic confirmation card. ## What Changed - Require explicit attention requests to create native review cards. Route each request independently and preserve pending or declined decisions. - Withdraw only pending system cards with matching old decision, assessment, effect, contract, and prompt provenance. Preserve history and explicit or answered requests. - Reassess an affected current result without overwriting later task edits, runs, contracts, or workspace failures. - Return pending approval links and required actions through the completion tool. Reject contradictory done reports and empty review requests before accepting a result. - Allow one corrective continuation for incomplete results, then expose a recovery action. - Update status fixtures, database regressions, runner tests, and the completion contract documentation. ## Verification - `pnpm -r typecheck` passed after merging current master. - `pnpm build` passed after merging current master. - The combined branch passed 89 completion and Agent Chat tests. Other targeted tests passed: 170 external-chat and reconciliation tests; 50 runner-resume and control-plane tests; 13 arbiter tests; 7 chat delivery tests; 21 runner completion and runtime-context tests. - The full local test attempt exposed old review fixtures and a missing fake-provider binary. The fixtures are fixed and the helper is built. All affected suites pass in fresh reruns. The timing-sensitive Discord test also passed on rerun. - All latest-head CI checks passed, including build, typecheck, general and serialized tests, runner verification, browser tests, and canary dry run. Greptile is 5/5 with zero unresolved comments. ## Risks - Cleanup changes existing pending cards. It requires exact system provenance and only applies to low-risk agent-claim contracts. It does not delete history or dismiss explicit requests. - Status still commits after the turn and workspace finalization. Completion feedback reports current constraints and does not claim an early status commit. - Incomplete reports now request a bounded corrective run instead of an automatic approval. Repeated failures expose recovery. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and test tools. The exact deployment identifier and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a12bbd1824 |
fix: stop completion reviews caused by policy upgrades (#13266)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runtime records completion assessments and task decisions. > - An application update can change the assessment policy version. > - The previous code treated that version change as a reason for human review. > - An upgrade alone does not give the user a new decision to make. > - This change keeps existing decisions and withdraws obsolete upgrade review cards. ## Linked Issues or Issue Description **What happened?** A rules version change moved unfinished tasks into review and created a card that said, "Review the superseding native policy assessment." The task did not need new work or a human decision. **Expected behavior** New runs use the current rules. An upgrade leaves existing task decisions alone. The saved policy version remains available in the audit history. **Steps to reproduce** 1. Save a native run assessment and a task decision. 2. Change the native status policy version. 3. Run finalization reconciliation without new task evidence. 4. The previous code created a review card. The corrected code keeps the saved assessment and decision. Related context: #13038 changed the native policy version. Searches found no duplicate fix for upgrade-only completion reviews. ## What Changed - Remove policy-version mismatch as a reconciliation trigger. - Withdraw pending cards only when their source decision, effect ledger, creator, key, and prompt match the old upgrade-only review. - Restore the previous status only while that decision and status version remain current and no other review gate is pending. - Preserve answered cards, real review requests, later task changes, and historical assessments and decisions. - Record cleanup activity and retire any corresponding chat review actions. - Isolate cleanup failures so one old card cannot block other cleanup or normal finalization. - Update the status conformance fixture, regression tests, and architecture documentation. ## Verification - Passed: `pnpm exec vitest run server/src/__tests__/native-status-arbiter-corpus.test.ts` (23 tests, including the 53-fixture status corpus). - Passed: `pnpm -r typecheck`. - Passed: `git diff --check`. - Passed: `pnpm build`. - Passed: fresh Greptile review at 5/5 on `d24e5ecaf`, with no open review threads. - Passed: the 23 focused tests in the isolated full-suite environment. - Local `pnpm test:run`: the server group finished with 10,638 passed and one missing-fixture failure. Built the required `fake-codex-app-server` fixture and reran the entire affected native session-resume suite: 37/37 passed. The initial full command exited on that server-group failure, so remaining groups are covered by CI. - Passed: the workspace-runtime-exposure suite (25 tests, 3 platform skips). - Passed: all CI gates, including every test shard, browser tests, typecheck, and build. The server shard passed on one retry after an unrelated host-port conflict in the unchanged workspace-exposure tests. - Cleanup tests cover repeat runs, real reviews, answered cards, later statuses, newer decision identity, status changes back to review, other pending requests, run scope, cleanup failures, retry, and live publication failures. ## Risks - Cleanup changes stored task state. It checks the exact obsolete decision and status version under database locks before restoring status. - Pending approvals, interactions, and execution stages prevent restoration out of review. - The cleanup handles at most 100 matching cards per reconciliation pass. It does not rewrite old decisions or accept an agent's completion claim. - New evidence and explicit task changes still use the existing reconciliation paths. This change does not reevaluate old work merely because Paperclip was updated. ## Model Used OpenAI Codex, GPT-6. The exact deployment identifier and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0bdbf61564 |
feat(runner): add secure remote transport (#12639)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The lower stack PR adds authenticated remote execution targets and provider ingress. > - The Rust daemon currently accepts only loopback plaintext WebSocket connections. > - Remote Codex needs authenticated WSS dialing and provider-ingress listener mode. > - This pull request adds the bounded Rust transport contract. > - The benefit is a secure transport layer for the Codex remote vertical slice. ## Linked Issues or Issue Description Refs #12638. Refs #12616. Refs #12352. **Subsystem affected** Paperclip Runner Rust transport and remote runner networking. **Problem or motivation** The runner daemon cannot connect to a public control plane with TLS. It also cannot accept a provider preview connection on the run-bound ingress path. **Proposed solution** Add WSS with native trust roots and an optional private CA bundle. Add a fixed authenticated listener mode for provider ingress. Advertise the exact transport contract through build metadata. **Alternatives considered** Plaintext public WebSocket connections would weaken the transport boundary. A general listener would expose more network surface than the run-bound provider ingress requires. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing native runs. ## Stack - Lower merged PR: #12638. - This PR contains only its 13-file delta against `master`. - Later stack PRs add the task workspace and administrator UI. ## What Changed - Added WSS dialing with rustls and native certificate roots. - Added an optional bounded private CA bundle that augments native roots. - Kept plaintext WebSocket dialing restricted to loopback addresses. - Pinned resolved dial addresses for the process lifetime. - Added a fixed `0.0.0.0:43127` listener with an exact run-bound path. - Rejected listener queries, ambiguous paths, and WebSocket extensions. - Kept frame and message size bounds. - Added bounded reconnect grace and exponential jitter. - Retried bootstrap failures only before authentication proof transmission begins. - Kept post-proof failures fail-closed and bounded the welcome exchange at two seconds. - Added runnerd build metadata for the versioned transport contract. - Updated Rust dependencies and `Cargo.lock` only for TLS and certificate handling. - Did not add provider dispatch, Pi, AWS, `pnpm-lock.yaml`, migrations, or workflows. ## Verification - GitHub Actions will run Cargo formatting, Rust tests, repository tests, typecheck, build, security, and policy gates. - Rust tests cover URL validation, listener path validation, build metadata, durable recovery, and the existing Codex provider path. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check master...HEAD` passes. - The delta contains 13 files. ## Risks - TLS and listener changes affect the runner trust boundary. - Public plaintext transport remains rejected. - The listener uses one fixed port and one exact run-bound path. - PRP authentication remains required after the WebSocket upgrade. - The optional CA file uses the existing private-file checks and a 4 MiB limit. - This PR does not enable another provider or change direct adapters. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0a422fda52 |
feat(runner): add remote execution substrate (#12638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner gives native runs a durable and governed execution path. > - The current native path runs on the control-plane host. > - Remote environments need an authenticated execution-target contract. > - The contract must not change direct adapters or enable new runtimes by default. > - This pull request adds the remote execution substrate and Daytona ingress. > - The benefit is a bounded base for later remote runner transport work. ## Linked Issues or Issue Description Refs #12616. Refs #12352. **Subsystem affected** Cross-cutting. This change touches runner transport, server orchestration, plugin contracts, and shared settings. **Problem or motivation** Native execution cannot resolve an authenticated runner ingress through a remote environment. The server also lacks one provider-neutral contract for remote execution targets. **Proposed solution** Add a default-off runner preview ingress capability. Add transport-neutral runner connectivity. Add remote execution target and lifecycle handling. Add a Daytona ingress implementation with redacted credentials. **Alternatives considered** A provider-specific server path would duplicate orchestration and authorization. A public endpoint without an environment contract would weaken the trust boundary. **Roadmap alignment** This work supports the Cloud and Sandbox agents milestone. It also supports self-healing runs and governed tool access. ## What Changed - Added execution-target traits for local, SSH, and sandbox environments. - Added plugin RPC contracts for runner ingress endpoints. - Added authenticated Daytona preview ingress. - Added transport-neutral PRP outbound connections. - Added remote runner artifact verification and fail-closed provider selection. - Added bounded native session resume, cancellation, and lifecycle recovery. - Preserved Codex-only selection for fresh experimental runner starts. - Preserved all direct adapter execution and finalization paths. - Removed stale Pi provider-pack requirements that security review rejected. - Kept the rollout controls off by default. - Did not change pnpm-lock.yaml, Cargo, database migrations, or GitHub workflows. ## Verification - GitHub Actions will run the repository test, typecheck, build, security, and policy gates. - Focused tests cover ingress validation, redaction, execution targets, remote lifecycle, cancellation, resume, and legacy adapter selection. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check origin/master...HEAD` passes. - The diff contains 52 files. ## Risks - Remote execution crosses a trust boundary. - The implementation validates target capabilities, artifact digests, provider-pack pins, and connection metadata. - The feature remains default-off. - Fresh native selection remains Codex-only. - Existing direct adapters remain on the legacy path. - This PR does not yet make remote Codex runnable. The next PR adds the Rust WSS and TLS transport. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
25cf079ec5 |
feat(runner): add Codex-native application integration (#12591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package is useful only when the application can start, observe, and recover a native Codex run safely. > - Existing direct adapters must keep their current execution and finalization paths. > - The application boundary therefore needs additive persistence, authorization, coordination, and recovery behind an explicit experimental adapter. > - This pull request adds that Codex-only boundary without activating generalized providers, remote environments, or the later task/SDK surfaces. ## Linked Issues or Issue Description **Subsystem affected** Shared contracts, database persistence, adapter utilities, server native-runtime services, and the experimental Paperclip Runner adapter. **Problem or motivation** The already-landed runner package has a qualified Codex path, but the application needs durable native-run state, guarded runtime selection, authenticated coordination, tool security, finalization, and recovery before the experimental adapter can be exercised safely. **Proposed solution** Add a Codex-only `paperclip_runner` application path behind the existing default-off native-runner setting. Bind native state and coordination to company/run identity, preserve persisted-run recovery, and leave every direct adapter on its existing legacy execution path. **Alternatives considered** The earlier stack boundary introduced a generalized executor and remote-environment lifecycle here. That made this PR depend on implementations in higher PRs and changed reusable sandbox behavior globally. Those pieces are now deferred together to #12592. **Roadmap alignment** ROADMAP.md does not list a conflicting native-runner integration project. This change adds the application boundary for the existing Runner architecture. ## What Changed - Added native run/result/finalization/provider-trace persistence, shared validators, and idempotent migration/replay coverage. - Added guarded Codex-only runtime selection, authenticated PRP coordination, recovery, finalization, and interaction services. - Added run/company-bound tool-gateway authorization, credential redaction, SSRF protections, and replay-safe behavior. - Added the explicit `paperclip_runner` adapter behind the default-off rollout setting. - Preserved legacy answered-question wake projection and direct-adapter execution/finalization paths. - Hardened cancellation so only owned in-memory child processes are signaled; persisted recycled PIDs/process groups are never trusted. - Retained the narrow Claude ACPX isolated-context security follow-up discovered after #12590. - Deferred the generalized executor, provider ingress, remote lifecycle, SDK/lab/eval work, release-process changes, and lockfile. ## Verification - Changed-file delta against `master`: 133 files. - GitHub Actions is the authoritative verification environment for this PR. - Full CI, security, and Greptile review will run on this lowest unmerged stack PR. - Local tests/build/typecheck were not run because this checkout is resource constrained. - Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged. ## Risks - This touches central heartbeat and agent-route code, so legacy compatibility is the primary risk. - Runtime selection remains Codex-only and explicit; direct Codex, Claude, OpenCode, process, HTTP, and plugin adapters remain on their existing paths. - Fresh native starts fail closed while the rollout flag is off; persisted native records remain readable and recoverable. - Cancellation, company/run binding, tool calls, status decisions, and completion writes are guarded or replay-safe. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI and security gates are green - [ ] Greptile is 5/5 with no open actionable findings - [x] I will address all Greptile and reviewer comments before merge ## Stack - Position: 3 of 5 overall; lowest of 3 currently unmerged - Base: `master` - Previous: [#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified Claude ACPX runtime — merged - Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592), generalized Codex executor, task experience, and developer SDKs --------- Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
41bf5cafa1 |
docs(runner): define architecture and compatibility (#12084)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent execution currently uses direct adapters inside the server process. > - The proposed Paperclip Runner adds a separate process and a new protocol boundary. > - This boundary needs clear trust, recovery, rollout, and compatibility rules before code lands. > - Large runner changes are difficult to review as one pull request. > - This pull request defines the first small boundary for the runner series. > - The benefit is a stable design contract for later implementation pull requests. ## Linked Issues or Issue Description **Issue type** Missing documentation. **Where is the issue?** The repository does not have a concise architecture decision or compatibility contract for Paperclip Runner. **What's wrong?** The available runner design material is too large for normal review. It mixes architecture, implementation history, test evidence, and deferred work. Reviewers need a short statement of the process boundary, trust model, rollout behavior, and direct-adapter compatibility rules. **Suggested fix** Add one architecture decision record and one compatibility document. Keep implementation details and campaign evidence out of this pull request. Related public work: Refs #11041, #11297, #11634, #11639, #11640, and #11962. This pull request is the first small replacement in the new review series for #11962. ## What Changed - Added an architecture decision for the runner process, PRP v1 transport, semantic tools, durable recovery, and additive server integration. - Added a compatibility and rollout contract for the default-off adapter, existing direct adapters, persisted native runs, and the task page. - Defined the initial package and provider limits. The first production provider is Codex only. - Defined acceptance checks for later implementation pull requests. ## Verification - `pnpm install --frozen-lockfile` passed with Node 24.19.0 and pnpm 9.15.4. - `pnpm check:node-version` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `git diff --check origin/master...HEAD` passed. - The pull request changes 2 files. - `pnpm test:run` completed with 4,686 passing tests and 30 failures in unchanged master paths. The failures reproduce macOS path aliases, invalid generated port values, and local listener behavior. This documentation-only change does not touch those paths. Linux CI must pass before this pull request is ready. - All applicable GitHub Actions and security scans passed. The Storybook visual job skipped because this documentation-only change does not match its paths. - Greptile completed at 5/5 with no actionable comments. ## Risks Low implementation risk. This pull request changes documentation only. A later implementation can still diverge from the contract. Each later pull request must prove its behavior against these compatibility rules. I checked `ROADMAP.md`. This design supports the governed tool and control-plane direction. It does not add an overlapping user feature. ## Model Used OpenAI Codex with GPT-5 was used. The exact serving model ID and context size were not exposed. The model used high reasoning, repository tools, GitHub tools, and local code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |