mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
345
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b2e9e82f05 |
fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The execution service owns each run and saves messages sent while it runs. > - Interrupt must stop the current executor before it delivers those messages. > - Remote Grok commands did not register the host cancellation control. > - A cancelled task run could still write Done and prevent queue recovery. > - This pull request connects remote cancellation and revokes cancelled run writes. > - Saved input can use the existing queue admission rules after verified cleanup. ## Linked Issues or Issue Description **What happened?** Interrupting a queued message marked a remote Grok run cancelled before its sandbox stopped. The old run could still post a reply and mark the task Done. Its saved follow-up remained deferred behind execution recovery. **Expected behavior** Stop revokes run write authority and waits for verified termination. Saved messages remain durable and enter one successor through normal admission after cleanup. **Steps to reproduce** 1. Run a task with `grok_local` in a remote sandbox. 2. Send a follow-up and use Interrupt while the command runs. 3. Let the old command attempt a task status update after cancellation. 4. Observe the task disposition and the saved message queue. **Paperclip version or commit** The gap is present in master at `d3e0f0a238`. **Deployment mode** Authenticated server with a Daytona sandbox. Related work: #14028 and #14046 handle bounded continuation. #13291 covers infrastructure interruption and verified remote cleanup. #13332 addresses atomic recovery holds. This change handles direct Grok operator cancellation and stale task writes. ## What Changed - Register remote Grok cancellation before preparation. Keep command ownership until the host confirms sandbox termination. - Reuse the sandbox cancellation boundary for the direct CLI invocation. Reject fresh attempts after cancellation and preserve workspace restore failure evidence. - Reject writes from cancelled task JWTs and runs with a pending stop. Preserve diagnostic reads and existing conversation error codes. - Recheck run authority under a database lock before task updates and interaction responses commit. - Preserve authorized handoffs that stop their own run. Only the server-issued stop receipt for that request permits the final task update. - Add tests for hung commands, unverified stops, early cancellation, copy-back failures, late Done, late interaction responses, authorized handoffs, exact lease receipts, and one queue successor across concurrent restart sweeps. - Document the cancellation and write-authority contract. ## Verification - Targeted adapter, cancellation-boundary, authentication, queued-message, interaction-service, and activity-route tests passed. The expanded run passed 214 tests; one new test had an incomplete fixture. After correcting the fixture, all 8 selected follow-up cases passed. - `pnpm -r typecheck`: passed on `179c86caf1bf0d89914a503d46e24af7e4b8c557`. - `pnpm build`: passed on the same commit. - `pnpm test:run`: the general-server group completed with 13,521 passed, 99 skipped, and 18 failed tests. It then stopped, so the remaining local groups did not run. Five Slack, email, and wake-batching failures passed on focused reruns after correcting the local environment. The remaining 13 failures reproduce as `EACCES` on rename in unchanged skill-cache code on macOS. Two custom-image suite setup hooks also failed to start embedded PostgreSQL after the machine exhausted shared-memory slots; all 31 tests in that file passed on rerun after the local resource issue was resolved. CI covers all test groups. - CI: 53 checks passed and 2 were skipped on the latest commit, including the aggregate verification gate. The last server shard passed on its single rerun after a preview-server startup timeout. The affected file also passed locally with 28 passed and 3 skipped. - Greptile: 5/5 on the latest commit. Both review threads are resolved. - No live deployment or staging task mutation has been performed. ## Risks - Stopping the sandbox can prevent file copy-back. The result preserves workspace restore failure evidence; termination does not imply restored files. - If provider termination fails, the adapter keeps ownership of its outstanding command and does not acknowledge Stop. - The write restriction now applies to ordinary cancelled tasks. Reads remain allowed. Task and interaction checks add a shared run-row lock to agent mutations. An exact server-issued receipt permits the task request that stopped its own run to complete its handoff. - Existing terminal tasks are not reopened automatically. An operator must correct a historical late Done before its saved queue can continue. - No schema migration or UI change. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ca404b49a |
fix: record safe sandbox restore failure diagnostics (#14064)
Record a bounded diagnostic for failed workspace and staged-asset restores. Preserve the original error, retry policy, and archive safety checks. Never copy raw provider messages, credentials, paths, or asset names into the log. Nested failures log once; safe fields survive throwing property getters. Verified 129 focused restore/Claude tests, typecheck/build, and green full PR CI. Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd6caf51bb |
fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget. Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f3a214fe77 |
Fix relative symlinks in secondary sandbox repositories (#13953)
## Thinking Path > - Paperclip manages agents and their task workspaces. > - A task can use several independent Git repositories. > - Sandbox staging copies secondary repositories from temporary clones. > - The copy changed relative symlinks into absolute host paths. > - Those links broke skill discovery and made workspace restore fail. > - This change preserves link targets while keeping extraction checks intact. ## Linked Issues or Issue Description **What happened?** Staging a secondary repository rewrites a link such as `.claude/skills/demo -> ../../skills/demo` to an absolute path in a temporary Git clone. That clone is then removed. The link is broken in the sandbox, and Daytona refuses the outbound archive during workspace restore. An agent can finish its turn but still have its run fail during restore. **Expected behavior** Repository links keep their original targets after staging. Links within a repository remain usable, and changes return to the local checkout. Unsafe outbound archive links still fail before extraction. **Steps to reproduce** 1. Create a project with a primary repository and a secondary repository. 2. Commit a relative skill directory link in the secondary repository. 3. Stage the workspace for sandbox execution and inspect the copied link. 4. Restore that repository through Daytona. Before this fix, the link points at a removed host temporary directory and restore rejects it. **Paperclip version/commit** Reproduced on `0f8750627f11d855552abce9a837d7f3b67c9ddf` in the multi-repository sandbox path. Related: #13442 introduced multi-repository provisioning. #13882 adds native Grok but keeps legacy adapters; #12991 addresses Grok instruction isolation and leaves skill staging unchanged. Searches of open/closed PRs and open issues found no direct fix for this copy behavior. ## What Changed - Set `verbatimSymlinks: true` when copying secondary Git clones. This [Node option](https://nodejs.org/api/fs.html#fspromisescpsrc-dest-options) preserves the stored link target instead of resolving it against the temporary source. - Test directory, file, chained and dangling links after temporary-clone cleanup. Assert the copied Git checkout remains clean. - Cover skill-link reads, edits and restore in fresh, warm-adoption and durable-seed workspace modes. - Extend Daytona checks for valid relative directory links and rejected absolute targets. - Document the staging behavior. ## Verification - Four regression cases fail without the source fix: one clone test and three staging modes. - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 301 passed. - `pnpm -r typecheck` and `pnpm build`: passed. - `pnpm test:run` was started locally, then stopped after the full Linux CI suite passed. It did not finish locally; this is not a full local-suite pass. - Full PR CI: 53 checks passed, two conditional skips, on `07b30ff298baab327a4e60708ade899935688a3f`. Three server shards were interrupted by runner shutdowns; the unchanged mobile repository test timed out waiting for a disabled Save changes button. One same-commit failed-job rerun passed. Original attempts remain in [run 36036538369](https://github.com/paperclipai/paperclip/actions/runs/36036538369). - Greptile: 5/5 on the same head, no review threads or actionable findings. - Diff scanned for secrets and private identifiers; no matches. ## Risks Low risk: the production change is one copy option. It preserves symlinks instead of following or materializing their targets. Daytona extraction guards, workspace exclusions, authentication and database behavior do not change. This prevents corruption in newly staged snapshots. It does not rewrite an already corrupted warm workspace or durable seed; those need fresh staging from the source checkout. No live provider run or customer-task replay was performed. The tests use real Git, filesystem and tar operations with mocked provider transport. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
|
||
|
|
8326e33ada |
Fix oversized sandbox process launch payloads (#13793)
## Thinking Path > - Paperclip manages agent work and preserves context across retries. > - Sandbox ACP runs encode their command and environment into one launch value. > - A long continuation can make that encoded value exceed Linux's exec limit. > - The launch shell then exits before the agent can initialize. > - This PR transfers large command envelopes through a private temporary file. > - The agent receives the complete environment and can start normally. ## Linked Issues or Issue Description Refs #13777. **Bug description** Sandbox tasks with long retry context fail during ACP initialization with exit 127. The streamed bridge also drops the shell error that explains the failure. **Steps to reproduce** Launch the streamed sandbox process bridge on Linux with one valid 110,000-byte environment value. Its base64 command envelope exceeds the limit on one exec argument or environment string. Daytona reports `argument list too long: env`, then exit 127. **Expected behavior** The command envelope must not make a valid child environment too large to launch. A shell startup failure must retain its diagnostic in the run log. ## What Changed - Keep envelopes up to 64 KiB on the existing launch path. Upload larger envelopes in bounded chunks inside a mode-0700 session directory. Set the final payload file to mode 0600. - Read the file without following symlinks and delete it before spawning the child. Remove incomplete uploads on failure. Both streamed and polled bridges use the same envelope. - Preserve stderr when the launch shell fails before the wrapper emits a terminal event. Emit a fixed terminal error and shutdown acknowledgement if a payload cannot be read or parsed, without exposing its contents. - Cover large environments, file permissions, payload deletion, interrupted uploads, missing or malformed payloads, and startup diagnostics. Document the transfer and cleanup behavior. ## Verification - Both large-envelope regressions fail before the fix and pass after it. - Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass: 370 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - A gated live Daytona probe on the current sandbox image reproduced exit 127 with the old bridge. The fixed bridge launched the same command successfully. A second probe completed real Claude ACP initialization with a 110,000-byte context value. It did not run an agent task. All temporary sandboxes were deleted. - `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner step and stop because this machine has no `cargo` executable. - Full CI passes on `ea816cf58d`: [run 35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754). All 53 checks pass; two optional checks are skipped. The local full-suite run was stopped after equivalent CI suites passed; it has no final local result. - Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads. The branch is mergeable. ## Risks Large envelopes require extra upload calls during startup. The temporary data stays inside the private session directory and is removed before child startup or during failure cleanup. Individual child environment values still obey the operating system's native limits. No migration or configuration change is required. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c221589b69 |
fix: close sandbox process proxies after remote exit (#13777)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox agents use a local proxy to exchange ACP messages with a remote process. > - ACP keeps the proxy input stream open while it waits for a reply. > - The proxy received a remote exit but kept that input stream open. > - This left the proxy alive and could block later task retries. > - This pull request closes the input handle after a terminal remote event and lets final output drain. ## Linked Issues or Issue Description **What happened?** A sandbox process could exit while its local proxy stayed alive. During startup, this could appear as a handshake timeout. A later task retry could then stop because the previous process was still alive. **Expected behavior** The proxy should exit with the remote process status, even when ACP has not closed its input stream. Final output and diagnostics should arrive before the proxy closes. **Steps to reproduce** 1. Start a sandbox process-session bridge with a child that writes output and exits. 2. Start the generated local proxy and keep its input stream open, as ACP does during startup. 3. Observe that the proxy stays alive after the remote process exits. **Paperclip version or commit** Reproduced on master at `846336e5a0`. Related: #13272 covers legacy sandbox startup recovery. #13765 covers the task retry UI. This change fixes the proxy process lifecycle. ## What Changed - Close the proxy input handle on remote exit or error, and preserve the exit status. - Stop forwarding input after the terminal event. - Wait for process close in the test helper so output is fully drained. - Test successful exit, failed exit, and spawn failure with input held open in both output modes. Check complete delivery of 208 KiB of final output. ## Verification - Before the fix, all four remote-exit regression cases timed out. - After the fix, all six terminal-event cases pass. - Targeted runtime suite: 362 tests pass across the sandbox bridge, stdin queue races, real ACP spawn, and ACP engine suites. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - Full workspace tests hit local platform failures: embedded PostgreSQL fails during initialization, and runtime skill-cache tests fail with `EACCES` while renaming read-only directories on macOS. The affected server code is unchanged in this PR. Those suites pass in CI. - `pnpm -r typecheck` and `pnpm build` stop at the existing Rust runner step because this machine has no `cargo` executable. The CI typecheck and build gates pass. - [CI run 35658635907](https://github.com/paperclipai/paperclip/actions/runs/35658635907) passes all required gates. The first browser shard run passed all 12 tests but hit a GitHub 403 during report upload; the rerun passed, including upload. - Greptile scored commit `9c3c89148e` 5/5 with no review threads. The branch is mergeable. ## Risks The proxy now closes its input immediately after a terminal remote event. Tests cover final output delivery and preserve nonzero exit codes. Existing orphan processes still require cleanup or a server restart. Live sandbox startup needs validation after deployment; this change addresses the confirmed proxy hang and does not establish why an earlier remote process stopped. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted checks; full-workspace limits described above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (existing behavior restored; no documentation change needed) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45c99a0d06 |
fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Legacy adapters launch provider CLIs and expose connected tools. > - Existing defaults did not consistently grant full automatic permission. > - Remote Claude used a fixed tool list that omitted MCP tools and future tools. > - Direct Codex launches and OpenCode configuration also used narrower defaults. > - This change gives all these paths the same full-auto default as native runners. > - Explicit restrictive settings continue to work. ## Linked Issues or Issue Description Refs #13686. This PR is stacked on that native-runner and task-reassignment PR. Merge #13686 first. Related: #831 (constructed Claude agents), #1935 (adapter-switching permission defaults). ## What Changed - Use actual Claude permission bypass for local and remote runs and probes. Remove the fixed tool list so MCP and future provider tools are included. - Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`. Do not mark ordinary host execution as a sandbox. - Default direct Codex execution to approval and sandbox bypass, matching agent creation. Preserve explicit false, CLI profiles, sandbox modes, approval policy, and network restrictions. - Set OpenCode's full-auto runtime permission to `allow` for every tool and connection. Preserve the existing explicit opt-out. - Default Gemini probes to the same YOLO mode as execution. Make the legacy ACP `default` alias use `approve-all` for fresh and resumed sessions. - Add default, opt-out, remote, probe, connected-tool, and resume regression tests. Update adapter configuration documentation. - Other adapter paths already request full automatic permission or have no provider approval gate. ## Verification - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server and legacy ACP tests passed, including defaults, explicit opt-outs, remote launches, connected tools, and fresh/resumed sessions. - Greptile reviewed current head `8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security reviewer acknowledged the documented full-auto requirement. Acknowledged discussions are resolved. - Current head has **54 passing checks**. [PR checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The process-adapter signoff browser shard passed on one retry after its first attempt exceeded a three-second issue-run wait. - **Six native Claude/Codex real-provider cases passed on their first attempt, with cleanup passing**, against the combined branch: plans, reassignment, and backlog creation/status. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). This does not claim a real-provider run of every legacy adapter. - The live-tested revision is `a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs only in the corrected heartbeat test expectation; application code is identical. - The campaign result-enforcement job passed. The separate report publisher failed during frozen dependency installation because the trusted workflow's patched-dependency configuration does not match its lockfile. Passing case evidence remains downloadable from the workflow. - Full-suite coverage comes from CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. ## Risks - Missing permission settings now grant all provider operations, including connected tools. OpenCode full-auto also overrides ambient provider permission rules. An explicit Paperclip permission opt-out preserves restrictive behavior. - Claude refuses full bypass as root outside an identified sandbox. Ordinary host deployments must run Claude as a non-root user. Managed sandbox launches include the required marker. - These defaults do not grant additional Paperclip roles, connections, or company access. Existing controller authorization and governance still apply. - This PR depends on #13686. Retarget it to master after that PR merges. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d54b750111 |
Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset time when the runtime reduced provider metadata to a generic category error. Inspect terminal metadata in memory and retain only safe recovery labels and a parsed reset timestamp. Preserve the existing handling of other limits. Verified real child processes on both pinned ACPX runtimes, adapter and server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate a pre-existing chat regression from unrelated fixtures’ retry work. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5442f2d869 |
fix: repair managed Git launchers in sandbox projects (#13588)
## Thinking Path > - Paperclip runs agents in local and remote execution environments. > - Managed GitHub launchers select credentials for each Git operation. > - Remote launchers are written inside the project checkout as extensionless CommonJS scripts. > - An ES module project makes Node interpret those launchers as ESM, so they crash before credential resolution. > - When the launcher can start, empty identity variables also override valid repository and command-line Git configuration. > - This change gives the launchers their own CommonJS scope and clears empty identity overrides while preserving managed credential isolation. ## Linked Issues or Issue Description **What happened?** In a repository with `"type": "module"`, the managed `git` and `gh` launchers fail immediately with `ReferenceError: require is not defined in ES module scope`. The launchers use CommonJS but inherited the enclosing project's module type. Sandbox agents also report empty `GIT_AUTHOR_NAME` and `GIT_COMMITTER_NAME` variables and try to unset them for each command. With no managed identity available, the Git launcher recreated those empty values. `git commit` failed with `fatal: empty ident name`, even with explicit `user.name` and `user.email` configuration. **Expected behavior** Managed `git` and `gh` start in both ES module and CommonJS projects. Local commits with an explicitly configured identity work without manual environment cleanup. Managed credentials and captured identity continue to take precedence. Missing identity does not silently select the host user's details. **Steps to reproduce** 1. Create a sandbox project whose `package.json` contains `"type": "module"`. 2. Stage the managed GitHub launchers and run `git --version` or `gh --version`. Before this fix, the launcher fails at its first `require()`. 3. In a CommonJS project with no available managed identity, configure repository `user.name` and `user.email`, or supply them with `git -c`. 4. Run `git commit --allow-empty -m test`. Before this fix, both identity configuration forms fail with empty identity. **Paperclip version or commit** Reproduced from master commit `165b10bd9`. **Deployment mode** Sandbox execution. The shared launcher is also used for managed local and SSH execution. Related work: #13094 introduced the local-operation fallback; #13053 changes launcher discovery on Windows. Neither fixes empty identity overrides. Related identity work in #8945 and #8946 configures worktree authorship and does not remove these environment overrides. ## What Changed - Stage `package.json` with `"type": "commonjs"` in the launcher directory before the Node scripts. Keep the project's package configuration unchanged. - Leave inherited author and committer variables unset in the real Git process. When credentials are absent, require explicit Git identity configuration with `user.useConfigOnly`. - Clear empty identity merge overrides in staged shell profiles after environment merging. Preserve nonempty captured identity values. - Exercise real Git commits with repository and command-line identity, broker failures, and managed-user switching. Verify startup in ES module and CommonJS projects, shell cleanup, and captured identity preservation. - Document launcher module scope and local identity behavior in the execution GitHub identity contract. ## Verification - Confirmed both new local-commit regression cases fail before the fix with `fatal: empty ident name`. - Confirmed the new ES module project regression fails before the fix with `require is not defined in ES module scope`. - Focused launcher and shell tests: 28 passed. - `pnpm exec vitest run --project @paperclipai/adapter-utils --exclude '**/dist/**'`: 1,216 passed, 11 skipped across 58 files. - `pnpm --filter @paperclipai/adapter-utils typecheck` and `pnpm --filter @paperclipai/adapter-utils build`: passed. - `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the unchanged native runner because Cargo is not installed on this machine. - Full `pnpm test:run`: started locally; stopped the duplicate run after the complete CI suite passed. No local full-suite success is claimed. - CI on `99ea8050e`: all 53 checks passed (2 skipped), including full tests, typecheck, build, native runner checks, and browser checks. - Greptile reviewed `99ea8050e`: 5/5 with no findings or unresolved comments. GitHub reports no merge conflicts with master. - No live sandbox or GitHub push probe performed. ## Risks - The new package scope is confined to the run-specific launcher directory. It does not change the project's module type, launcher names, or credential selection. - Without a managed identity, an explicitly configured repository author can now create local commits. GitHub access remains subject to the existing credential broker. Global/system Git configuration, ambient credentials, and SSH identity remain isolated. - Managed identity still wins over repository settings. Missing local identity still fails instead of guessing host details. - New or resumed executions must stage the updated launcher and shell profiles. Existing processes retain their prior files and environment until refreshed. No database migration or sandbox image rebuild is required. - Revert this change to restore the prior behavior. ## Model Used - OpenAI GPT-6 via Codex, with code inspection, implementation, and local test execution. The hosted model variant and context window were not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (the affected adapter-utils package) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e26d787928 |
Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must continue tasks using user answers without losing earlier requirements or approval gates. > - The wake prompt mixed human decisions with prior tool evidence and repeated detailed question instructions. > - Those instructions belong with the question tool, with a short routing hint in the wake. > - The Runner evals need to prove that answers, approvals, and completed work survive later turns. > - This PR shortens the prompts, separates authenticated answers, and adds continuation tests with useful screenshots. ## Linked Issues or Issue Description Refs #13517. This is a follow-up to the merged onboarding skill and Runner E2E work. Related #13539 covers responses received while a run is active; this PR preserves its cases and adds continuation coverage. Existing continuation/recovery and question PRs were searched; none covers this prompt/documentation and eval change. **What existing behavior does this improve?** The instructions sent when an agent continues a task, the native human-input tool documentation, and the evidence captured by Runner full-stack E2E. **Current behavior** The wake repeats a long question-tool guide. Human answers appear alongside untrusted prior results. Screenshot capture can finish at DOM load while the task still shows a spinner, even when backend behavior checks pass. **Proposed behavior** Keep earlier requirements unless the user changes them. Treat clarification as distinct from approval. Give authenticated human responses a scoped field. Keep tool and agent results as evidence. Put detailed question behavior in the tool descriptor and retain one routing sentence in the native wake. Wait for the correct task and loaded conversation before taking screenshots. **Reason and benefit** Reduce repeated prompt text and make authority boundaries clear. Test that real question cards, later answers, approval gates, and completed child tasks still work. Make screenshots useful for human review. ## What Changed - Shorten shared continuation instructions for legacy and native runners. Separate authenticated user responses from tool results and agent summaries. - Remove the detailed question guide from native wake prompts. Keep its behavior in the canonical `request_human_input` descriptor and existing payload schema. Regenerate semantic contracts and fixture hashes. - Add five continuation cases across four local profiles. Add a dedicated choice-then-text case for native Codex and native Claude. All 22 cells join the shared full E2E campaign. - Cover revised scope, clarification without approval, hostile instructions in a handoff file, and reuse of a completed child after restart. Keep production instructions and fixed user facts. - Capture continuation screenshots only when the intended task and conversation have rendered. Add provider-free browser regressions for loaders and wrong-task capture. - Preserve current master’s extra tool and onboarding cases. The default campaign now contains 166 cells; 35 manual everyday cells remain separate. ## Verification - `pnpm -r typecheck`: passed after replay on current master. - `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed. - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm test:e2e:runner:browser-support`: 4 passed. These tests failed against immediate screenshot capture and passed after the fix. - Focused continuation and native-input tests: 36 passed locally. The tool-authority suite could not initialize embedded PostgreSQL locally, including one isolated retry; its 17 assertions did not run locally. The full remote server shards passed on this PR commit. - `pnpm build`: passed after replay on current master. `pnpm test:run` was attempted locally but hit the same embedded PostgreSQL initialization failure; the remaining local run was stopped after complete remote CI passed. This is not claimed as a full local test pass. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685): passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All server/chat/workspace/serialized shards, browser shards, Runner checks, typecheck, build, canary and policy checks passed. The isolated native Runner build and security checks also passed: 57 successful checks, with two expected Storybook skips. - Greptile reviewed the exact PR head at 5/5, with no findings or unresolved review threads. The PR has no merge conflicts. - [Live question-docs report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/): 3/3 passed at source `83dd132f2` before replay on master. Native Codex and Claude each asked a choice, waited, asked a text question, and saved both answers. Claude also passed a completed-child restart case. All three native turns are checked for absence of the old question block. - [Earlier continuation report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/): all five continuation cases passed on native Claude. The report retains campaign and revision provenance and separately shows two unresolved onboarding behavior failures. - [Before/after prompt report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/): full text, current recorded Claude inputs, and reproducible reference-token counts. The controlled wake comparison removes 401 reference tokens; the net counted input reduction is 339 after charging the larger tool description. These are text-size estimates, not measured billing savings. ## Risks - Prompt wording affects model behavior. Live results cover the stated cases, not every provider or conversation. Legacy profiles are registered but were not rerun for this change. - The optional continuation field changes prompt data only; there is no database migration or new production API. - Authenticated answer projection excludes generated summaries and agent-resolved interactions. It preserves the answer’s question or approval scope. - The screenshot guard can expose UI loading failures that earlier runs hid. Backend grading alone no longer makes those captures valid. - The two prior onboarding failures remain separate product issues: work before acceptance and a missing saved plan. This PR does not claim the entire onboarding suite passes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, code execution, and browser verification. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted tests above; the full local database-startup limit is documented - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d49f168381 |
fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must publish generated files so users can inspect their results after a sandbox stops. > - Legacy sandbox bridges blocked attachment listing and could not carry multipart binary uploads through the queue transport. > - The native runner has a separate verified file registration path that needs the same durable result. > - This pull request repairs legacy binary transport, makes native download receipts explicit, and reveals new outputs in the task Artifacts tab. > - Users can open generated files from either runner without a transport flag change. ## Linked Issues or Issue Description Refs #13355 for the existing native file publication path. Related filename fixes: #2615 and #4788. Related sandbox persistence work: #13376. This change repairs attachment delivery through the existing API; it does not add workspace persistence. **What happened?** The upload helper first lists task attachments to avoid duplicates. Both legacy bridge allowlists rejected that GET request with 403. A direct multipart upload also failed: the queue bridge accepted only JSON, excluded attachment uploads, and converted bytes to UTF-8 text. Enabling HTTP/2 alone did not fix the missing listing route. These failures occurred before attachment storage. **Expected behavior** Both runners can publish a workspace file, register its work product, bind it to a response, and return a working download. The file stays accessible after sandbox deletion. A new output opens the task Artifacts tab. The agent receives accurate errors and decides how to retry or report a failure. **Steps to reproduce** 1. Run a legacy agent in Daytona with the duplex bridge disabled. 2. Invoke the bundled upload helper with Bash on a PNG or PDF. 3. Repeat with the duplex bridge enabled. 4. Register the same file through the native runner with generic API tools disabled. 5. Retry registration, delete the sandbox, and compare the downloaded bytes with the original file. **Paperclip version or commit** The failing baseline was `f2c5e54dc`. This branch is rebased onto `6cfe4acff`. **Deployment mode** Source checkout with a local API and real isolated Daytona sandboxes. ## What Changed - Allow authenticated attachment listing, upload, and content download through both legacy bridge transports. - Add optional base64 body encoding to queue envelopes. Preserve the existing UTF-8 contract when the encoding field is absent. Decode binary bodies before forwarding them. - Preserve multipart headers. Bound raw bytes, encoded envelopes, and in-flight reservations. Retain timeout and uncertain-write behavior. - Preserve helper deduplication and return structured uncertain-write failures. Document explicit Bash invocation in live skills. - Add attachment IDs and content/download paths to native registration receipts. Reuse verified local and remote file reads, attachment storage, work-product registration, and response binding. - Preserve Unicode upload filenames and provide a valid Content-Disposition header. - Open the task Artifacts tab when new stored outputs arrive, including a closed desktop panel or mobile drawer. Deduplicate upload and registration events by object ID. Preserve manual selection on refetches, edits, and panel remounts. - Remove task artifact filters, the company Artifacts footer link, and the unassigned group heading and timestamp. ### Screenshot  This is the local display fixture. The image was generated separately and published through the attachment and work-product APIs. ## Verification - Post-rebase `pnpm -r typecheck` and `pnpm build` pass. - The post-rebase local `pnpm test:run` passed 12,369 tests before one existing conversation reset test timed out; all 33 tests in that suite pass when rerun with isolated test configuration. The aggregate command stopped before its remaining groups. GitHub runs the complete suite in separate shards. - All [GitHub verification checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350) pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks and two configured skips. The native-session recovery assertion initially raced its fire-and-forget Sentry report; all 13 tests pass locally, and the same-commit CI rerun passes all 170 suites (3,079 tests). - Live post-rebase Daytona: all three file-delivery tests pass. They cover the real Bash helper with the queue bridge, the helper with HTTP/2, and native `register_deliverable` with generic API tools disabled. - Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate registration, response binding, authorization controls, and byte-for-byte downloads after sandbox deletion. - Local focused coverage includes transfer bounds, malformed encoding, interrupted transfers, remote path containment, and native file verification. The attachment route suite passes all 32 tests, including an eight-case filename-header matrix for Unicode and special characters, inline and forced downloads, and full and partial responses. - Browser verification confirms image previews, persisted downloads, automatic Artifacts selection, and preserved manual selection after edits and reloads. Desktop/mobile component coverage passes. The latest UI cleanup passes its 10 affected tests and token gates. - Coverage limit: the Daytona tests call the real helper and native registration path directly. They do not replay a complete model-led image-generation task through the browser. Live command (requires a configured Daytona credential): ```sh PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts ``` ## Risks - Binary queue bodies use more memory because base64 adds encoding overhead. Transfer and process limits must remain aligned. - An interrupted write can have an unknown result. The bridge reports this state and preserves stable retry identities. - New artifacts intentionally change the active task tab. Existing history and repeated updates must not take focus again. - Transport flag defaults, server authorization, frozen skill snapshots, and completion policies remain unchanged. No schema migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f1905d34d |
fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects can now attach several source repositories. > - Task preparation still treated these sources as alternative workspaces. > - Sandbox sync preserved Git history only for the selected repository. > - A task needs every attached repository to complete work across the project. > - This pull request prepares all distinct project repositories and preserves their separate Git histories through sandbox restore. ## Linked Issues or Issue Description **What happened?** A user reported that a project with two repositories received only the first repository in Daytona. Repository-only project rows also reached the agent with null local paths. Managed checkouts with matching repository names could resolve to the same directory. **Expected behavior** Local and sandbox tasks receive every distinct repository attached to their project. Repository-only sources work without preconfigured local folders. Each repository keeps its own Git history and working files. **Steps to reproduce** 1. Create a project with two repository sources and no local folder paths. 2. Assign a task to the project and run it in Daytona. 3. Inspect the task workspace and the repository paths exposed to the agent. 4. Observe that the original implementation supplies only the selected checkout. Related change: #13010 added multiple repository selection. The open repository-catalog proposals #11234 and #11228 cover a different data model. This fix uses the existing project workspaces. ## What Changed - Materialize each additional distinct repository as an editable checkout inside the task root. Seed configured local sources with their current working files and retain task edits across runs. - Pass materialized repository paths to local agents and native sandbox task prompts. Apply existing run-scoped Git credentials to each remote clone. - Preserve each repository's Git history, dirty files, and restore baseline during sandbox staging and durable recovery. Apply each repository's ignore rules and the operator's workspace exclusions. - Keep same-name managed repositories in separate directories. Report additional clone failures before the task starts. - Add task-level, checkout, sandbox round-trip, environment-hint, and recovery-descriptor regression coverage. Document checkout and restore behavior. ## Verification - Red: the original implementation fails the sandbox test because the second repository has no Git directory. It also fails the same-name checkout test and both real-database task tests because repository hints have no local path. - Green: focused tests pass for one and two repository-only sources, local source edits, clone failures, per-repository credentials, separate Git histories, ignored files, and recovery from remote or durable seed state. - Live Daytona smoke passed with two disposable repositories through the production provider sync functions. Both repositories arrived with Git history. Commits from both restored locally. Ignored files stayed excluded. The disposable sandbox was deleted. - Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm build`. - Final focused coverage: 254 assertions across the six changed test areas passed across the serial run and an isolated rerun of the existing process-kill timing test. The live Daytona smoke also passed. - The local `pnpm test:run` overlapped source edits and retained stale transformed code. Its first phase reported 12,240 passed assertions, nine failed assertions, three hook failures, and one worker error; later phases did not run locally. This run is not claimed as green. Fresh focused tests verify the changes, and every general/workspace and serialized-server CI shard passes on the final commit. - Final CI is green on `93ab76763`: all test shards, all three browser shards, typecheck, build, runner verification, canary dry run, and security checks. The initial unrelated chat-delivery browser timing failure passed in the final CI run. Optional Storybook visual checks were skipped. - Greptile is 5/5 on the final commit with no unresolved review threads. Its checkout-race finding was reproduced with a failing test, fixed, and rechecked. ## Risks - Additional repositories need disk space and clone time. Access failure for an attached repository stops preparation. - Additional checkouts live under `.paperclip-repositories/` and keep independent histories. Changes stay in those task copies; they do not overwrite configured source folders. - Detached or reconfigured repository copies are retained under `.paperclip-runtime/detached-repositories/`. Sandbox recovery retains per-repository merge baselines. - No database migration, UI contract change, or new credential delegation is required. Referenced projects retain their separate read-only behavior. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and tool use. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5054c9ef9b |
fix(grok): stage the environment test from a host directory, and survive an absent workspace (#13416)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent runs through an adapter. Before you use an environment, the
adapter's environment test probes it and reports structured checks
> - The grok adapter's environment test stages the managed account into
a remote environment first. It gave the runtime its resolved `cwd` as
the host workspace directory
> - For a remote target that `cwd` is a path inside the sandbox. It does
not exist on the host that runs the test
> - A new workspace ignore scan reads that host directory. The scan
failed on the missing directory, and the error escaped the environment
test
> - The user got a server error instead of a result with checks
> - This pull request stages from a host temporary directory, sends the
remote path separately, and reports a staging failure as a check
> - The benefit is that the grok environment test always answers with
checks, and a workspace directory that does not exist no longer stops a
runtime preparation
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
The grok environment test failed with `Error: Workspace ignore scan
failed: git-ignore-scan-failed`. The route returned a server error, so
the user saw no checks at all.
**Expected behavior**
The environment test always returns `{status, checks}`. Every other
problem it finds (an invalid working directory, a command it cannot
resolve, a probe that times out) becomes a check with a level. A
credential staging problem must do the same.
**Steps to reproduce**
1. Configure a grok agent with a managed AI connection.
2. Point the agent at a remote (sandbox) environment.
3. Run the environment test for that environment.
**Paperclip version or commit**
Present on master. Both halves landed on 2026-09-12: the call site in
#13247, and the scan that it trips in #13353.
## What Changed
- `packages/adapters/grok-local/src/server/test.ts` stages the managed
account through a fresh host temporary directory and passes the remote
path as `workspaceRemoteDir`. This is the split `codex-local` and
`opencode-local` already use.
- A failure while staging becomes a `grok_environment_unprepared` check
with level `error`, instead of an exception that escapes the function.
The probe does not run after it, because there is no prepared
environment to probe.
- The temporary directory is removed in the existing `finally` block.
- `packages/adapter-utils/src/sandbox-managed-runtime.ts` treats a
`workspaceLocalDir` that does not exist as "nothing to sync". A
directory that is not there has no files for ignore rules to govern,
nothing to stage, and nothing to restore.
- Tests: the grok environment test now covers the managed-connection
remote branch, which had no coverage. `sandbox-file-sync.test.ts` covers
a preparation whose workspace directory does not exist.
## Verification
```
npx vitest run packages/adapters/grok-local/src/server/test.test.ts # 9 passed
npx vitest run packages/adapter-utils/src/sandbox-file-sync.test.ts \
packages/adapter-utils/src/sandbox-managed-runtime.test.ts # 90 passed
cd packages/adapter-utils && npx tsc --noEmit # clean
cd packages/adapters/grok-local && npx tsc --noEmit # clean
```
The two new grok tests fail against the old call site: the first asserts
the runtime never receives the remote path as its host workspace
directory, and the second asserts a staging failure becomes a check
instead of an exception.
## Risks
Low risk, and limited to environment preparation.
- The adapter change only affects the managed-connection remote branch
of one adapter's environment test.
- The `adapter-utils` change makes a preparation that used to throw now
continue with no workspace sync. A real workspace is unaffected, because
the directory exists in that case and the scan runs exactly as before.
- No migration. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
02c7175e72 |
feat(agents): hired agents inherit provider credential references (#13268)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters need provider credentials to start work > - A hired agent can lack the credential reference that its hiring agent already uses > - The child agent then cannot authenticate, even when the company has a valid credential > - This pull request copies matching credential references from the hiring agent to the hired agent > - The benefit is that hired agents can start with the provider access that the hiring agent already uses ## Linked Issues or Issue Description This change relates to [PR #9920](https://github.com/paperclipai/paperclip/pull/9920), which covers credential inheritance for other agent creation paths. This pull request covers hiring-specific inheritance and fixed Claude OAuth binding checks. ### What existing behavior does this improve? The agent hire route builds the child adapter configuration from the hire request only. ### Current behavior A hired agent does not receive matching provider credential references from its hiring agent. The child agent cannot run when the request omits the credential. ### Proposed behavior The hire route inherits matching credential references from the hiring agent. The request keeps priority. A Claude hire that supplies any Claude credential inherits none. ### Reason and benefit The child agent can use the provider access that the hiring agent already uses. The change copies references only and never copies raw token values. ### Breaking changes None. The change affects only hires that need an inherited reference. ## What Changed - Copy matching credential references from the hiring agent into the hired agent adapter configuration. - Preserve pinned versions, `required`, and `allowMissingOverride` fields on each copied reference. - Keep hire-request credentials ahead of inherited credentials. - Reject inherited fixed Claude OAuth bindings unless the parent agent passes company, adapter, and exact-binding checks inside the same transaction. - Add route and service tests for inheritance, precedence, and binding validation. ## Verification - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-auth-inheritance-routes.test.ts src/__tests__/agents-claude-oauth-binding.test.ts` — 60 passed. - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-idempotency-routes.test.ts src/__tests__/agents-service-secret-bindings.test.ts src/__tests__/secrets-service-user-secret-owner-scoped.test.ts` — 28 passed. - `tsc --noEmit` in `server/` — the error count matches the merge base, with no error in either changed source file. - `git diff --check` — clean. ## Risks Low risk. The route copies references, not raw tokens. The request keeps precedence. Company, adapter, and exact-binding checks protect the inherited Claude OAuth path. ## Model Used Codex, OpenAI GPT-5, with code execution and review support. The implementation commit predates this pull request handoff. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78ce96a48b |
fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner supplies each agent with instructions, assigned skills, and tools. > - Native Claude lost its skill snapshot before provider launch. It also changed exact model IDs to aliases. > - The read permission mode denied ordinary Paperclip reads because this runner has no interactive approval handler. > - This pull request restores the missing context and permits only assigned tools that Paperclip defines as reads. > - Other requests that need approval stop with a clear action for the operator. They do not retry automatically. > - Agents can complete read tasks. Users can see why a restricted task stopped. ## Linked Issues or Issue Description **What happened?** Native Claude could not find assigned skills. The ACP layer could replace an exact model ID with an alias and then fail identity verification. The default `approve-reads` mode denied Paperclip read tools. A denied write showed a generic transport failure. The tool descriptions also called live operations “mock” operations. The completion prompt said to call a completion tool once, although the protocol can reject a claim and require a corrected call. **Expected behavior** Load assigned skills before launch. Keep the selected model ID. Allow assigned Paperclip reads. Stop an operation that needs approval with clear instructions when no approval handler exists. **Steps to reproduce** 1. Configure a native Paperclip Runner agent with ACPX Claude and an exact model ID. 2. Assign a skill and ask the agent to use it. 3. Select the read permission mode and ask the agent to read task context and list documents. 4. Ask the agent to write a document. Check the task state and recovery message. **Paperclip version or commit** Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`. **Deployment mode** Local native runner with real Claude, Rust runnerd, the Paperclip server, and embedded PostgreSQL. Related: #13196 fixed remote skill staging in the legacy `claude_local` adapter. The native runner uses a separate path, which this pull request fixes. ## What Changed - Carry the runtime context through Rust and the ACPX sidecar. Load assigned Claude skills after the provider lifetime lease is held. Refresh the files on each open. - Write the exact requested model ID into the isolated Claude settings. - Grant exact MCP permissions for the intersection of assigned tools and Paperclip's read catalog. Provider hints cannot grant access. Existing task-control permissions stay in place. - Stop requests that need an unavailable approval handler. Preserve the typed error through the server. Show “Approval required” on the task and require operator action without automatic retry. - Label the setting “Allow Paperclip reads.” Remove “mock” from live tool descriptions and regenerate the contracts. - Change one completion-prompt sentence to require one accepted result. Add regression tests and update the runner documentation. ## Verification - Red/green regression tests cover read admission, unavailable approval handling, server recovery, and the task error message. - A real Claude read trial failed before the fix and completed after it: [red trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1), [green trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1). - Full browser tests used real Claude and the production tool authority. Reads completed. A write stopped with “Approval required.” No document was created and no automatic retry was scheduled. All nine checks passed: [events and screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc). - A separate full browser test assigned a skill, invoked it, and completed with a marker absent from the task prompt: [skill evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52). - Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests, 27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm build`. The server/UI red-green checks and `pnpm check:token-gates` also passed. The full suite passed in CI, including all general and serialized test shards, all three browser shards, and runner `check:all`. The duplicate unsharded local full-suite run was stopped after CI passed. Braintrust links require project access. The traces contain provider/runner events and app outcomes. They do not contain raw model HTTP requests. ## Risks - `approve-reads` now allows assigned Paperclip reads. Other operations that need approval stop the turn. Users must review the operation and change permissions before retrying. - The Claude settings depend on the pinned ACP and SDK behavior. Automated tests and real Claude trials cover this boundary. - The native Codex path is unchanged. The ACPX Codex fallback shares the clearer approval failure handling. - Local execution was tested end to end. Remote execution was not run. This change has no database migration. ## Model Used OpenAI GPT-6 through Codex. The agent used reasoning, code execution, and browser tools. The exact deployment ID and context-window size were not exposed in the session. Live acceptance tests used `claude-sonnet-5`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2c5e54dca |
fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - A task keeps its instructions, progress, and files when its assigned agent changes. > - The replacement runner lost the interrupted run's context and could overwrite an existing draft. > - A saved message also stayed attached to the former agent and could reopen the task after the replacement finished. > - File tasks could report Done with only a local path that the user could not open. > - This pull request transfers handoff context and saved messages, and makes requested files accessible through the existing attachment contract. > - Users can change agents and collect completed work without repeating instructions or confirming bookkeeping. ## Linked Issues or Issue Description Refs #13338. Builds on merged #13354 for queue admission and #13353 for remote workspace retry. #10123 concerns restricted recovery-model escalation; this change instead covers ordinary native handoff and file completion. **What happened?** Codex wrote a draft before a user assigned the task to Claude. The replacement lacked continuation context and replaced the draft. A queued user message could later restart the former agent and reopen the completed task. Separately, a runner could finish a requested file but return only a machine-local path. Remote native runs had no bound file publication tool. **Expected behavior** The replacement reads and preserves existing work, receives saved messages once, and keeps each message's author. The former agent stays stopped. A requested file has a working attachment or accessible work product before Done. Text-only tasks do not require attachments. **Steps to reproduce** 1. Ask Codex to save three newsletter names and then wait. 2. Queue an instruction to keep those names and expand the draft. 3. Use Interrupt and assign to select Claude. 4. Verify the original names survive, the result has a working download, and only the source and replacement runs exist. 5. Ask either provider for a Markdown checklist and open the file from its completed response. ## What Changed - Carry the exact same-task interrupted run's summary, semantic receipts, and history into handoff context. Tell the replacement to inspect existing files before editing. - Adopt saved ordinary task comments into the successor's receipt under the task lock. Preserve authors and separate mention, chat, and interaction contracts. - Prevent a former-assignee comment wake from reopening a completed task or starting a stale execution. - Reject workspace-only, fabricated, and cross-task file completion references with actionable runner feedback. New file output also needs a matching current-run publication receipt and asset filename/size/hash, or an accessible work product registered by the current run. Prior output can remain context alongside a current file, or be verified and re-registered internally. Authorized chat attachment reuse retains its verified current-run clone receipt; older receipt shapes require an intact matching source. - Bind remote file reads to the active environment runner and reuse the existing attachment and work-product publication path. - Enforce workspace confinement, regular single-link files, stable identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the native session fingerprint for the updated tool contract. - Contain rejected remote signals and protocol-failure cleanup, including logging failures. Preserve the original cleanup rejection for its owner; a rejected operation never supplies stop acknowledgement or cleanup proof. - Allow exactly one maximum-size base64 file through the native SSH command adapter, preserving a finite output cap. - Document handoff and accessible file completion rules. ## Verification - Each observed bug has a failing regression before its fix. Final post-rebase integration passed 732 tests across 13 files before the final receipt and signal guards; final affected results are below. - Publication provenance and compatibility: 8 provenance regressions and 2 compatibility regressions failed before their fixes; the final four affected suites pass 53 tests, including mixed old/new references and real authorized chat reuse. Controls cover old attachments and work products, filename/size/hash/origin mismatch, missing/wrong receipts, current-run publication, same-run durable proof, internally re-registering preserved bytes, and no-new-file follow-ups. - Remote signal rejection: the real Node subprocess previously exited 1 when the production launcher signalled a deleted sandbox. It now stays alive for both a failed signal and failed logging; all 349 executor tests pass. The failed signal still provides no termination proof. - Remote file reader and SSH command boundary: 33 tests passed, including real Linux descriptor reads and the actual SSH adapter subprocess output cap (network executable replaced by a deterministic fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap fails. - Live local Claude and Codex Stop journeys preserve the saved file, deliver queued instructions once, and reach Done with two total runs. The handoff journey preserves the original names and download with exactly two runs. Both local providers deliver exact checklist files without a completion confirmation. - Live combined Daytona verification passed: the original failed task's Retry reused its sandbox; a selected Git subfolder produced an exact downloadable file; warm and deliberately resumed Claude runs took about 33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134 seconds of contention/backoff. Both cloud downloads retained exact bytes after the two owned sandboxes were deleted. - The sandbox-deletion retest identified a separate ignored promise in protocol-failure cleanup. Two real Node subprocess regressions failed under fatal unhandled-rejection policy before the fix; all 36 protocol, lifecycle, and integrity tests now pass. The original close promise still rejects to its owning runtime. The final live retest passed: a normal Claude Daytona task completed in 132.352 seconds, then its sandbox was deleted. Thirteen samples over 361 seconds confirmed the same controller stayed healthy, the task stayed Done with unchanged run IDs, and its attachment retained exact bytes. The post-deletion browser download passed with zero page errors; all five owned sandboxes are confirmed absent. - Full repository typecheck and build passed on final commit `de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758) passed: 32 successful checks and two conditional skips. The earlier mixed-source full local test invocation was deliberately stopped before rebase, so no pristine green full local aggregate is claimed. Its known failures passed in later affected suites. ## Risks - Handoff may adopt only ordinary comments from its validated former owner. Other delivery contracts must remain independent. - File verification fails closed if a remote file changes during reading. The runner must retry publication or explain a blocker. - The updated session fingerprint starts a fresh provider process where needed to install the new tool contract. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI `gpt-6-astra` through Codex, with reasoning, code execution, browser testing, and tool use. The context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6809314a3f |
fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox tasks need a workspace and a provider session before work can start. > - A selected Git subfolder is a valid workspace, but it is not a repository fetch source. > - A resumed sandbox must keep one resource identity when the provider fills in its region. > - Workspace reuse does not prove that a provider session has started. > - A recorded continuation must not leave an old failure blocking Retry for a newer attempt. > - This pull request fixes those setup and recovery boundaries while preserving ownership checks. ## Linked Issues or Issue Description **What happened?** A Daytona task failed before provider startup when its project folder was inside a parent Git repository. After the folder was repaired, Retry resumed the same sandbox and verified its workspace sentinel, but workspace preparation reported that the lease was no longer active. A later attempt could fail because the resumed workspace had no provider checkpoint. The old run also kept a recovery projection after the server recorded its explicit successor, hiding Retry for the new failure. **Expected behavior** Sync the selected folder without importing parent files or history. Keep the existing sandbox identity stable. Allow a new provider session only with proof that its exact session has never started. Let the current failed attempt retain Retry once the old recovery has a recorded successor. **Steps to reproduce** 1. Select a subfolder of a Git repository as a project workspace and start a Daytona task. 2. Leave the target region unset. Release its reusable sandbox after setup fails, then resume and realize the workspace. 3. Retry a provider setup that failed before any session directory or checkpoint was created. 4. Record an explicit successor for a native failure, fail that successor during setup, and inspect Retry in the task thread. **Paperclip version or commit** Reproduced against `8d1f0c20a` with new failing regressions before each production fix. **Deployment mode** Local controller with a Daytona environment. Automated tests use deterministic provider fixtures and real filesystem operations. Related: #13338, #13163, #13264, #13349. The open work-folder stack in #13264 includes a broader fresh-session authority change. This patch addresses the independently reproduced startup failure with existing durable bootstrap proof and atomic directory creation. It does not include the work-folder migration or credential changes from that stack. ## What Changed - Classify only the selected Git repository root as a fetch source. Sync subfolders as directories and preserve their enclosing Git ignore rules on upload and restore. Recognize the shared scheduler’s completed non-repository result so ordinary folders still sync; timeout, cancellation, and output-limit failures remain closed. - Remove the placement target from Daytona account cache identity. Keep API endpoint, credential digest, company, environment, driver, and sandbox ID boundaries. - Preserve closed-lease admission until a sentinel-verified resume reopens that same resource. - Show the already-recorded explicit successor of a resolved recovery. Keep the old failure evidence and unresolved holds. No historical status writes occur. - Permit a resumed workspace to create a new session directory only with matching durable identity, zero connections and events, untouched bootstrap commands, no backup, and an absent remote session. Claim the directory atomically. Existing, partial, or ambiguous state still blocks startup. ## Verification - Final head `c43c7403f`: [CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2), with 32 successful checks and 2 conditional skips. This includes every server, UI, package, serialized-route, browser, native-runner, typecheck, and build gate. Greptile reviewed the same head at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168). - New regressions failed before each of the four production fixes. Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git regression also failed before its fix; 135 affected Git/sync/Codex tests then passed. - Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native executor, projection, and TaskChatThread: 494 passed, including existing/partial state, wrong identity, prior connections or turns, backups, unavailable proof, and unresolved recovery controls. - Local full-repository typecheck, build, and token gates passed on the final head. Local CLI: 485 passed. Complete single-worker package rerun: 3,224 passed, 19 skipped. - Local verification is an aggregate with recorded retries, not one pristine green invocation: the general-server run began on `e721a920a` and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex scheduler failure is fixed above; the socket and workspace-runtime timeout failures passed unchanged in focused reruns. Both Inbox failures passed unchanged in the full 27-test Inbox file; database/shared-package failures passed in the single-worker package rerun. The supplemental local serialized-route rerun remains in progress; all five corresponding final-head CI lanes passed. - The first final-head CI attempt hit a Daytona fixture-readiness race and a signoff-browser heartbeat receipt timeout. One supported unchanged failed-job rerun passed both and the aggregate gates. Live combined user-journey verification is tracked in the related follow-up; this PR's provider tests use deterministic fixtures and real filesystem checks. ## Risks - A selected subfolder uses directory sync and does not carry parent Git history. Its ignored files stay local. - The target region remains a creation setting and part of workspace reuse policy; it does not split the account identity of an existing sandbox. - Incomplete or conflicting provider state still fails closed. This change does not erase a session, infer completed work, bypass a user decision, or replay uncertain actions. - A later remote setup failure can leave a claimed partial session directory. It remains blocked rather than being overwritten. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9031516a7e |
fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy conversation adapters can run in Daytona sandboxes. > - A server restart during provisioning can occur before the invocation event exists. > - Recovery then lacks the old adapter identity and leaves a hold that ordinary user retries cannot clear. > - A remote launch can also fail when its host relay looks for Node in the sandbox PATH. > - This pull request records the adapter at claim time and restores explicit user continuation after verified cleanup. > - Users can recover from the task or inbox while the failed run and uncertain action history remain intact. ## Linked Issues or Issue Description Refs #13237, #13239, #13254. Those changes cover recorded conversation runs, native user continuation, and explicit remote Stop. This change covers legacy failure before `adapter.invoke` and exact task/inbox Retry. Refs #9771 for overlapping generated-command quoting. This change also supplies the absolute host Node executable. Refs #13163 and #13264 for the separate native restart and retained-workspace work. **What happened?** A legacy Daytona run interrupted during provisioning became `process_lost` without an invocation event. Recovery preserved an execution hold, and Retry or a new task reply could not resume it. Cleanup could also run before the Daytona plugin was ready. On a macOS host, a subsequent ACP relay launch failed with `env: node: No such file or directory` because the remote launch environment did not contain the host Node path. **Expected behavior** An interrupted conversation can continue after its previous execution stops. Explicit Retry and new user replies should start a fresh turn with the task history. Cleanup failures must remain visible and recoverable. The host relay must use the host Node executable. **Steps to reproduce** 1. Use a legacy Claude adapter with a Daytona environment. 2. Interrupt the server after it acquires the sandbox lease and before it records `adapter.invoke`. 3. Restart and inspect the task hold. 4. Retry from the task or inbox, or send a new task reply. 5. Confirm the old sandbox has stopped and one new response arrives. **Paperclip version or commit** Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`. The branch is rebased on `51b0e01ea`, including #13261 and #13270. **Deployment mode** Built from source on macOS with a real Daytona sandbox and the legacy Claude ACP adapter. ## What Changed - Count new browser specs with the scheduler's median duration in the shard-balance check. This fixes a false policy failure after new specs arrive from both branches. The balance threshold is unchanged. - Persist server-owned adapter identity in the queued-to-running claim before provisioning starts. - Wait for provider plugin startup before restart cleanup. Keep failed cleanup leases as active ownership blockers. - Admit exact board retries and new user comments after verified termination. Retain the old run, task history, approvals, and unknown action outcomes. - Adopt repeated Retry requests. Permit one scoped cleanup attempt per explicit user Retry after the automatic limit, with an activity record. A later user Retry can recover after a transient provider failure; automatic attempts remain capped. - Resume replies deferred during cleanup, including historical legacy startup failures. - Launch the host ACP relay through the absolute host Node executable. - Add a task-level Retry button and return actionable blockers when retry admission is refused. - Add database regressions and three browser recovery journeys. Exclude installed third-party dependency skills from the shipped-skill audit. ## Verification - Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution retains the saved-message recovery, local stop receipts, and wait reasons from #13270 alongside exact legacy Retry support. - Real Daytona: interrupted the server after lease acquisition and before adapter invocation. Restart cleanup confirmed provider termination. Task Retry cleared a seeded historical hold and a real Claude agent returned `Recovery verified.` in the task. Removed the disposable sandbox and environment after testing. - All three browser recovery journeys passed again after the final rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh successor, completed the task, preserved the failed run, and retained the answer after reload. - All 29 e2e/server shard-partition tests passed. The balance check now uses the scheduler's median fallback for unmeasured specs, with the same balance threshold. - Server typecheck passed after rebuilding the generated runner dependencies. The combined recovery/route run passed 136 of 137 tests. Its remaining route test timed out during the first cold module import at its explicit 10-second limit; an isolated rerun reproduced that timeout and passed the other 51 route cases. The complete CI suite passed on this head. The same route file passed all 52 cases in CI, including the first cold import in 7.5 seconds. - Before the final rebase, recursive typecheck, full build, UI token gates, 132 targeted server tests, and the complete [CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085) passed. The subsequent CI failure was the shard-balance accounting mismatch fixed here. - Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable findings. The complete [current CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949) passed on attempt 2. All test, typecheck, build, and canary jobs passed on the first attempt. Docker setup timed out fetching BuildKit from Docker Hub; retrying that job and its dependent aggregate succeeded. ## Risks - Recovery admission changes executable authority. Company, task, agent, user, approvals, process ownership, and provider termination checks remain required. - Explicit continuation starts a fresh conversation with history. It does not certify unknown external action outcomes or rerun non-conversation adapters automatically. - Changing task status alone does not clear an execution hold. The task now offers an explicit Retry action. - Historical adapter claims and invocation events take precedence over current agent settings. Known process or webhook runs retain their hold. Pre-upgrade rows with no adapter evidence may receive only a new explicit user turn after termination proof; they do not become eligible for automatic replay. - No schema migration or sandbox-image change is required. This branch has not been deployed to production. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution, browser automation, and test execution. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad4f0b5867 |
Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtime settings can bind organization secrets to an adapter environment > - Paperclip redacts plain environment values when it returns a saved agent to the UI > - A saved-agent test sent the redacted `CODEX_HOME` value back to the server > - Codex ACP also received the API key without an ACP API-key authentication request > - This pull request restores saved environment values for tests and selects API-key authentication for Codex ACP runs > - The benefit is that Codex agents can test and run with an organization-scoped OpenAI API key ## Linked Issues or Issue Description **What happened?** Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`. Secret normalization rejected that placeholder. Remote Codex ACP runs received `OPENAI_API_KEY`, but session creation stopped with `Authentication required`. **Expected behavior** Paperclip must use the saved `CODEX_HOME` value when it tests an existing agent. Codex ACP must select API-key authentication when `OPENAI_API_KEY` is available. **Steps to reproduce** 1. Create an organization-scoped secret named `OPENAI_API_KEY`. 2. Give a Codex agent access to the secret. 3. Save the agent runtime settings. 4. Test the saved agent again. 5. Run the agent in a remote sandbox through ACP. **Paperclip version or commit** Reproduced on master before commit `68c17709d7c051a804a416263e2e08920f1dfcb1`. **Deployment mode** Self-hosted server with a remote sandbox environment. **Installation method** Built from source. **Agent adapter(s) involved** Codex. ## What Changed - Send the saved agent ID with adapter environment tests. - Restore redacted plain environment values from the saved agent before test-time secret resolution. - Select the Codex ACP `api-key` authentication method when `OPENAI_API_KEY` is present. - Add focused regression coverage for saved-agent tests and remote ACP launch configuration. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/test-agent-setup.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `git diff --check` ## Risks - Low risk. The test route reads saved configuration only when the request supplies a compatible agent ID and the caller can update that agent. - The Codex ACP change applies only when `OPENAI_API_KEY` exists and no explicit `DEFAULT_AUTH_REQUEST` exists. - There are no schema migrations or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. The context-window size is not exposed in this runtime. The model used reasoning, repository search, file editing, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
87b3e5fc61 |
fix(adapter-utils): stage selected skills into the sandbox for a remote Claude ACP run (#13196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A user selects skills for an agent, and the host materializes those skills into a bundle the agent reads > - An agent can run in a remote sandbox, where the host must stage every file the agent needs > - On the Agent Client Protocol lane the host built that bundle and then named its host path in the prompt, but it never staged the bundle into the sandbox > - The agent therefore read a path that does not exist inside the sandbox, and the run failed on the missing skill file > - The command-line lane of the same adapter already stages a `skills` asset and reads the in-sandbox directory back from the staged runtime > - This pull request carries that proven pattern to the Agent Client Protocol lane, so a selected skill reaches the agent in a remote run ## Linked Issues or Issue Description No public issue exists for this change. The description follows. **What happened?** A remote run of the Claude adapter on the Agent Client Protocol lane could not read any selected skill. The host builds the skill bundle in its own state directory, then writes that host path into the prompt as `Skill root: <path>`. The remote seam of that lane staged one asset only, the configuration seed. It staged no skills asset, so no skill file crossed into the sandbox. The agent then tried to read the skill file at the host path, and the read failed with a missing-file error. **Expected behavior** A remote run receives the skills the user selected, and the prompt names the directory that holds those skills inside the sandbox. **Steps to reproduce** 1. Select one or more skills for an agent that uses the Claude adapter. 2. Start a run for that agent in a remote sandbox on the Agent Client Protocol lane. 3. Ask the agent to read the skill file at the path the prompt names. The file is not there. **Agent adapter(s) involved** The Claude local adapter, on its Agent Client Protocol lane. The shared engine in `packages/adapter-utils` carries the prompt rewrite. **Additional context** The command-line lane of the same adapter already stages a `skills` asset and remaps onto the staged directory. This change reuses that mechanism instead of adding a new transport. One other adapter shows the same host-path shape on its own Agent Client Protocol lane. That lane is tracked separately and this pull request does not change it. ## What Changed - Return the host skill bundle directory from the Claude skill runtime step, and carry it through the remote managed-home context to the staging seam. The value is null for a non-Claude agent, for a run that selects no skill, and for a run whose selected skills all fail to materialize. - Stage that bundle as a `skills` asset on the Claude Agent Client Protocol remote seam, and only when the run selected a skill. - **Stage that asset with `followSymlinks: false`.** The bundle holds an owned copy of each selected skill, and the copy step never copies a symbolic link at the root or at any depth. So the bundle contains no symbolic link, and staging has none to follow. Refusing to follow one also stops a link planted in the bundle directory after the copy from pulling an unrelated host file into the sandbox. A regression test walks the real adapter sources and pins the reviewed `followSymlinks` value at every skills staging site, so a new or changed site fails the test. - **Drop a skill whose staged copy has no usable `SKILL.md`** from the prompt, the skill identity, the command notes, and the bundle, and log which skill was dropped and why. Without this, a skill whose copy failed, or whose `SKILL.md` is a symbolic link the copy step skips, stayed advertised in the prompt while its file was absent — the same missing-file symptom this change exists to fix. - Rewrite the `Skill root:` prompt line, the skill identity, and the command notes onto the in-sandbox directory. The rewrite runs in the engine, after the workspace placement returns the staged runtime. A compatible session resume reuses the cached staged runtime, so the rewrite runs on that path too. - Keep the session fingerprint on the host-independent skill identity. A change to the selected skill set still invalidates a warm session, and the volatile sandbox path stays out of the hash. - A local run, and a run with no selected skill, keep their current behaviour. ## Verification - `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` — 28 of 28 pass. - `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` — the new engine tests and the staging-site tests pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`, and the same check on the adapter package — both exit 0. - The end-to-end test drives the lane against a local sandbox stand-in. It reads the skill root out of the prompt the runtime received, and then opens the skill file at that path. That is the reported symptom, proved closed. - The new tests carry a sensitivity control. Restoring only the production files to their previous content fails 7 of the 9 new tests. The other 2 do not depend on production code: one is a parser unit test for the source scanner. ## Risks Low risk, and the change is a two-way door. A revert restores the previous behaviour exactly. - **Scope.** The change touches one adapter lane. It does not change the local lane, and it does not change any other adapter. No existing staging site changes its `followSymlinks` value. - **The staged bundle and the workspace.** The staged skills land under the runtime directory inside the workspace. The workspace restore excludes that whole runtime directory, so the staged skills never return to the host worktree. A test proves the exclusion end to end. - **Session reuse.** The rewritten path never enters the session fingerprint, so it cannot invalidate a warm session, and a compatible resume applies the same staged path. - **Direction of data.** Files move from the host into the sandbox only. The change adds no path that writes sandbox content onto the host. - **A dropped skill.** A skill with no usable `SKILL.md` is now absent from the prompt instead of named but unreadable. The run logs the skill and the reason, so the cause is visible. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking and tool use, through Paperclip agents. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` passes — 28 tests. - [x] `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` passes. - [x] `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - [x] All continuous-integration gates are green. - [x] The automated review reports no open finding against the current head. ## Required Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant bug report template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect this change - [x] I have considered and documented the risks above - [x] All continuous-integration gates are green - [x] The automated review score is 5/5 with no open current-head findings - [x] I have addressed every reviewer comment that applies to the current head **Note on the branch history.** This branch first carried a different change: a filename-based admission filter that refused to stage files such as `.env` from a skill directory, together with a switch from symbolic-link bundles to copied bundles. That approach was rejected and **reverted** on this branch. It does not match the documented trust boundary, because the host already delivers credentials into the sandbox on purpose, and replacing the symbolic-link bundles broke live editing of a skill. The revert is in this branch's history. The file that work changed, `packages/adapter-utils/src/server-utils.ts`, is byte-for-byte identical to `master` here and is not part of this diff. Earlier review findings that name that file target the reverted code. All of them are resolved, and the automated review passes on the current head. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d1ba17eeca |
fix(adapter-utils): fail fast when the sandbox control channel is lost mid-turn (#13158)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter utilities run agent turns and report their results to the control plane > - A lost sandbox control channel can leave an agent turn without a result > - The host then waits for the full adapter timeout instead of reporting the loss > - This pull request adds a push loss signal and a bounded host wait > - The benefit is a prompt failure terminal when the agent stops answering ## Linked Issues or Issue Description **What happened?** A sandbox control channel loss during an Agent Client Protocol turn left the host waiting for the four-hour adapter execution timeout. **Expected behavior** The host should detect the terminal channel loss, stop the turn, and report a safe failure without waiting for the agent. **Steps to reproduce** 1. Start an Agent Client Protocol turn through a sandbox adapter. 2. Close the duplex control channel while the turn remains active. 3. Observe the host response before the adapter timeout expires. **Paperclip version or commit** Test the pull request commit set at `10b6bbc5525a79fd575298607dd5a25ae448fc8a`. **Deployment mode** The change applies to sandbox-backed adapter execution. ## What Changed - Add `onLoss(listener)` to the duplex bridge handle. - Register the loss listener at turn start and read losses latched before turn start. - Cancel the turn on loss and arm a 30-second host deadline. - Close the stream locally when the deadline wins and create a host terminal. - Derive the public error from the closed `DuplexLossReason` enum. - Add tests for loss order, cancellation, timeout, and safe error output. ## Verification - Run `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Run `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts -t "run-disposition seam"`. - Confirm that the full pull request workflow passes. ## Risks The new deadline changes a lost-channel path from a long wait to a host-built failure after 30 seconds. Orderly completion keeps its existing behavior. The deadline race against a pending `turn.result` has no direct test. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The runtime does not expose a more specific deployment version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018ca5daaf |
fix: verify ACP Stop and preserve safe continuation (#13119)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task controls coordinate provider execution and queued user messages. > - Stop could finish before an embedded ACP provider stopped its tools. > - A later request could be held for reconciliation without a clear task response. > - A restored provider could also retain the stopped run's API credential. > - This pull request verifies provider termination and preserves safe session continuation. > - Operators can continue known-safe work and see why uncertain work cannot start. ## Linked Issues or Issue Description **What happened?** Stop could leave an embedded ACP provider running. A queued follow-up followed by “go” could fail before it reached the provider. Task chat could show a generic missing-response message. Even a restored session could use the previous run's credential and fail its task update. **Expected behavior** Stop waits for confirmed provider termination. A later explicit wake continues the same compatible session only when recorded actions have known outcomes. It carries pending comments and the current run's environment. Uncertain actions retain a visible reconciliation hold. Composer Stop preserves the existing pause rule: conversation can continue while paused, but task work requires Resume. **Steps to reproduce** 1. Start an embedded ACP task. 2. Send a second request while the provider is running. 3. Interrupt the run, then send “go”. Also test composer Stop followed by Resume work. 4. Check that the request is delivered once and that the provider can complete the task through the current run's API credential. 5. Repeat with an unfinished write. Confirm that the write stops and that further execution stays blocked with a visible reason. **Paperclip version or commit** Built from source on master at `3bc60dd8b` plus this branch. **Deployment mode** Local source build with an isolated embedded PostgreSQL instance. Refs #11183. Refs #12552. Those changes address recovery after operator cancellation. This change also covers embedded ACP termination, session proof, pending-comment delivery, and task feedback. ## What Changed - Propagate Stop into embedded ACP and wait for bounded adapter cleanup and provider exit. Retain the actual ChildProcess object for forced termination on all platforms; never signal a recycled numeric PID. - Preserve interrupted checkpoints only for acknowledged, local, persistent sessions with settled reads or no tools. Keep writes, incomplete actions, and forced termination blocked. - Restore the same compatible provider session with the current run's environment. Reject fresh-session fallback for an interrupted checkpoint. - Adopt pending comments on the next explicit wake. Stop alone does not dispatch them. - Share the execution-blocker rule across dispatch, Resume, and task detail. Show Stopped or Couldn't start with the recorded reason. Resolve the stopped agent for the run link, including reviewer runs. - Keep execution reconciliation holds intact when generic recovery sees queued comments or healthy child tasks. - Add process, service, component, and browser regression coverage. Fix disposable database cleanup and React test settling exposed by the full suite. ## Verification - Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed all three `acp-stop-continuation.spec.ts` browser journeys. They use an actual ACP child process and require task completion through the agent API. - Passed 165 adapter execution, operator-stop, and child-process control tests, 17 queued-comment route tests, and 65 tests in the two adjusted UI suites. Earlier focused recovery, heartbeat, and task-control tests also passed. - Manually used the browser to queue a request, Stop, send “go” while paused, and Resume. The same session answered once and moved the task to Done with the current run's credential. - Manually interrupted an unfinished write. Its file size stayed fixed for five seconds. “Go” showed the reconciliation reason and did not start another provider prompt. - Separate live Claude ACP smoke checks confirmed that Stop ended a disposable local write and that a no-tool interruption could resume the exact provider session. The browser fixture does not call Drive or another external app. - Passed all 5,615 UI tests and 3,090 other workspace tests. The CLI and general server groups pass with targeted retries: two transient server failures passed together on retry, and two embedded-database startup failures passed after removing abandoned shared-memory segments from this task's completed browser fixtures. All 144 serialized server suites completed, with 2,189 tests passing after two transient HTTP socket failures passed on retry. - Passed all 135 heartbeat process/recovery tests, including a deterministic regression that failed before the recovery-sweep fix. - Passed 18 dispatch integration tests, including stopped-reviewer links, company boundaries, and malformed run IDs. - Greptile is 5/5 on `7dd170d83`, with zero unresolved review threads. The security scan and all required CI gates pass for the same commit. ## Risks - Safe continuation depends on complete tool reporting and a restorable local provider session. Unknown outcomes remain blocked and require reconciliation. - Provider cleanup can take time. A timeout does not grant replay permission. - The change adds optional adapter context fields and an optional issue projection. It does not change the database schema or require a migration. - Test cleanup truncates company data only in a disposable test database. ## Model Used OpenAI GPT-6, running as Codex with repository tools, code execution, and browser interaction. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3bc60dd8bf |
fix(adapters): probe Git context in the remote workspace (#13116)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Remote environments realize the workspace at a provider-owned path. > - Run startup probes Git and network context before it starts the agent. > - The probe used the controller path inside the remote environment. > - A missing directory stopped the run before any provider work began. > - This pull request uses the remote execution target's working directory. ## Linked Issues or Issue Description Refs #13094, which introduced this probe. Related: #8997 and #10419 address other workspace-directory handoffs; this patch fixes the newer Git-context probe. **What happened?** Remote runs failed with `setup_failed: Could not read execution-target Git context`, including tasks that did not use Git. The provider shell could not enter the controller's workspace directory. **Expected behavior** Startup must inspect Git and network context in the realized remote workspace. **Steps to reproduce** 1. Select a remote sandbox whose workspace is `/home/daytona/paperclip-workspace`. 2. Start a task whose controller workspace is under `/paperclip/instances/default/workspaces/`. 3. The startup probe exits when the controller directory does not exist in the sandbox. **Paperclip version or commit** Reproduced on `622376e99` and in the regression test before this patch. **Deployment mode** Docker controller with a remote Daytona sandbox. The same helper also serves SSH targets. ## What Changed - Use `remote.remoteCwd` for the remote Git-context probe. - Cover a missing controller directory in both credential modes. - Verify that SSH reads Git metadata from the remote workspace even when the caller directory exists. ## Verification - Before the fix, both new missing-directory tests failed with the reported error. - The launcher environment suite passes: 18 tests. - The adapter-utils suite passes: 1,085 passed, 11 skipped. - Adapter-utils typecheck passes. - A disposable sandbox with the affected deployment's image reproduced the old `cd` failure and passed with the corrected helper. The sandbox was deleted afterward. - Full repository typecheck and build pass locally. All CI gates pass at `61b522c`: regular test shards, serialized server suites, browser shards, Runner verification, build, policy, and security checks. - The first CI Build attempt hit a Runner suspension-acknowledgment timeout outside the changed code. The affected 13-case recovery group passes locally with its Rust fixtures built, and the unchanged CI job passed on its single retry. - The full serial local test run was stopped in favor of the complete CI matrix; it is not claimed as a local pass. ## Risks The probe now uses the provider-resolved target directory for sandbox and SSH execution. Local execution keeps its existing directory. There are no schema or API changes. Existing local credential and Git metadata tests pass. ## Model Used OpenAI GPT-6 in Codex, with tool-assisted code analysis, implementation, and live and automated testing. The exact model identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
82f662656a |
fix(runner): restore legacy Git access and independent networking (#13094)
## Thinking Path
> - Paperclip runs agents for people with different GitHub accounts.
> - Managed operations must use the intended person's eligible
connection.
> - A failed duplicate connection must not hide a healthy grant for the
same account.
> - Legacy hosts also need their existing Git configuration when managed
access is not configured.
> - Runner networking and local Git operations must not depend on GitHub
broker availability.
> - This pull request separates those policies and improves failure
diagnostics.
## Linked Issues or Issue Description
**What happened?** New runs always cleared host Git credentials and
installed managed launchers. Network permission depended on GitHub
environment variables. A launcher failure could stop even local `git
status`. A newer unhealthy duplicate could take precedence over a
healthy connection, and generic health errors were shown as reconnect
requirements.
**Expected behavior:** Use a healthy eligible managed connection for the
intended account. Preserve host authentication only for unconfigured
standard-trust local or SSH execution. Permit local Git during broker
failures and keep network permission independent of GitHub credentials.
**Steps to reproduce:** Configure healthy and unhealthy grants for one
GitHub account, dispatch an agent, and execute Git commands. Separately
run an unconfigured legacy host with existing GitHub CLI authentication.
Stop the broker and run local `git status`.
**Paperclip version or commit:** Master at
|
||
|
|
35fdc0c66b |
fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e200104727 |
feat: review connection actions from tasks (#13063)
Bring governed connection reviews into task history and composer approvals. Share resolution with Connections, add scoped remembered permissions, and resume agents through durable outcome receipts. Keep cards compact, collapse raw results, isolate untrusted provider output, bound continuation payloads, and reconcile missed live events. Add Storybook coverage, browser journeys, and service regression tests. Verification: all PR CI gates passed, Greptile 5/5, security scans passed, five connection-review browser journeys passed, and real native Codex approval/continuation was verified against the local MCP fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2043e0c735 |
fix: repair runner configuration, macOS execution, and artifact galleries (#13062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters select a provider, a model, and a runtime. > - Runner conversion rejected existing Claude agents. The model list mixed providers. > - The native Claude runner rejected custom models and could not launch on macOS. > - This pull request fixes conversion, model selection, and verified macOS execution. > - It also groups configuration fields consistently across adapters and opens artifact images in the task gallery. > - Operators can change an agent configuration and run the selected model on their Mac. ## Linked Issues or Issue Description **What happened?** Converting an existing Claude agent to Paperclip Runner failed with a Codex-only restriction. ACPX Claude showed unrelated models and required `claude-sonnet-5`. Its native runtime rejected macOS. Configuration mixed common model settings with process controls. Artifact cards labeled “Open gallery” navigated to attachment URLs instead of opening the task gallery. **Expected behavior** Conversion keeps agent identity and compatible settings. ACPX Claude uses the normal Claude catalog and accepts typed model IDs. Codex uses the native runner. The verified Claude runtime can launch on macOS ARM64 and x64. Common configuration sections place the same fields together across adapters. Artifact images open in the shared task gallery with navigation and downloads. **Steps to reproduce** 1. Open the configuration of an existing Claude agent. 2. Convert it to Paperclip Runner. 3. Select ACPX Claude and a different catalog model or a typed model ID. 4. Save the agent and run a disposable task on macOS. 5. Inspect configuration and advanced run-policy controls across adapters. **Paperclip version or commit** The bugs were reproduced on `165ca56a22adb60e5fda56045442d9c8498116a8`. This branch was rebased onto `7ed122911`. **Deployment mode** Built from source. Local test-drive instance on macOS ARM64 with an isolated database. Related work: #11798 addresses unsupported ACP session options in the existing adapter path. #13048 addresses working-folder preservation. This change fixes native runner configuration and launch behavior. ## What Changed - Remove the Codex-only conversion restriction. Preserve agent identity, instructions, directories, credentials, and compatible model settings. Reset incompatible sessions while retaining history. - Show ACPX Claude and native Codex as distinct provider choices. Remove ACPX Codex from advertised configuration. Normalize legacy configurations before fresh runs without rewriting historical run descriptors. - Select model catalogs and cache entries by provider. Support refresh and typed model IDs. Pass exact Claude IDs through session creation, model changes, and recovery. - Add verified macOS ARM64 and x64 Claude SDK snapshots. Bound executable allocation and total snapshot size. Preserve package checks, dependency isolation, process ownership, cancellation, and Linux descriptor loading. - Probe local runtime readiness. Report remote platform checks as incomplete until the remote runner verifies its runtime. - Surface actual model rejection and allow correction and retry. - Repair missing ACPX goal-capability helpers exposed by the post-rebase live test. Persist and restore the optional capability without breaking session startup. - Put Agent identity first and intentionally remove the Capabilities editor, as requested. This is removal of UI editing, not relocation: preserve existing capability metadata and API compatibility without adding another editor. Use the themed select for configurable permission modes, with normal text instead of monospace. - Put model and provider under Adapter. Give environment variables their own section. Fold command and arguments under Configuration. Fold lifecycle, timeout, and interrupt grace under Advanced Run Policy. Hide single-option permission controls. - Open image and video artifact cards in the existing task gallery, including cards in the artifacts panel. Chat attachment images use the same gallery. Preserve standalone media previews and download links. ## Verification - Rebased focused UI/API/database suites: 293 tests passed. - Rebased native runtime and ACPX suites: 242 passed, 7 skipped. - Repository typecheck, build, and token gates passed for the runner changes. Gallery follow-up UI typecheck, build, and token gates also passed. - Follow-up UI suites passed (86 tests), packaging checks passed (14 tests), and the final focused runtime suites passed (126 passed, 7 skipped). - Linux container isolation and lifecycle fixtures passed before rebase (57 passed, 2 skipped). Rust ACPX provider-session tests passed after rebase (8 tests). - Browser tests completed actual Claude and native Codex tasks on macOS ARM64. They covered conversion, catalog refresh, a non-default catalog model, a typed `haiku` ID, save/reload, cancel, follow-up session continuity, invalid-model errors, and recovery. - Final-revision live tests completed a typed Claude task, a follow-up with the same provider session, and a native Codex task on macOS ARM64. - Browser tests confirmed the moved interrupt-grace field saves and survives reload. Cross-adapter tests cover Claude, Codex, Gemini, process, gateway, and schema forms. - Full local run: 7,080 passed, 30 skipped, and two timeouts. Both timeout suites passed on isolated rerun (84 tests); the failures were the plugin login-worker exit diagnostic and the runner real-server vertical slice. - Final follow-up checks: 50 registry tests and 45 snapshot/installation tests passed (6 platform-specific skips). Oversized executable rejection is covered before allocation or reading; unsupported-platform tests invoke the real installation probe. - Runner head `ddb5101c483a297f74875ab96b3c66035b002d50`: all CI gates green, including full runner verification, repository build, typecheck, general/serialized server suites, browser tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34286178670). - Greptile: 5/5 on that runner head. All four review threads resolved. Superagent, Socket, and Snyk checks green. - After snapshot hardening, another real Claude task completed on this Mac using the rebuilt runtime. - Gallery follow-up: 148 focused tests passed, covering artifact selection, shared attachment collections, deduplication, image/video cards, standalone previews, downloads, and closing. Live browser verification completed on the settings follow-up: artifact selection, 6-image pagination with wrapping, download action, and closing all stayed on the same task URL. All checks passed on gallery head `96136da58ff195bf6ca00b281eb3022ad12d7bd8`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34287987536). Greptile returned 5/5 on that exact head with no unresolved threads. - Final settings polish: 96 focused tests, UI typecheck/build, and token gates passed. A real browser walkthrough verified readable permission options, identity placement, Capabilities removal, and permission save/reload. Original test-agent permission mode restored. All 31 checks passed on final head `e46540d6bf32bfb0566dca16b2f4a75ba437618c`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34292797886). Greptile returned 5/5 with no unresolved threads. ## Risks - Capabilities intentionally has no editable UI field after this change. Existing values remain readable and API-compatible; removing the field does not erase stored metadata. - macOS launch now copies verified package files into private snapshots. The implementation must retain isolation and clean up snapshots on exit. - Runtime provider or model changes reset the current session. Historical runs remain available. - The macOS x64 SDK executable digest was verified, but a live Intel Mac run was not available. Linux verification used container fixtures, not a real Claude task. - Remote environment tests report a warning when only the platform has been checked. They do not claim package readiness from the server host. ## Model Used OpenAI Codex, based on GPT-6. The exact served model identifier and context-window limit are not exposed in this session. Used reasoning, repository inspection, code execution, Rust and TypeScript tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites and both timeout suites on rerun; full-run counts above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6019e2bd6e |
feat(adapter-utils): carry binary bodies and attachment routes over the HTTP/2 sandbox bridge (#12923)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents in local and remote sandboxes through adapter utilities > - The HTTP/2 sandbox bridge decoded every body as UTF-8 text and rejected non-JSON content > - This stopped agents from uploading or downloading issue attachments through that bridge > - This pull request carries raw bytes, permits the two attachment routes, and enforces a shared body limit > - The benefit is correct attachment transfer with a process-wide memory guard ## Linked Issues or Issue Description **What existing behavior does this improve?** The HTTP/2 sandbox bridge forwards request bodies between an agent sandbox and the Paperclip host. It now supports binary bodies and the issue attachment routes. **Current behavior** The bridge decodes each body as UTF-8 text. It returns HTTP 415 for content types outside the JSON route list. An agent cannot upload or download an issue attachment through this transport. **Proposed behavior** The bridge carries raw bytes through the forward path. It permits the attachment upload and content routes. The queue transport and file gateway keep their existing route behavior. A shared 10 MiB body limit and process-wide byte reservation protect memory use. **Reason and benefit** Attachment clients need byte-preserving transfer. The shared limit keeps the gateway and host aligned. The reservation prevents concurrent streams from exceeding the accepted process memory ceiling. **Breaking changes** The HTTP/2 bridge accepts two attachment routes and permits binary content. The queue transport and file gateway keep their previous route lists and HTTP 415 behavior. No schema or external endpoint changes. ## What Changed - Carry request and response bodies as raw bytes through the HTTP/2 bridge. - Permit attachment upload and attachment content routes on the HTTP/2 bridge only. - Raise the resolved per-body limit to 10 MiB and share it between the gateway and host. - Reserve body bytes before allocation and release each stream reservation on every terminal path. - Document the body limit, process ceiling, and reservation behavior. ## Verification - Run `pnpm exec vitest run packages/adapter-utils/src/http2-bridge-server.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-callback-bridge.test.ts`; 226 tests pass. - Run `pnpm --filter @paperclipai/adapter-utils typecheck`; it passes. - Run the direct server TypeScript check with `tsc --noEmit` in `server/`; it passes with zero errors. - Verify multipart upload and binary download round trips over HTTP/2 without corruption. - Verify the queue transport and file gateway return HTTP 415 for the same routes. - Verify the host rejects bodies over the resolved limit. - Verify a denied reservation returns HTTP 503 and allocates no copy. - Verify stream cleanup releases reservations after completion, error, abort, timeout, and close. ## Risks The bridge now accepts larger bodies and binary content. The process-wide reservation limits total live body bytes to 1 GiB. Route behavior changes only for the HTTP/2 bridge. The security review found no blocking issue for this commit range. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code execution. The runtime did not expose the context window size. No model-generated code changes were made for this pull request. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ad494aacd |
fix(adapter-utils): preserve legacy sandbox PATH with managed GitHub (#13051)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Remote agents need both their installed tools and managed GitHub credentials. > - The GitHub launcher replaces a missing remote PATH with a small system path. > - Legacy images install agent CLIs outside that path, so those agents cannot start. > - This pull request preserves the remote toolchain and puts managed GitHub commands first. > - Command checks now use the same environment as execution. ## Linked Issues or Issue Description Refs #13005. Related: #10239 fixes a separate Cursor environment path issue. Searched open issues and PRs for sandbox PATH and GitHub launcher changes. No duplicate of this launcher fix was found. **What happened?** A remote Claude or Codex run passes command discovery, then fails with `command not found` and exit code 127. Managed GitHub launchers use only their own directory and `/usr/local/bin:/usr/bin:/bin`. This drops NVM and other toolchain directories from the sandbox path. **Expected behavior** Agent CLIs remain available on legacy and current sandbox images. Managed `git` and `gh` still resolve first and use the responsible person's credentials. **Steps to reproduce** 1. Use a sandbox whose agent CLI is installed in an NVM or other non-system bin directory. 2. Start an agent run with managed GitHub launchers and no explicit PATH override. 3. Observe that command discovery succeeds but the agent command exits with code 127. **Paperclip version or commit** Observed on `b97101893f0926f57ed0ce9ef1f8d3e4780c62c2`. The same launcher behavior remains on the base commit `be6bb768b`. **Deployment mode** Hosted server with remote sandbox execution. The shared launcher also supports SSH targets. ## What Changed - Read the remote target's effective PATH when no remote override is set. Do not copy an inherited controller PATH. - Prepend the managed launcher directory and retain the combined path in shell startup files. - Stop startup if path discovery fails. Frame the response so login banners cannot contaminate PATH. - Pass the sanitized launch environment to sandbox command checks, installation, and the second check. - Add real shell tests for legacy and current CLI layouts, quoted paths, managed GitHub command execution, explicit overrides, SSH, and failure cases. - Document the remote path contract. ## Verification - Four focused adapter utility suites passed: 153 tests. - The regression suite passed: 13 tests, including the Linux stdin handling fix. - `pnpm --filter @paperclipai/adapter-utils typecheck` passed. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. - The regression suite fails on the unchanged base revision (12 failures, 1 pass) and passes with this change (13 passes). The baseline ran in an isolated scratch copy. - Full local test coverage was attempted using the official CI shards. The run was stopped after macOS Postgres shared-memory exhaustion and CLI timeouts under load. The affected server database suite passed in isolation (31 tests), as did the five affected DB/CLI suites (89 tests). - [Full Linux CI](https://github.com/paperclipai/paperclip/actions/runs/34258333833) passed on `4e1426f5b`: all general and serialized test shards, all browser shards, typecheck and release registry checks, native runner verification, application build, and release canary dry run. - Greptile scored the latest commit 5/5 with no unresolved findings. - Shell tests use isolated local fixtures and make no provider or model requests. No live sandbox qualification is claimed. ## Risks - Remote startup adds one bounded path query when no explicit override exists. A failed query stops startup. - Explicit remote path overrides still control which tools are available. Invalid overrides now fail the command check earlier. - Credential selection and GitHub broker policy are unchanged. Tests verify managed wrappers stay first and can invoke underlying commands. - No database migration or sandbox image replacement is required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, and terminal tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
297d8741f5 |
fix: resolve duplicate connections to the same GitHub account (#13022)
## Thinking Path > - Paperclip lets people share agents while keeping GitHub access personal. > - Each managed Git or GitHub operation selects an eligible connection grant. > - Connecting the same GitHub account twice creates two grants. > - The old resolver counted grants and rejected them as competing identities. > - Managed commands then ran anonymously and reported a misleading login failure. > - This change compares GitHub account IDs and selects one eligible grant for the same account. > - The benefit is reliable access after reconnecting, with clear diagnostics for real failures. ## Linked Issues or Issue Description Refs #13005. **What happened?** Two active connections owned by one Paperclip user pointed to the same GitHub account. Managed Git refused both as ambiguous. The agent could not push, although the account was connected and had repository access. **Expected behavior** Multiple grants for the same GitHub account resolve to one eligible authorization. Different accounts remain ambiguous. Unavailable access explains its cause without blocking unrelated work. **Steps to reproduce** 1. Connect the same GitHub account twice for one Paperclip user and allow the shared agent through both connection audiences. 2. Start an instruction as that user. 3. Run managed gh or git push. Before this fix, no credential is provided. ## What Changed - Compare stable GitHub account IDs when more than one eligible grant exists. Never deduplicate by login alone. - Prefer an available grant, then the newest authorization with a stable ID tie-breaker. Refresh and webhook timestamps do not change the selection. - Keep the selected credential and connection policy together. Do not combine permissions or fall back from a dedicated account to a personal account. - Print the redacted unavailable reason in managed command output. Unrelated local operations still work anonymously. - Add database and executable launcher regressions, and document selection behavior. ## Verification - Final `pnpm -r typecheck` and `pnpm build` passed. - Fourteen operation credential integration tests passed, covering duplicate personal/dedicated grants, incomplete credentials, distinct accounts with the same login, missing identity metadata, revocation, membership, connection audiences, and A → B → A steering. Existing Git credential and gateway suites and both executable launcher tests also passed. - The local broad test run encountered three embedded-Postgres lifecycle timeouts and stale modules from edits made during that run. A fresh process rerun of all four affected suites passed all 35 tests. The full Node 24 CI test matrix passed on the final commit. - CI passed all 31 checks on `797973b30beb16ba5fa69ed281835e1ab812b449` (Storybook visual regression was correctly skipped). An unrelated Company Settings UI test failed once; the focused local reproduction and rerun of its CI shard both passed without code changes. - Fresh Greptile review of the final commit: 5/5, with no open findings. Security checks passed. - Live acceptance passed with both duplicate connections enabled: managed `gh api user` returned the expected account, managed `git push` succeeded, and the agent created #13023 and pushed its review fixes. No host login or credential changes were used. - Applied the final source/compiled patch to the affected instance with backups, after confirming no runs were active. Restarted service health and the final resolver selection were verified. The patch is an overlay on the existing deployment; this PR supplies the upstream fix. ## Risks The resolver selects one authorization for an already permitted GitHub account. It does not combine repository permissions across connections. If the selected authorization has narrower access, that operation can still be denied by GitHub. Different provider account IDs and unknown duplicate identities continue to fail closed. No schema, host credential, or connection permission changes are included. ## Model Used OpenAI GPT-6 through Codex assisted implementation and verification with shell, database, and browser tools. The exact model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1cc45086d3 |
feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Several people can send instructions to the same agent and task. > - A fixed GitHub token in the provider process can keep the first person's access after another person's message is accepted. > - Task ownership cannot select credentials for each accepted instruction or preserve the identity of an operation already in progress. > - This pull request records ordered execution identity contexts and resolves credentials when managed Git, gh, or GitHub tools start. > - The benefit is automatic personal GitHub access for shared agents, with durable continuation rules and no teammate credential fallback. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: orchestration, connection grants, database, runtime adapters, native runners, and run details. **Problem or motivation** A shared agent must use the person whose instructions it has accepted. A queued message must retain its author. A retry or approval without new instructions must retain the originating identity. GitHub must remain optional for ordinary work. **Proposed solution** Persist execution identity separately from task ownership. Give new processes a run-scoped broker capability and token-free managed launchers. Capture identity at operation start. Keep an explicit dedicated-agent grant as an override. Show redacted diagnostics in run details. **Alternatives considered** Per-task ownership, fixed provider tokens, and mutable repository author configuration do not handle accepted steering or concurrent operations. A manual account-selection action would add unnecessary setup to each turn. **Roadmap alignment** This completes the existing Multiple Human Users, MCP Tool Gateway & Apps, Secrets Manager, and Self-healing Runs capabilities. The implementation follows the maintainer-approved plan. Related work: Refs #12843, Refs #12907. Existing proposals #4618 and #8945 cover per-agent or per-worktree author configuration. This change instead follows the accepted human instruction across runtime types. Refs #11831 for governed personal connection delegation; this change preserves connection audience checks and does not use standing delegation as a personal credential fallback. ## What Changed - Add durable, ordered identity contexts and active run references. Preserve message authors through consolidation, steering, retries, delegation, approvals, routines, and restart. - Add an authenticated operation-time GitHub credential broker and local/remote managed git and gh launchers. Keep personal tokens out of the long-lived provider process. - Resolve GitHub gateway and server-side Git operations through the same responsible-person or dedicated-grant selection rules. - Make absent and unavailable GitHub credentials non-blocking at generic startup. Clear host and prior-person credentials. Keep anonymous Git access where supported. - Add run-detail identity history and the dedicated-account warning. Keep task ownership and queue-versus-steer decisions unchanged. - Preserve personal OAuth declarations through connection edits. Retain exact selected grants in the gateway. - Fix continuation races found during real acceptance: verify a warm owner before credential rotation, and wait for bounded durable runner suspension before the next run starts. - Make migrations replay-safe. Retain identity through agent/run deletion, remove it with its company, and clean terminal launcher directories before releasing execution environments. Document coordinated release and rollback. ## Verification - Full workspace typecheck, build, and token gates passed. The complete local suite passed in its normal test groups: 17,120 passing tests, including all 143 serialized server suites. After integrating the newly merged runner API work, full local typecheck and build passed again, along with 890 focused integration tests. All 31 checks on the integrated revision passed, including build, browser E2E, release registry, canary dry run, typecheck, security and all test suites. Greptile is 5/5 with all review threads resolved. - Current focused checks passed: 142 native executor tests, 67 runtime lifecycle tests, 9 durable identity tests, 75 credential/routine tests, 19 low-trust/resumption tests, and the executable migration replay test. - Authenticated browser acceptance with two Paperclip users and two GitHub accounts on one shared native agent passed. Real commits and pushes followed A → B accepted steering → queued A continuation in the same saved conversation. GitHub commit author and committer identities matched all three operations. Both runs succeeded and task ownership stayed unchanged. - Real GitHub MCP calls switched from A to B after accepted steering. A delegated subtask retained its originating identity across a server restart. - Disabling B's GitHub connection left ordinary work successful. Managed gh was unauthenticated and the provider had no inherited GH_TOKEN or GITHUB_TOKEN. - The browser displayed run-detail diagnostics and the exact dedicated-account warning. A final controller-restart check followed by another-person continuation retained the conversation, selected the correct GitHub login and Git author, and removed each terminal launcher directory. - Company-lifetime migration and all five previously failing CI suites passed locally (167 tests). Same-token gateway A → B → A and six broker/launcher boundary tests passed. - Remote callback, launcher, sandbox, and runtime contract tests passed. Both native and legacy Codex completed actual Daytona executions on the integrated revision ([campaign results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)). The remote package-manager shim staging regression also passed locally. ## Risks - Deploy the migrations, server broker, launchers, and runner artifacts together. Existing processes finish with their original contract. New managed processes need the broker endpoint for GitHub operations. - Finish or stop new managed executions before rolling application code back. Keep the additive schema and identity history during rollback. - Scripts that require a persistent raw GH_TOKEN must use managed git, gh, or GitHub gateway tools. Run capabilities authorize code executing within that run to acquire its current identity; this is not hostile-code isolation within one execution principal. Managed commands prevent automatic credential carryover; arbitrary code deliberately copying a credential is outside that boundary. - Uncertain steering acknowledgement deliberately holds new credential acquisition until reconciliation. Already-started operations retain their captured identity. - GitHub private access and provider outages can still fail the specific operation that needs them. Dedicated grant failure does not fall back to personal access. ## Model Used OpenAI GPT-6 through Codex assisted implementation, review, shell execution, and browser acceptance. The exact model variant and context-window size are not exposed in this session. Tool use included TypeScript and Rust tests, database integration tests, GitHub CLI, and authenticated browser control. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ee2a190626 |
Unify Paperclip Runner experimental controls (#12666)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is an experimental execution adapter. > - The adapter and its required sandbox ingress had separate settings. > - A user could enable one setting and still have an unusable runner configuration. > - The runtime already makes one durable native or legacy decision for each run. > - This pull request uses that runtime decision for ingress authorization. > - The benefit is one clear opt-in with safe recovery for existing native runs. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the experimental settings and transport authorization for Paperclip Runner. **Subsystem affected** Cross-cutting. This change affects the React settings UI, shared settings contracts, adapter utilities, and server runtime selection. **Current behavior** Settings shows separate Paperclip Runner and Runner Preview Ingress controls. A user can enable the runner but leave required sandbox ingress disabled. **Proposed behavior** Settings shows only Paperclip Runner. Its native runtime decision also authorizes provider WebSocket ingress when the execution target requires it. A persisted native run keeps its recovery transport after the setting is disabled. **Reason and benefit** Paperclip Runner is one experimental capability. One opt-in removes an invalid partial configuration and makes the rollout boundary easier to understand. **Breaking changes** The Runner Preview Ingress card is removed. The old `enableRunnerPreviewIngress` key remains accepted in stored settings and managed configuration, but it has no server runtime effect. The public adapter-utils input remains compatible through a deprecated alias. **Additional context** Refs: #12638, #12641, #12656. ## What Changed - Removed the separate Runner Preview Ingress card from Experimental Settings. - Made resolved native runtime selection authorize required provider ingress. - Preserved ingress recovery for persisted native runs after the rollout flag is disabled. - Kept the old settings key and adapter-utils input as deprecated compatibility contracts. - Added focused UI, runtime policy, transport, stored-settings, and managed-config regression tests. - Updated deployment documentation and feature descriptions. ## Verification - GitHub Actions will run typecheck, tests, build, policy, and browser shards. - Focused tests cover the single settings control, runtime authorization, fail-closed transport selection, the deprecated public input, and old managed configuration. - No local tests were run, per the maintainer request to use GitHub Actions for verification. - `git diff --check` passes. ## Risks Low to moderate risk. The effective ingress gate changes from a separate stored flag to the resolved native run decision. Fresh runs still require `enableNativeRunner`. Persisted native runs remain recoverable. Legacy adapters never receive ingress authorization. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |