mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
codex/native-completion-current-master-baseline
205
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d9b64ee28e |
fix(codex-local): order Codex models the way the ChatGPT app does (stacked on #14917) (#14918)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each agent runs on an adapter, and the operator picks the agent's model from the list the adapter advertises. > - The `codex_local` adapter advertises a static, curated list. The server returns that list as written, but the list itself is not in a useful order: the default `gpt-5.6-sol` sits above the whole GPT-6 family (#14878). > - This pull request orders the list the way the ChatGPT app orders Codex models: newest model version first, then by decreasing capability inside each version, with older models at the end. > - It is stacked on #14917, which makes the model dropdown show a hand-ordered list as the adapter advertises it. Without that change the picker shows every list alphabetically. > - The benefit is that a user who knows Codex finds the right model at once, and older models sit at the end of the list. ## Linked Issues or Issue Description Fixes #14878. Related: #14917 (the Claude counterpart, #14877) carries the dropdown change this PR relies on. This PR contains that commit until #14917 lands; after that it rebases to the Codex commit alone. ## What Changed - `packages/adapters/codex-local/src/index.ts`: reorder the advertised `models` list. GPT-6 (`astra`, `sol`, `luna`) first, then GPT-5.6 (`sol`, `terra`, `luna`), then `gpt-5.5`, `gpt-5.4`, `gpt-5.4-mini`, `gpt-5`, `gpt-5-mini`, `gpt-5-nano`, then the o-series and `codex-mini-latest` in their existing relative order. `DEFAULT_CODEX_LOCAL_MODEL` is unchanged. - `packages/adapters/codex-local/src/index.test.ts`: the metadata test now asserts the full order of the GPT entries. The ChatGPT app also lists GPT-6.1 Sol first. That model is not in Paperclip's list today. Adding a model needs reasoning-effort and fast-mode entries as well, so it is out of scope for an ordering fix. ## Verification Run from the repository root: ```sh pnpm exec vitest run packages/adapters/codex-local server/src/__tests__/adapter-models.test.ts pnpm --filter @paperclipai/adapter-codex-local typecheck ``` Red on `master`, green here: the updated metadata test fails against the unmodified list because `gpt-5.6-sol` comes first. Manual check: open an agent that uses `codex_local` and open the model dropdown. The list reads gpt-6-astra, gpt-6-sol, gpt-6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5, gpt-5.4, gpt-5.4-mini, gpt-5, gpt-5-mini, gpt-5-nano, o3, o4-mini, o3-mini, Codex Mini. ## Risks - Low risk. The list content and the default model do not change; only the order does. The server returns this list as-is for the built-in Codex adapter and the server test compares against the exported list, so no server change is needed. - The first entry is no longer the default model. Nothing reads the first entry as a default: `DEFAULT_CODEX_LOCAL_MODEL` is resolved separately. ## Model Used Anthropic Claude Fable 5.1 (`claude-fable-5-1`) through Claude Code, extended thinking on, with tool use for reading the repository, running vitest and tsc, and editing files. The account holder reviewed the change and owns the commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: abderbj <115119179+abderbj@users.noreply.github.com> |
||
|
|
2ec82c5774 |
fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use tools through company-scoped connections and provider sessions. > - Resolving a tool connection currently forces a fresh session even when the provider can load new tools into the existing conversation. > - A fresh provider conversation can receive too little history to continue the task. > - This pull request adds explicit tool-refresh capabilities and uses them in both runner paths. > - Fresh attempts receive bounded task history with source IDs and retrieval instructions. > - The benefit is that agents can continue the same task after a connection changes. ## Linked Issues or Issue Description **What happened?** A resolved tool connection forced a fresh provider conversation. The new conversation could lose the original goal and prior answers. Claude also rejected resume when only the MCP server set changed. **Expected behavior** Resume the provider conversation when its harness can refresh tools. When a fresh session is required, supply enough bounded history to continue the task. Preserve company, agent, task, workspace, model, instruction, and skill checks. **Steps to reproduce** 1. Start a conversation and agree on a task and its constraints. 2. Request and connect a tool needed for the task. 3. Continue the conversation after the connection resolves. 4. Check that the agent remembers the task and can use the new tool. **Paperclip version or commit** The bug was reproduced on master at `c46e41e81`. This branch is rebased on current master. **Deployment mode** Self-hosted server. Both legacy adapters and the native runner are affected. Related public work: Refs #13282 for task-backed conversations. Refs #13057 for the broader session-compaction proposal. Refs #14659 for another report about local CLI session continuity. This change fixes tool-connection continuation. Provider authentication repairs keep their existing recovery behavior. ## What Changed - Expose tool-refresh support in native harness descriptors and legacy adapter metadata. - Request tool refresh after connection resolution. Keep provider authentication repair as a fresh-session wake. - Reload current tools and credentials while retaining supported Claude, Codex, Grok, and other provider conversations. - Allow MCP-only changes during qualified native recovery. Keep all other compatibility checks. - Refresh managed-provider and ACPX tool bindings when attaching a new run. - Add a fresh-session handoff for both runner paths. Bound database reads, excerpts, and the final packet to 24,000 bytes. - Include the original request, recent messages, decisions, plans, prior answers, and source IDs. Mark omitted content. Apply reset boundaries, wake cutoffs, quarantine, and secret redaction. - Add regression tests and document the capabilities and handoff behavior. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed. - Final review fixes passed 314 server tests, 333 adapter utility tests, 139 native-session runtime tests, and 12 managed-provider Rust tests. They verify historical quarantine, raised budgets across attachment, no history reads on successful resume, and handoff delivery on fresh retry. - Broader branch verification also passed 1,401 adapter utility tests, 1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core tests. - Live Claude CLI and Grok ACP probes preserved the provider session ID, recalled a prior task constraint, and called a newly added read-only MCP tool. - GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55 successful checks and 4 skipped checks. This includes all test shards, all eight browser shards, runner checks, and the Grok clean public npm install canary. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37057514976). - A full local test attempt encountered a separate Git snapshot timeout. All affected local suites passed after the final edits, and the full CI test gates passed. - Review the capability matrix in `packages/paperclip-runner/README.md`. Repeat the four reproduction steps with a supported provider and with an unsupported harness. ## Risks - Provider tool refresh can fail. Existing recovery falls back to a fresh conversation where policy permits it. - A new transport can replace an old process while preserving the provider conversation. Tests cover current credentials and unchanged identity. - Long history can omit older context. Explicit markers and source IDs let the agent retrieve needed context within task scope. - Unknown and unqualified harnesses use the fresh-session path. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live provider testing. The exact model ID and context-window size are not exposed in this session. Claude and Grok also ran as test subjects. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d432dc7fa3 |
Add GitHub-synced skill sources (#14713)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company skills supply instructions and files to those agents. > - GitHub imports already exist, but users cannot manage repositories as skill sources. > - Repository refresh also needs caller-authorized access and complete local packages. > - This pull request adds Sources inside Skills and reuses GitHub connections from Apps. > - Installed snapshots let agents use skills without fetching GitHub during a run. > - Manual refresh preserves skill identity and leaves failed imports on their last good version. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: skills UI, server, database, shared contracts, and runtime materialization. **Problem or motivation** Users keep skills in GitHub repositories. They need a clear way to select, import, and refresh those skills. Existing imports do not expose repository management or consistently preserve supporting files. **Proposed solution** Add company-scoped skill sources. Browse repositories from all accessible GitHub connections, or paste a public repository or branch URL. Select whole skill packages, inspect included files and reference warnings, and install complete, immutable snapshots. Refresh each source manually. **Alternatives considered** Project repository settings hide the workflow from Skills. A second GitHub connector would duplicate credentials and grants. Upstream editing and PR creation are separate work. **Roadmap alignment** This implements the Skills Manager direction in ROADMAP.md. The maintainer requested this scope and reviewed the component and full-app journey stories before implementation. Related reports: Refs #10285, Refs #10949, Refs #13464. Related work: #14356, #13656, #9268. ## What Changed - Add source and entry records, an idempotent migration, company-scoped APIs, and legacy GitHub import adoption. - Reuse current caller grants and credential refresh. Combine and deduplicate repository inventories across accessible connections. Pasted public URLs also prefer the active user’s authorized connections. Tokens stay in the Git child environment, never argv or disk. - Fetch a shallow Git snapshot at one immutable commit. Scan the full local tree, including hidden and nested folders. Read Git objects without checkout or archive transformations and enforce nested package boundaries. - Bound Git downloads to 128 MiB and three minutes. Cancel active process groups and remove incomplete downloads. Preserve cancellation and deadlines while progress drains; close stalled HTTP progress streams after 30 seconds. Reuse caller-scoped temporary snapshots for preview/import after reauthorization. - Index repository package boundaries once and cap expanded work at 1,000 packages, 10,000 files, and 100 MiB, including repeated copies of shared blobs. Bound path depth and the shared path index. Discovery keeps audited manifests without retaining all package bodies. - Resolve moving branches before fetching so unchanged discovery reuses caller-scoped snapshots. Limit active scans, scan frequency, and new downloads per caller and company; quotas apply before metadata reads and across connections, and cached scans do not consume the download quota. - Store complete versions with script content, binary bytes, and executable modes. Preserve these through copies, runtime caches, and runner packaging. - Stage downloads before publication. Use source leases, revision checks, and transactional activity records. Keep installed versions after failures, upstream deletion, deselection, and disconnect. - Add the approved import flow, Sources page, selection tree, provenance, read-only Studio behavior, and saved return from GitHub setup. - Add package manifests, commit-pinned file previews, and separate runtime requirements and reference warnings. Supporting files are included together; nested skills remain independently selectable. Preview requests reauthorize the caller and re-audit package content. - Show installed skills as compact links beneath each source. Repository titles open GitHub. Keep Refresh, Select skills, and Disconnect source in a three-dot menu. Source rows omit the branch, imported count, and refresh timestamp; action alignment and repository titles work at narrow widths. - Stream discovery metadata over an opt-in NDJSON response. Show measured Git download progress and real package/file counts, animate newly checked skills, support cancellation, and require a complete scan before selection. Keep the existing JSON API. - Retain component stories and add a separate full-app journey story group. Include fixed progress states and interactive scan, large-repository, interruption, and saving stories. - Update Skills documentation and product contracts. Suppress private GitHub skill references in telemetry. Privacy review requested for the telemetry changes. ## Verification - Local repository typecheck, full build, token gates, and Storybook build passed during this work. Focused transport, authorization, scanner, persistence, route, and UI tests pass. The final UI refinement passes all eight focused UI tests, UI typecheck/build, and token gates. The scanner resource and repeated-discovery fixes pass 132 focused scanner, transport, authorization, source-service, route, and rate-limit tests, plus server typecheck/build. Full-suite verification comes from CI; the older full local Vitest run was stopped after unrelated chat failures and a font-test failure, all of which passed in fresh focused runs. At commit `1098d5996`, all 54 active checks pass; two optional Storybook jobs are skipped. CI covers repository typecheck, build, the full test suites, browser shards, and the canary dry run. Greptile is 5/5 with no open findings; the security scan also passes. - Adversarial scanner tests verify repeated-blob byte accounting with and without declared sizes, package/file/path caps, one-time repository indexing, metadata-only discovery audits, and nested package boundaries. Additional tests cover branch movement, snapshot reuse, caller/company quotas, isolation across connections, active-lease cleanup, quota recovery, and rejection before any metadata API call. - Real Git tests verify hidden paths, exact binary bytes, executable modes, export-ignore preservation, symlink/submodule reporting, pinned commits, caller-scoped cache reuse, cancellation, cleanup, and credential isolation. Regression tests hold both download slots with permanently blocked progress callbacks, verify timeout/cancellation cleanup and retry, and exercise HTTP backpressure cancellation. Access tests cover automatic public-URL connection selection and revoked grants. Database tests verify company and grant audiences. - Live isolated browser test: the public `anthropics/skills` scan now completes and discovers all 20 skills without connecting an account. Imported canvas-design with all 83 files, opened it from Sources, and verified the installed binary-font preview/download control. Package previews also expose the complete file inventory before import. Cancelled an active Git download and retried successfully to all 20 discovered skills; the browser displayed measured download progress. The current audits reject four other packages; eligible selections remain importable. - Browser checks verify the simplified source rows at desktop and narrow widths, keyboard navigation into the actions menu, Refresh from the menu, selection, and fixture disconnect with installed skills retained. Storybook includes a menu-open checkpoint and a 320px layout. - Storybook includes receiving/preparing download checkpoints and a timed full-app import journey, plus cancellation, retry, large-repository, and saving states. Streaming tests cover split UTF-8 frames, incomplete streams, late responses, cross-company requests, HTTP errors, and JSON compatibility. - Earlier live acceptance on this PR imported `stitch-skill` with `DESIGN.md`, assigned it to an agent, disconnected its source, and ran a successful Studio test that read both installed files. An editable copy changed independently. Both Skills variants, mobile selection, and return from GitHub setup were exercised. - Private access, revoked credentials, OAuth success return, binary/script preservation, concurrent refresh, transaction rollback, version pins, and legacy adoption have automated coverage. A real private-repository OAuth grant was not created during this test. ## Risks - The migration groups recognizable legacy imports without provider calls. Their first successful refresh completes the local package snapshot. - Reference checks are advisory. They cover Markdown links and explicit relative resource paths, not arbitrary runtime dependency graphs. Preview text is capped at 64 KiB; imported bytes remain complete. - Git must be installed on the server. Shallow fetches still download the branch snapshot, including files outside selected packages. Downloads have size/time/concurrency limits. Temporary caches are bounded and caller-scoped. GitHub API quota still applies to repository metadata and the connection picker; content no longer uses per-file API requests. Failed scans retain installed content. - Sources depend on the current caller's GitHub access. A saved connection does not grant access to another person's token. - GitHub script support and immediate manual refresh are explicit maintainer-approved requirements. The operator trusts the selected repository and accepts upstream script and executable-mode changes on refresh. Static audits are not a sandbox or a guarantee of safe code; agents may later invoke installed helpers under their runtime permissions. Import and refresh do not execute scripts, hooks, package installation, or builds. Raw URL and skills.sh imports keep their prior script restrictions. - Source originals remain read-only. Refresh affects subsequent unpinned runs; explicit pins and active runs retain their versions. - The telemetry change removes source-managed GitHub identifiers from skill-reference events. It introduces no event or field. Please review the privacy boundary. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbc5132e6c |
fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters own local or remote provider processes. > - Run cleanup must know when the final provider process has stopped. > - A remote timeout or a lost transport does not prove that the provider stopped. > - Retried provider invocations also make an earlier stop signal stale. > - This pull request adds a verified final-invocation stop callback before workspace restoration. > - The dependent instruction revision change uses that boundary to preserve private instruction edits safely. ## Linked Issues or Issue Description **What happened?** Adapter completion did not expose a reliable point between provider shutdown and workspace restoration. Cleanup could lose provider-written files, or treat a remote timeout as proof that a process stopped. **Expected behavior** Cleanup runs once after the final provider invocation has a verified stop receipt and before workspace restoration. An incomplete remote command keeps collection pending. **Steps to reproduce** 1. Run a remote provider that returns a timeout without a numeric exit code. 2. Let the adapter return or retry the provider. 3. Attempt to collect provider-written files during cleanup. The adapter has no verified final-process boundary to use. This is the prerequisite for the stacked canonical instruction revision pull request. It has no database or UI dependency. Related #13291 verifies remote termination for later recovery; this change exposes the earlier adapter-owned stop boundary before workspace restoration. The scopes do not duplicate each other. ## What Changed - Add a stop callback to adapter execution context and a final-invocation fence. - Confirm local child closure and complete remote exit receipts. Reject remote timeouts, missing exits, transport failures, and SSH exit 255 as stop proof. - Invoke collection once before workspace restoration in eight CLI adapters and at the confirmed ACP stop boundary. - Preserve the stop observation when a later log flush fails. - Run bridge and workspace cleanup in `finally` even when collection rejects. ACP records a safe error without exposing a raw filesystem path. Grok keeps collection errors separate from workspace restore failures and preserves completed provider results when both cleanup steps fail. - Correct the existing Cursor test shell fixture so bounded remote file reads run against real fixture files. ## Verification - Three stop-boundary regressions failed before the callback implementation and passed after it. - Independent prerequisite branch: 377 tests passed across 24 adapter, process-target, and ACP suites. Two added collector-rejection tests failed before the cleanup fix and passed after it. - A third regression reproduced Grok misclassifying a collection failure as failed workspace restoration. Two additional cases covered completed and failed provider turns when collection and restore both fail. The Grok and restore-classifier suites passed 48 tests. - All nine affected package typechecks and affected package builds passed; Grok checks passed again after its classification fix. - Integrated instruction branch: native local, legacy local, and native Daytona each passed three browser tasks with exact persisted bytes, fresh-task readback, history/restore, and explicit conflict resolution. Unchanged warm Daytona passed three turns. Legacy Codex passed all five checks again after the exception-safe cleanup fix. - Alternate staging passed the same three-task native Daytona flow: exact stopped-run save, independent downloaded readback, browser history/restore, and explicit resolution of a real concurrent edit. The deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which covers the initial adapter callback. Later Grok cleanup failures are qualified by the adapter tests above. - The final combined native Daytona flow passed again on deployed `2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved ordinary instruction edits, independent readback, History/Restore, concurrent board conflict, and explicit candidate resolution. This native staging flow does not claim to exercise the Grok adapter. ## Risks - An unverified remote stop intentionally does not trigger collection. A later controller with verified stop evidence must recover it or report the copy unavailable. - The callback is optional. Callers that do not register it retain their existing behavior. - The callback runs before workspace restoration and can delay cleanup if its caller does not bound its own work. The dependent instruction collector uses bounded reads and retries. A rejected callback still permits bridge and workspace cleanup; it cannot claim an instruction save. ## Model Used OpenAI Codex, GPT-6, with tool use and code execution. The runtime does not expose the exact deployment variant or context-window size. Multiple Codex agents implemented and verified the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f8750627f |
fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path > - Paperclip runs agents and recovers failed tasks. > - Codex ACP can report usage exhaustion as a typed terminal failure. > - The shared engine removes provider text before it stores the result. > - Codex did not use the existing terminal classifier hook, so a quota failure became a generic turn failure. > - This change classifies explicit usage exhaustion before that text is removed. > - Recovery can then use quota backoff and the supported reset clock. ## Linked Issues or Issue Description **What happened?** A Codex ACP `limit` failure with explicit usage-exhaustion text produced `acpx_turn_failed`. Recovery could schedule the ordinary short retries because the quota classification and reset time were lost. **What did you expect to happen?** Keep the failure visible and use the existing provider-quota wait. Preserve a supported reset timestamp without storing provider text. **Steps to reproduce** Return a terminal ACP session failure with category `limit` and title `You've hit your usage limit for GPT-5. Switch to another model now, or try again at 4:30 PM (America/Chicago).` The real child-process regression tests exercise both pinned ACPX versions in persistent and oneshot modes. **Paperclip version** Base commit: `32573876d4`. Related work: #13651 and #13831 provide the shared hook and Claude classification. #11854 includes quota handling as part of optional credential rotation, but reads the already-sanitized result; this patch handles the typed terminal boundary without adding rotation. #13549 reads recovery text after this boundary and cannot recover discarded provider text. #9011 concerns the Codex CLI backoff. This change leaves those other mechanisms in place. ## What Changed - Register a Codex terminal-failure classifier with the existing ACP engine hook. - Recognize explicit usage exhaustion only in a typed `limit` failure. - Reuse the Codex reset-time parser and existing provider-quota recovery fields. - Leave context, turn, rate, budget, storage-capacity, and unknown failures on their existing paths. - Test real ACP children, both dependency patches, privacy, and recovery classification. Document the boundary. ## Verification - 88 focused tests passed across Codex ACP, parsing, and server recovery classification. - An initial cross-adapter run passed 53 tests, including the existing Claude quota suite. - Removing only the classifier registration makes five new integration tests fail; restoring it passes all 19 new tests. - `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `b10d60e002`, including all unit/integration shards, browser shards, typecheck/build, runner checks, and release/package gates. - The duplicate local `pnpm test:run` reported three skill-cache failures and two runner-suite failures. It is not claimed as a full local pass. - The three skill-cache failures reproduce in a clean worktree at the unchanged base commit (3 failed, 107 passed, 3 skipped across the skills and runner files). The cache rename reports `EACCES` on macOS. - An isolated runner-suite run reports two embedded PostgreSQL startup failures before its assertions (37 tests pass). No runner source is changed; the Linux CI runner and server suites passed. - Greptile reviewed the current head at 5/5 with no actionable findings or unresolved threads. ## Risks - Explicit usage-exhaustion failures now wait for quota recovery instead of short generic retries. - Unknown wording retains the existing behavior. The generic historical terminal-limit message alone cannot establish quota exhaustion. - Reset parsing keeps the existing Codex clock formats. A missing or unsupported reset uses the existing quota backoff. - No schema, credential, UI, dependency, or deployment changes. Provider text stays in memory and is absent from results and logs. ## Model Used OpenAI GPT-6 via Codex. Exact runtime model variant and context-window size were not exposed. Used reasoning, repository inspection, editing, and local tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
|
||
|
|
cdf04a33fa |
feat(adapters): refresh current coding models and reasoning controls (#13829)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its adapters supply model catalogs and reasoning controls to agent setup. > - Several provider releases are missing from the fallback catalogs. > - Some newer models also have effort levels that the UI does not offer. > - Operators need the exact supported IDs and controls when discovery is unavailable. > - This pull request updates the existing adapters from current provider documentation. > - Operators can select current coding models without entering custom IDs. ## Linked Issues or Issue Description **What existing behavior does this improve?** Model selection and reasoning controls across the existing coding-agent adapters. **Subsystem affected** Claude, Codex, Grok, Gemini, Cursor, Kimi, and OpenCode adapters; model discovery tests; agent creation and editing. **Current behavior** The catalogs omit Opus 5.5, GPT-6 Sol/Luna, Grok 4.7/4.6/4.5, current Gemini Flash models, and several Cursor/Kimi choices. Bedrock has obsolete IDs. The UI omits supported effort levels and saves Grok effort under a key the runtime does not read. **Proposed behavior** Offer verified current model IDs and model-specific efforts. Remove retired Gemini 2.0 choices. Keep configured defaults and saved model IDs. Keep runtime discovery for account-specific choices. **Reason and benefit** Catch up with provider releases through September 22, 2026. Correct the picker and runtime controls together. **Breaking changes** No database or API change. Gemini 2.0 options leave the picker after their June 1 shutdown. Existing saved IDs remain unchanged. Corrected Bedrock catalog IDs do not rewrite saved configuration. **Additional context** Supersedes the separate GPT-6 Sol PR #13830. Fable 5.1 was already merged in #12730, and GPT-6 Astra in #12851. The Grok 4.6/4.5 proposal #11324 was closed and parked by its author. This change retains the default-sentinel fix from #12062. Related discovery proposals #13127 and #13565 do not supply these catalog and effort updates. Searches found no open PR for the additional model IDs. See [the dated audit](https://github.com/paperclipai/paperclip/blob/feat/claude-opus-5-5/doc/adapter-model-audit-2026-09-22.md) for exact scope, primary sources, runtime observations, and account-specific limits. This updates existing adapters and does not duplicate planned core work. ## What Changed - Add Opus 5.5 for direct Claude and Bedrock, with a Claude Code 2.1.280 gate. Correct and extend Bedrock model IDs. - Add GPT-6 Sol/Luna and Fast mode. Offer Ultra for Astra/Sol and GPT-5.6 Sol/Terra, and Max for both Luna generations. - Add Grok 4.7/4.6/4.5, expose supported Extra High effort, and save Grok edits under `reasoningEffort`. - Add Gemini Flash 3.8/3.7/3.6/3.5, Flash Lite 3.5/3.1, and 3 Flash Preview. Remove retired 2.0 choices. - Add the current documented Cursor fallback models, including Fable 5.1, Composer 2.5, and Muse Spark 1.3. - Refresh OpenCode fallback IDs used in remote environments from its installed provider registry. - Add Kimi K3 256K. Update the existing coding alias to K2.8 Preview and enable its CLI effort settings. - Use model-specific Claude/Grok efforts in creation and editing. Clear unsupported effort when switching models. - Add catalog, CLI/ACP forwarding, compatibility, and UI persistence coverage. Record the audit and sources. ## Verification - Latest head `6e63c9ef53b54ba869cd4fb431a8570bebe289f4`: 53 CI checks passed, 2 skipped. This includes full workspace typecheck, build, and all test shards. Greptile is 5/5 with zero unresolved threads. GitHub reports no merge conflicts. - 340 focused tests passed across adapter metadata, CLI/ACP arguments, Claude version checks, Kimi effort, Grok execution, server model discovery, and UI effort selection/persistence. - `pnpm --filter @paperclipai/adapter-claude-local --filter @paperclipai/adapter-codex-local --filter @paperclipai/adapter-grok-local --filter @paperclipai/adapter-gemini-local --filter @paperclipai/adapter-kimi-local --filter @paperclipai/adapter-cursor-local --filter @paperclipai/adapter-opencode-local typecheck` — passed. The same filters with `build` passed. - `pnpm check:token-gates` and `git diff --check` — passed. - Full workspace and UI typechecks were attempted locally. They stop on existing missing `three` dependencies in `packages/shared/src/cliplab`. - Full `pnpm test:run` and `pnpm build` were not run locally. Worktree creation exhausted disk space, so a clean dependency install is not feasible on this host. Focused checks reuse existing dependencies. CI supplies full workspace verification. - No provider inference was run. Account-specific runtime model lists were inspected where available. - Manual check: select the new models in agent setup and editing. Confirm Luna has Max but no Ultra, Grok 4.7 has Extra High, and Fable 5.1 has Extra High/Max. Save Grok effort and confirm `adapterConfig.reasoningEffort` contains the selection. ## Risks - Catalog presence does not grant account access. Older CLIs and restricted accounts can reject a model. Opus 5.5 has an explicit upgrade check. - Higher effort can increase cost and latency. Existing agent defaults are unchanged. - Cursor fallback IDs come from public model documentation; the local account exposed no live catalog. Runtime discovery still adds account-specific variants. - Kimi effort remains supported only on its explicit CLI engine. This does not add effort support to its default ACP engine. - Saved obsolete Bedrock or retired Gemini IDs are not migrated automatically. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45c99a0d06 |
fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Legacy adapters launch provider CLIs and expose connected tools. > - Existing defaults did not consistently grant full automatic permission. > - Remote Claude used a fixed tool list that omitted MCP tools and future tools. > - Direct Codex launches and OpenCode configuration also used narrower defaults. > - This change gives all these paths the same full-auto default as native runners. > - Explicit restrictive settings continue to work. ## Linked Issues or Issue Description Refs #13686. This PR is stacked on that native-runner and task-reassignment PR. Merge #13686 first. Related: #831 (constructed Claude agents), #1935 (adapter-switching permission defaults). ## What Changed - Use actual Claude permission bypass for local and remote runs and probes. Remove the fixed tool list so MCP and future provider tools are included. - Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`. Do not mark ordinary host execution as a sandbox. - Default direct Codex execution to approval and sandbox bypass, matching agent creation. Preserve explicit false, CLI profiles, sandbox modes, approval policy, and network restrictions. - Set OpenCode's full-auto runtime permission to `allow` for every tool and connection. Preserve the existing explicit opt-out. - Default Gemini probes to the same YOLO mode as execution. Make the legacy ACP `default` alias use `approve-all` for fresh and resumed sessions. - Add default, opt-out, remote, probe, connected-tool, and resume regression tests. Update adapter configuration documentation. - Other adapter paths already request full automatic permission or have no provider approval gate. ## Verification - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server and legacy ACP tests passed, including defaults, explicit opt-outs, remote launches, connected tools, and fresh/resumed sessions. - Greptile reviewed current head `8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security reviewer acknowledged the documented full-auto requirement. Acknowledged discussions are resolved. - Current head has **54 passing checks**. [PR checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The process-adapter signoff browser shard passed on one retry after its first attempt exceeded a three-second issue-run wait. - **Six native Claude/Codex real-provider cases passed on their first attempt, with cleanup passing**, against the combined branch: plans, reassignment, and backlog creation/status. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). This does not claim a real-provider run of every legacy adapter. - The live-tested revision is `a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs only in the corrected heartbeat test expectation; application code is identical. - The campaign result-enforcement job passed. The separate report publisher failed during frozen dependency installation because the trusted workflow's patched-dependency configuration does not match its lockfile. Passing case evidence remains downloadable from the workflow. - Full-suite coverage comes from CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. ## Risks - Missing permission settings now grant all provider operations, including connected tools. OpenCode full-auto also overrides ambient provider permission rules. An explicit Paperclip permission opt-out preserves restrictive behavior. - Claude refuses full bypass as root outside an identified sandbox. Ordinary host deployments must run Claude as a non-root user. Managed sandbox launches include the required marker. - These defaults do not grant additional Paperclip roles, connections, or company access. Existing controller authorization and governance still apply. - This PR depends on #13686. Retarget it to master after that PR merges. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ed55f6931 |
fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts local command-line sessions and stores provider credentials through managed connections > - One OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run > - A second run then waited for the first run, and concurrent runs could overwrite a newer credential > - The write-back must compare fresh credentials while it holds the row lock > - This pull request removes the run lease, keeps the revocation guard, and bounds Codex timestamps against the host clock > - The benefit is safe concurrent use of one subscription connection with newest-credential selection ## Linked Issues or Issue Description **What happened?** A managed OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run. A second run waited for the first run to finish. The write-back gate also rejected any row change before it compared credential freshness. **Expected behavior** Concurrent runs should start on one subscription connection. The server should keep the newest valid credential and reject a credential write after a person revokes the connection. **Steps to reproduce** 1. Start two runs that use one OpenAI or xAI subscription connection. 2. Let both provider tools refresh the credential. 3. Finish the runs in either order. 4. Confirm that the newest valid credential remains in the connection. **Paperclip version or commit** `ec25bf1e4a81d1729a6d7276e6a486587cff4a0b` **Deployment mode** Local dev (`pnpm dev`), built from source. **Agent adapter(s) involved** Codex. The server path also covers xAI Grok subscription connections. **Database mode** Embedded PGlite for local development, and external Postgres for deployments. **Access context** Both board and agent runs can use managed connections. **Additional context** The provider command-line tool refreshes credentials inside the sandbox. The server copies the result back after the run. Two long runs can still refresh one token hours apart, so the provider can reject the second refresh. The server cannot observe that provider call. ## What Changed - Remove the full-run credential lease for OpenAI Codex and xAI Grok subscription connections. - Lock and re-read the connection row before credential write-back. - Accept only a strictly newer credential, while keeping the connection revocation guard. - Reject Codex freshness timestamps more than five minutes ahead of the host clock. - Add tests for both completion orders, xAI cleanup, revocation, the Codex time bound, and agent hiring. - Update the connection and run-log documentation. ## Verification - `server/src/__tests__/ai-connections.test.ts` passes with 42 tests. - `packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts` covers the five-minute boundary and the one-millisecond overflow. - `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` passes. - `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI and Anthropic. - The project type check reports no new error in changed files. - GitHub Actions must pass on this pull request. ## Risks The write-back now permits concurrent runs, so the provider may reject a later refresh when both long runs use one token. The server keeps the revocation guard and rejects future-dated Codex timestamps. No schema change occurs. ## Model Used OpenAI Codex, GPT-5. Context window and exact deployment build are not exposed in this run. The model used tool calls, code inspection, and test verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f12b647ae8 |
fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can collect more messages while its agent works. > - Legacy runners must stop the active process before they can receive those messages. > - The old Interrupt action cancelled the run but could leave the queue idle and hidden. > - Codex could also classify a cancelled run as successful or start a fresh process after cancellation. > - This pull request joins cancellation, preserves the provider session, and dispatches the current queue after cleanup. > - The benefit is reliable interruption with the saved message order, edits, and deletions. ## Linked Issues or Issue Description **What happened?** Interrupt could strand a legacy message queue. The UI could hide pending messages after the run stopped. A Codex signal exit could race the cancellation write. A stale session warning could also trigger a fresh process after an interrupted resume. **Expected behavior** Interrupt stops the active turn and sends the remaining messages once, in their saved order. Deleted messages stay deleted. An interrupted Codex turn keeps its session and does not restart itself. **Steps to reproduce** 1. Assign a task to a legacy Codex agent that runs a long command. 2. Queue three messages. Edit one, discard another, and move the last message first. 3. Click Interrupt in the queue. 4. Repeat the interruption while the resumed session runs another command. Related work: Refs #13160, which moves native queue steering into the wake-queue module. This change fixes legacy interruption and keeps native steering unchanged. ## What Changed - Add a revision-checked, company-scoped endpoint for legacy queue interruption. - Promote only the requested queue after the provider stops and releases its lease. Retry its persisted interrupt intent from the scheduler after a promotion error or server restart. - Keep pending legacy queues visible after a run stops. Use server state for the interrupt result. - Serialize owned process cancellation before classifying the adapter result. Preserve late session and log metadata. Acknowledge cancellation only when an actual process or process group was owned; scheduler placeholders retain their normal release policy. - Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the session has started. - Add cancellation race, multi-actor queue order, durable retry, resume fallback, and stale request regression tests. Document the behavior. ## Verification - Real browser tests passed with legacy Codex CLI and ACP engines, using Codex 0.153.4 and gpt-5.6-sol. - All three automated ACP browser scenarios passed locally: immediate Interrupt delivery, no replay of an unfinished write, and pause requiring Resume. Updated the old test expectation that required a separate “go” after Interrupt. - Browser tests covered queued edits, deletion, reordering, deleting the final message, and repeated interruption. - Two consecutive CLI interrupts kept one provider session. Both stopped processes exited. The final message arrived once. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed. - All 346 post-review scheduling, recovery, queue-route, archived-company, worktree-suppression, and stale-queue regression tests passed. - All 318 process-recovery and durable-chat tests passed after the final cancellation guard. - Codex adapter, queue UI, issue-page, and OpenAPI contract tests passed. - `pnpm build` passed. - Full local suite coverage completed with `PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI shards: 618 general server suites, all 145 serialized server suites, and all workspace groups. Every failing suite passed a targeted rerun after the fixes, rebuilding the native test fixture, correcting macOS temporary-path setup, or retrying setup/timing failures. Existing skips remain. - The original monolithic run reported failures before the final fixes; its failed suites were rerun rather than rerunning all 618 suites again. The final process-recovery/durable-chat regression run passed all 318 tests. - All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: [run 34654820774, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2), including typecheck, build, all test shards, E2E, and canary. The signoff and Cursor sandbox tests each hit a timeout in the initial attempt; both suites passed locally, and both failed shards passed their single CI rerun. All three corrected ACP browser scenarios passed in CI. - Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no open review threads. ## Risks Cancellation order affects local adapters. The tests cover signal exits, graceful exits, adapter exceptions, termination errors, and cancellation write errors. Embedded adapters keep their cancellation controls. Ordinary run cancellation and task pause keep their distinct queue policies. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code execution. The exact serving model ID and context-window size are not exposed in this session. The live test runner used OpenAI gpt-5.6-sol. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bce976d60d |
feat: bind an agent to a Codex login whose account differs from the company default (#13067)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter signs agents in to OpenAI, and a company keeps one default Codex identity in its shared company home > - A login with a DIFFERENT account than the company default is deliberately kept out of the shared home — one agent's sign-in must not switch every unbound agent's credentials — but that left the cross-account login inert: nothing connected the agent the operator was configuring to the credential the login stored > - The stored credential and its company secret already exist; only the last mile — an agent actually using them — was missing > - This pull request reports a non-secret binding claim on the authenticated login and lets the agent page bind that one agent's `CODEX_HOME` to the account's secret, exactly and only when the identities differ > - The benefit is that multi-account Codex becomes one click on the agent that needs it, with company-wide identity untouched ## Linked Issues or Issue Description **What happened?** On an agent's detail page, "Sign in with Codex" using a different OpenAI account than the company default succeeds but changes nothing for that agent. The credential lands in the per-identity store and a company secret names it, but the agent keeps using the company default. The Test keeps reporting that authentication is needed, and no repeat login helps. **Expected behavior** When the operator deliberately signs an agent's page in with a different account, that agent starts using that account. Agents that were not part of the action keep the company default. A same-account login keeps working through the shared company home with no per-agent pinning. **Steps to reproduce** 1. Configure a company whose Codex home holds account A. 2. Open a `codex_local` agent's detail page with a sandbox environment and complete "Sign in with Codex" using account B. 3. Press Test. Before this change the agent still resolves account A and the authentication-needed check returns. ## What Changed - `packages/adapters/codex-local` — the prerequisite shield: `isCodexAuthCachePath` recognizes per-identity credential-store entries, and `seedManagedCodexHome` refuses to symlink, heal, or API-key-overwrite an entry's `auth.json`. The seeding pass runs before every probe and execute; without the shield, an agent bound to an entry would have its stored login silently swapped for the host credential. Static shared config files still copy in. Rotation already survives binding: the sandbox copy-back writes rotated credentials into the identity-keyed store slot. - `server` — the promotion records whether the company default home ended on a different account than the login (any read failure degrades to `false`, so the client can never be told to bind wrongly). After the terminal commit, the routes layer remembers a non-secret claim — the opaque account-home secret id plus that verdict — in a bounded in-memory map, and merges it into the owner read of an `authenticated` `codex_local` session. A restart drops the claim; the panel then shows plain success. - `packages/shared` — `CodexAccountBindingClaim` on the owner session response. It carries no account identifier and no credential byte. - `ui` — the login panel reports the claim upward once. The edit-mode form binds the agent's `CODEX_HOME` to the secret and saves in one step, only when `companyIdentityDiffers` is true. Same-account logins bind nothing on purpose: the company-home refresh already carried them, and an unbound agent keeps following the company default across rotations. Create mode is unchanged. ## Verification - Adapter suite: 381 passed, 1 skipped (includes the new store-entry shield tests and the path-predicate cases). - Server suites (8 files): 130 passed, 15 skipped — including two new route tests that drive a login to `authenticated` and assert the claim with both identity verdicts. - UI render suite: 85 passed — including a panel test that the claim is reported upward exactly once. - `tsc --noEmit` clean in `packages/shared`, the adapter package, and `ui`; `server` clean for the touched file. ## Risks - The bind changes one agent's configuration through the normal agent-update patch, initiated by the operator's own login on that agent's page. The failure direction of every fallback is "offer nothing": a missing claim, a restart, or an unreadable company home all degrade to no bind. - The seed shield narrows what the seeding pass may touch; homes outside the credential store behave exactly as before, covered by the existing seed tests. - Builds on the sign-in credential-resolution fix (#13064), now merged; this branch is rebased onto master and the diff contains only the binding feature. Supersedes #13066, which GitHub auto-closed when its stacked base branch was deleted on merge. ## Model Used Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use in Claude Code (terminal). **Related PRs (searched; no duplicates found):** #12740, #12082, and #9621 touch adjacent Codex credential sync paths; #8495 is the standing hardening effort for probe auth seeding. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no standalone docs cover this flow; the behavioral contracts are documented in-line at each changed site) - [x] I have considered and documented any risks above |
||
|
|
2991a59b17 |
fix(adapters): prevent engine fallback and preserve usable runtime defaults (#13105)
## Thinking Path > - Paperclip manages agents that must write work and report task outcomes through its API. > - Local adapters select an execution engine and its permission settings. > - A higher ACP Node requirement can make an unchanged installation lose access to its default engine. > - The adapter then silently selects CLI, which can change permissions and block API access. > - This pull request keeps the engine choice fixed and reports missing prerequisites before work starts. > - It also gives explicit Codex CLI runs usable defaults and keeps managed services on a supported Node runtime. ## Linked Issues or Issue Description Refs #12215. Related changes: #11792 raised the Node requirement; #13094 addressed separate runner networking behavior. This change fixes the engine-selection and managed-launcher paths. **What happened?** An unchanged agent could switch from ACP to CLI after an upgrade. Codex CLI then used read-only permissions with networking disabled. The run could finish without updating its task. Repeated recovery attempts used the same unavailable setup. Managed updates also skipped the Node check and did not refresh old launchers. **Expected behavior** An unavailable engine must fail with a clear setup error. It must not silently select another engine. Explicit CLI runs must be able to write workspace files and call the API unless the operator configures stricter settings. Managed updates must validate Node and keep child tools on that runtime. **Steps to reproduce** 1. Run an ACP-default agent under Node 22 after the ACP minimum rises to 24.11. 2. Leave the engine unset and disable the approval/sandbox bypass. 3. Observe the old adapter select CLI and fail to write task disposition through the API. 4. Start a managed service with an old launcher and a supervisor PATH that selects a different Node for child tools. ## What Changed - Remove automatic engine fallback for Codex, Claude, Gemini, and Kimi. Check prerequisites for default and explicit ACP selections. - Return a configuration error with proof that provider work did not start. Stop automatic continuation retries for this error. - Enable Codex ACP workspace networking at the actual turn boundary. Upstream mode presets otherwise force it off even when config.toml enables it. Preserve explicit network denial and read-only mode. - Set workspace-write and network access defaults for explicit Codex CLI runs. Preserve explicit sandbox modes, profiles, and network restrictions. - Pin the validated Node directory in managed launcher PATH. Refresh legacy launchers during installs and npm/Git updates. - Reject updates on unsupported Node. Keep update checks, dry runs, and rollback available. - Synchronize the qualified Codex ACP executable identity across server, TypeScript runner, Rust runner, and provider-pack launch paths. - Add regression tests and update engine and installation documentation. ## Verification - [Full CI passed on the final head](https://github.com/paperclipai/paperclip/actions/runs/34387099695): typecheck, build/native runner verification, all general and serialized test shards, all browser shards, release registry, canary dry run, and policy checks. - Greptile: 5/5 on `2c1d6e2815830a5cd39e36c8a082cc0c4441b6c0`, with no unresolved review findings. Security gates are green. - Full workspace typecheck and build also passed locally. The final deployed Linux build passed. - Full Codex, Claude, Gemini, and Kimi source test suites: 804 passed, 2 skipped. Installer, updater, and launcher tests: 47 passed. Installed ACP turn-boundary tests: 3 passed. ACP packaging tests: 14 passed. Focused recovery classification tests also passed. - Real Linux Codex CLI runs, both fresh and resumed, wrote a workspace file and reached the control-plane health API with the new defaults. - Explicit read-only and network-disabled control probes retained those restrictions. - A real ACP run on the final deployed Linux build wrote a file and reached the control-plane API with HTTP 200, without engine fallback. The same probe failed DNS before the turn-policy patch. - Executable-identity and installed-policy contracts: 12 passed. Affected native server tests: 197 passed. Runner factory tests: 21 passed. Rust qualification and native provider integration tests: 11 passed. - Deployed the production changes to a Linux service on Node 24.20 after a verified database backup. Health, bootstrap readiness, static UI, executable/cwd identity, and guarded restart checks passed. The restart lost no runs. - Corrected stale Kimi skill-default and Gemini remote-archive fixtures; both suites pass. ## Risks - Default or legacy auto engine settings now fail when ACP is unavailable. Operators who intend to use CLI must select it explicitly. - Codex CLI now permits workspace writes and networking by default, and ACP workspace-write turns permit networking by default. Explicit operator sandbox settings remain authoritative. - Old managed launchers keep their pinned Node until they are reinstalled under a supported runtime. An old updater cannot repair itself; the documentation gives the current installer command. - Custom service wrappers and global/source installations must configure their runtime PATH. No database migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, shell execution, and test tools. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fe5e68d7a5 |
fix: make Codex sign-in and the environment test agree on the credential a run uses (#13064)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter signs agents in to OpenAI with a device-code login, and the agent page has a Test button that probes the sandbox with the credentials a real run would use > - The login stored its credential where the Test never looked: the company Codex home kept an old shape-valid credential, so the Test failed with "authentication needed" right after a successful sign-in > - The Test also staged a different Codex home than a real run resolves, so the Test and real runs could disagree in both directions > - This pull request makes the login, the seeding pass, and the Test probe agree on one credential resolution > - The benefit is that a sign-in from the agent page fixes the Test on the next click, and a green Test means the same thing a real run experiences ## Linked Issues or Issue Description **What happened?** Sign in with Codex works during onboarding but not on the agent detail page. The operator completes the device-code login. The panel reports success. The Test button still reports that authentication is needed. No number of repeat logins changes the result. Three defects combine to cause this: 1. The device-login promotion only wrote the company default Codex home when that home held no shape-valid credential. A stale credential (for example a symlink to an old host login) blocked the write forever, so the fresh login stayed invisible to the Test. 2. The seeding pass that runs before every probe and execute replaced a same-identity regular-file `auth.json` with a symlink to the host credential, with no freshness comparison. Even a freshly promoted credential was deleted on the next Test. 3. The sandbox hello probe always staged the company default home. An agent with a configured `CODEX_HOME` was tested against one credential and ran with another. **Expected behavior** A completed sign-in updates the credential the Test probes. The Test stages the same Codex home a real run resolves. A stale credential never outranks a strictly newer one from an interactive login. **Steps to reproduce** 1. Configure a company whose Codex home holds a shape-valid credential that no longer authenticates (for example an old host login symlink). 2. Open a `codex_local` agent's detail page with a sandbox environment and press Test. The result shows the authentication-needed check. 3. Complete the "Sign in with Codex" device-code flow from the panel. 4. Press Test again. Before this change the result still shows authentication needed. ## What Changed - `packages/adapters/codex-local/src/server/adapter-auth-promotion.ts`: the promotion writes the company default home unconditionally. The shared `last_refresh` merge predicate scopes the write. It seeds an absent or unusable slot, refreshes a same-identity slot only with a strictly newer credential, and keeps a slot a different account or an API-key file holds. The atomic rename replaces a symlinked `auth.json` at the link itself. It never writes through into the host home. - `packages/adapters/codex-local/src/server/codex-home.ts`: the same-identity heal in `seedManagedCodexHome` is freshness-aware. A regular-file credential is swapped for the shared symlink only when the shared source is strictly fresher by `last_refresh`. Ties and unparseable timestamps keep the file, which matches the predicate's fail-closed direction. A genuine stale copy still heals as soon as the host credential rotates past it. - `packages/adapters/codex-local/src/server/test.ts`: the sandbox hello probe prepares and stages the same home a real run resolves. The identity-anchored cache vend runs first. A configured managed `CODEX_HOME` is seeded in place and staged. A genuine external override is staged as-is and never seeded or mutated. - Tests: new pins for the strictly-newer company-home refresh, the different-account keep, the symlink-replaced-without-writing-its-target property, the freshness-aware heal (newer kept, older healed, ties kept), the configured-home staging, and the external-home no-mutation proof. The remote-probe suite now pins `CODEX_HOME`/`PAPERCLIP_HOME` to scratch directories so no test can touch a real `~/.codex`. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src --root packages/adapters/codex-local` — 378 passed, 1 skipped. - Server suites for device login, reconciliation, and the codex adapter (8 files) — 128 passed, 15 skipped. - `tsc --noEmit` clean in the adapter package. ## Risks - Behavioral shift is scoped by the shared merge predicate: only a strictly newer same-identity login can displace a company-home credential, so a second account still never takes over the company slot, and API-key files are never displaced. - The heal keeps ties and unparseable timestamps instead of swapping. A kept file self-corrects on a later seed once the source is provably fresher; deleting a promoted credential is irreversible, so the failure direction is chosen deliberately. - The probe change makes the Test exercise the credential a run uses. A Test that previously passed against the company home while the agent's configured home was broken now fails honestly. ## Model Used Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use in Claude Code (terminal). Diagnosis traced through the live promotion locks, the on-disk Codex homes, and the adapter's credential-resolution code paths. **Related PRs (searched; no duplicates found):** #12740, #12082, and #9621 touch adjacent Codex credential sync paths; #8495 is the standing hardening effort for probe auth seeding. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no standalone docs cover this flow; the behavioral contracts are documented in-line at each changed site) - [x] I have considered and documented any risks above |
||
|
|
2043e0c735 |
fix: repair runner configuration, macOS execution, and artifact galleries (#13062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters select a provider, a model, and a runtime. > - Runner conversion rejected existing Claude agents. The model list mixed providers. > - The native Claude runner rejected custom models and could not launch on macOS. > - This pull request fixes conversion, model selection, and verified macOS execution. > - It also groups configuration fields consistently across adapters and opens artifact images in the task gallery. > - Operators can change an agent configuration and run the selected model on their Mac. ## Linked Issues or Issue Description **What happened?** Converting an existing Claude agent to Paperclip Runner failed with a Codex-only restriction. ACPX Claude showed unrelated models and required `claude-sonnet-5`. Its native runtime rejected macOS. Configuration mixed common model settings with process controls. Artifact cards labeled “Open gallery” navigated to attachment URLs instead of opening the task gallery. **Expected behavior** Conversion keeps agent identity and compatible settings. ACPX Claude uses the normal Claude catalog and accepts typed model IDs. Codex uses the native runner. The verified Claude runtime can launch on macOS ARM64 and x64. Common configuration sections place the same fields together across adapters. Artifact images open in the shared task gallery with navigation and downloads. **Steps to reproduce** 1. Open the configuration of an existing Claude agent. 2. Convert it to Paperclip Runner. 3. Select ACPX Claude and a different catalog model or a typed model ID. 4. Save the agent and run a disposable task on macOS. 5. Inspect configuration and advanced run-policy controls across adapters. **Paperclip version or commit** The bugs were reproduced on `165ca56a22adb60e5fda56045442d9c8498116a8`. This branch was rebased onto `7ed122911`. **Deployment mode** Built from source. Local test-drive instance on macOS ARM64 with an isolated database. Related work: #11798 addresses unsupported ACP session options in the existing adapter path. #13048 addresses working-folder preservation. This change fixes native runner configuration and launch behavior. ## What Changed - Remove the Codex-only conversion restriction. Preserve agent identity, instructions, directories, credentials, and compatible model settings. Reset incompatible sessions while retaining history. - Show ACPX Claude and native Codex as distinct provider choices. Remove ACPX Codex from advertised configuration. Normalize legacy configurations before fresh runs without rewriting historical run descriptors. - Select model catalogs and cache entries by provider. Support refresh and typed model IDs. Pass exact Claude IDs through session creation, model changes, and recovery. - Add verified macOS ARM64 and x64 Claude SDK snapshots. Bound executable allocation and total snapshot size. Preserve package checks, dependency isolation, process ownership, cancellation, and Linux descriptor loading. - Probe local runtime readiness. Report remote platform checks as incomplete until the remote runner verifies its runtime. - Surface actual model rejection and allow correction and retry. - Repair missing ACPX goal-capability helpers exposed by the post-rebase live test. Persist and restore the optional capability without breaking session startup. - Put Agent identity first and intentionally remove the Capabilities editor, as requested. This is removal of UI editing, not relocation: preserve existing capability metadata and API compatibility without adding another editor. Use the themed select for configurable permission modes, with normal text instead of monospace. - Put model and provider under Adapter. Give environment variables their own section. Fold command and arguments under Configuration. Fold lifecycle, timeout, and interrupt grace under Advanced Run Policy. Hide single-option permission controls. - Open image and video artifact cards in the existing task gallery, including cards in the artifacts panel. Chat attachment images use the same gallery. Preserve standalone media previews and download links. ## Verification - Rebased focused UI/API/database suites: 293 tests passed. - Rebased native runtime and ACPX suites: 242 passed, 7 skipped. - Repository typecheck, build, and token gates passed for the runner changes. Gallery follow-up UI typecheck, build, and token gates also passed. - Follow-up UI suites passed (86 tests), packaging checks passed (14 tests), and the final focused runtime suites passed (126 passed, 7 skipped). - Linux container isolation and lifecycle fixtures passed before rebase (57 passed, 2 skipped). Rust ACPX provider-session tests passed after rebase (8 tests). - Browser tests completed actual Claude and native Codex tasks on macOS ARM64. They covered conversion, catalog refresh, a non-default catalog model, a typed `haiku` ID, save/reload, cancel, follow-up session continuity, invalid-model errors, and recovery. - Final-revision live tests completed a typed Claude task, a follow-up with the same provider session, and a native Codex task on macOS ARM64. - Browser tests confirmed the moved interrupt-grace field saves and survives reload. Cross-adapter tests cover Claude, Codex, Gemini, process, gateway, and schema forms. - Full local run: 7,080 passed, 30 skipped, and two timeouts. Both timeout suites passed on isolated rerun (84 tests); the failures were the plugin login-worker exit diagnostic and the runner real-server vertical slice. - Final follow-up checks: 50 registry tests and 45 snapshot/installation tests passed (6 platform-specific skips). Oversized executable rejection is covered before allocation or reading; unsupported-platform tests invoke the real installation probe. - Runner head `ddb5101c483a297f74875ab96b3c66035b002d50`: all CI gates green, including full runner verification, repository build, typecheck, general/serialized server suites, browser tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34286178670). - Greptile: 5/5 on that runner head. All four review threads resolved. Superagent, Socket, and Snyk checks green. - After snapshot hardening, another real Claude task completed on this Mac using the rebuilt runtime. - Gallery follow-up: 148 focused tests passed, covering artifact selection, shared attachment collections, deduplication, image/video cards, standalone previews, downloads, and closing. Live browser verification completed on the settings follow-up: artifact selection, 6-image pagination with wrapping, download action, and closing all stayed on the same task URL. All checks passed on gallery head `96136da58ff195bf6ca00b281eb3022ad12d7bd8`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34287987536). Greptile returned 5/5 on that exact head with no unresolved threads. - Final settings polish: 96 focused tests, UI typecheck/build, and token gates passed. A real browser walkthrough verified readable permission options, identity placement, Capabilities removal, and permission save/reload. Original test-agent permission mode restored. All 31 checks passed on final head `e46540d6bf32bfb0566dca16b2f4a75ba437618c`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34292797886). Greptile returned 5/5 with no unresolved threads. ## Risks - Capabilities intentionally has no editable UI field after this change. Existing values remain readable and API-compatible; removing the field does not erase stored metadata. - macOS launch now copies verified package files into private snapshots. The implementation must retain isolation and clean up snapshots on exit. - Runtime provider or model changes reset the current session. Historical runs remain available. - The macOS x64 SDK executable digest was verified, but a live Intel Mac run was not available. Linux verification used container fixtures, not a real Claude task. - Remote environment tests report a warning when only the platform has been checked. They do not claim package readiness from the server host. ## Model Used OpenAI Codex, based on GPT-6. The exact served model identifier and context-window limit are not exposed in this session. Used reasoning, repository inspection, code execution, Rust and TypeScript tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites and both timeout suites on rerun; full-run counts above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
77312ee2d9 |
feat(codex): add GPT-6 Astra support (#12851)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The Codex local adapter supplies model metadata to the server and the user interface. > - OpenAI now lists `gpt-6-astra` as a supported Codex model. > - Paperclip did not list this model or its model-specific controls. > - This pull request adds the model through the existing adapter metadata path. > - The benefit is that agents and task overrides can use the exact model ID and supported controls. ## Linked Issues or Issue Description **Subsystem affected** `packages/adapters` and `ui` **Problem or motivation** Paperclip does not expose `gpt-6-astra` in Codex model selectors. Operators cannot select and save the model through the normal agent and task forms. **Proposed solution** Register the exact model ID in the Codex local adapter. Use the adapter as the source for the model-specific reasoning options. Preserve the current default model. Forward the saved model, reasoning effort, and fast-mode controls through both Codex execution lanes. **Alternatives considered** A user-interface-only model list would duplicate adapter metadata. A model alias would not match the official model ID. Both options were rejected. **Roadmap alignment** This is a small adapter compatibility update. It does not duplicate a planned item in `ROADMAP.md`. ## What Changed - Added `gpt-6-astra` to the Codex local adapter model registry and fast-mode support list. - Added the official Astra reasoning efforts: `low`, `medium`, `high`, `xhigh`, `max`, and `ultra`. - Used the adapter metadata in agent and task model selectors. - Preserved supported effort choices when the model changes. Cleared an effort only when the new model does not support it. - Added tests for registration, user-interface selection, configuration persistence, and CLI and ACP forwarding. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/lib/codex-reasoning-effort.test.ts ui/src/components/AgentConfigForm.render.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/issue-assignee-overrides.test.ts` passed 245 tests. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed all four gates across 939 files. - `pnpm --filter @paperclipai/ui build` passed and supplied isolated user-interface build proof. - `pnpm build` passed. - `pnpm test:run` passed 5,812 tests and failed 24 workspace-runtime tests in this isolated host. The failures use invalid generated ports above 65,535, incomplete nested-worktree fixture configuration, or `/tmp` path aliases. The focused tests for this change all pass. GitHub CI must pass before review handoff. - GitHub CI run `33918372718` passed all required checks and the aggregate verify gate on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - Independent engineering review approved the exact remediation head after 170/170 reviewer tests passed. - Greptile reported 5/5 with no open review threads on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - The model ID and capabilities were checked against the [official OpenAI Codex model list](https://developers.openai.com/codex/models). ## Risks - Low risk. The change adds one model and model-specific selector options. It does not change the default model. - OpenAI can change model capabilities later. The adapter metadata must stay aligned with the official Codex metadata. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model ID `gpt-5.6-sol`, a 272,000-token context window, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9064cfd09e |
feat(codex-local): give each Codex account its own home and path secret (#12709)
## Thinking Path > - Paperclip is the control plane for companies that use AI agents for work > - Local adapters connect Paperclip agents to provider command line tools > - The Codex adapter stores login data in a shared company home > - A shared home cannot keep credentials for more than one Codex account > - This pull request gives each account a safe home and a matching company secret > - The benefit is that one company can use multiple Codex accounts at the same time ## Linked Issues or Issue Description **Problem or motivation** A company can hold only one Codex subscription credential because device login uses one shared home. A second account cannot log in without replacing or conflicting with the first credential. **Proposed solution** This change validates the vendor account identifier, stores each credential in its own home, and creates a company secret that points to that home. Repeat login calls return success when the matching secret already exists. **Roadmap alignment** The change supports the roadmap goal for centrally managed secrets with scoped access and audited resolution. **Additional context** The security review returned approve with no blocking finding. The branch adds shared account-handle validation and tests for device login and the Codex local adapter. ## What Changed - Add strict allowlist validation for Codex account handles. - Store each Codex account credential in a separate home under the Codex cache root. - Verify that the resolved account home stays inside the cache root. - Create the `CODEX_HOME_<handle>` company secret for each account. - Keep repeat and concurrent login calls safe and idempotent. - Add shared helper and route, adapter, and validation tests. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local test` passes with 343 tests. - `pnpm --filter @paperclipai/server test src/__tests__/agent-device-login-routes.test.ts` passes with 25 tests. - The adapter suite passes with 23 tests. - The shared package and Codex adapter typechecks pass. - Continuous integration must pass on every check before merge. ## Risks The account handle becomes part of a directory path and secret name. The strict allowlist and root containment check reduce path traversal risk. Existing single-account homes remain unchanged unless a new device login creates an account-specific home. ## Model Used OpenAI GPT-5 (exact runtime model ID: gpt-5), with tool use and code execution. The runtime context window is not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
84bedd4ca1 |
feat(runner): activate qualified OpenCode and ACPX providers (#12691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is the experimental native runtime for governed agent work. > - The runtime contracts already describe Codex, OpenCode, and ACPX providers. > - The merged control plane still rejected OpenCode and ACPX for new runner agents. > - Runnerd also selected only the Codex provider implementation. > - This pull request activates the qualified OpenCode and ACPX paths from the form to runnerd. > - The benefit is one durable runner path with provider-specific permissions and recovery. ## Linked Issues or Issue Description Refs #12685 **Subsystem affected** This change affects the runner package, server orchestration, adapter configuration, and UI configuration. **Problem or motivation** Paperclip Runner stores provider contracts for OpenCode and ACPX. New agents cannot select those providers. Runnerd cannot execute those stored provider descriptors. The UI also shows only Codex. **Proposed solution** Accept the qualified OpenCode 1.18.17 profile and the fixed ACPX Claude and Codex profiles. Route them through runnerd. Keep provider selection, model selection, permissions, credentials, events, and recovery inside closed provider-specific boundaries. **Alternatives considered** One option was to keep the contracts dormant. That option leaves stored configuration and runtime behavior out of sync. Another option was to enable every ACPX agent. That option is not safe because Pi does not yet have the same verified launch path. **Roadmap alignment** This change supports the completed cloud and sandbox agent milestone. It also supports self-healing runs and governed agent execution. It does not add a new roadmap surface. ## What Changed - Add one server profile resolver for Codex, OpenCode, and qualified ACPX descriptors. - Keep `adapterConfig` as the provider and permission authority for fresh runs. - Add Paperclip Runner provider, ACPX agent, and provider-specific permission controls to the UI. - Reset the model to a compatible qualified value when the provider changes. - Route Codex, OpenCode, and ACPX through the durable runnerd provider selector. - Add a durable ACPX executor with bounded state, recovery, events, tool receipts, and identity checks. - Remove Codex labels from OpenCode events, results, evidence, and recovery diagnostics. - Pass only provider-specific credential names to child processes. - Keep ACPX Pi unavailable and reject it before process launch. - Keep the existing Paperclip Runner experimental flag unchanged. ## Verification - `pnpm exec vitest run packages/paperclip-runner/src/backends/native-backend-factory.test.ts packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/adapters/codex-local/config-fields.test.tsx server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/__tests__/agent-adapter-validation-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/services/native-runtime/runtime-mode.test.ts server/src/services/native-runtime/native-session-executor.test.ts server/src/services/heartbeat-runner-provider-config.test.ts` - The focused TypeScript, server, and UI suites passed 274 tests. - `cargo test -p paperclip-runner-core --test native_provider_backend` - The executable native provider integration suite passed 4 tests. - `cargo test -p paperclip-runner-core --lib` - The Rust unit suite passed 91 tests. - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - `git diff --check codex/runner-parity-task-runtime...HEAD` ## Risks - This changes provider process selection and durable recovery. The experimental flag still gates every fresh Paperclip Runner run. - OpenCode requires a model in `provider/model` form and stays pinned to version 1.18.17. - ACPX accepts only exact Claude and Codex profile versions and models. Pi stays unavailable. - ACPX steering stays unavailable and reports that limit through the driver capabilities. - Child processes receive explicit environment allowlists. They do not inherit the full server environment. - This pull request has no database migration. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
131f5c4065 |
feat(runner): add administration and observability (#12641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Administrators need bounded controls for experimental native execution. > - The lower stack adds remote Codex execution and the task workspace. > - Operators need to configure Codex safely and inspect provider traces. > - Unsupported providers must not appear as runnable choices. > - This pull request adds Codex-only administration and observability. > - The benefit is a default-off operational surface for production diagnosis. ## Linked Issues or Issue Description Refs #12640. Refs #12616. Refs #12352. **Subsystem affected** Agent configuration, instance experimental settings, run ledger, provider trace inspector, and administrator actions. **Problem or motivation** The native runner lacks one safe operator surface for Codex permissions, lifecycle, raw trace capture, and run inspection. The integration branch also contains provider choices that the production backend cannot execute yet. **Proposed solution** Expose only the qualified Codex controls. Keep Paperclip Developer Mode and runner preview ingress off by default. Gate raw trace actions by administrator access and existing trace authorization. **Alternatives considered** Exposing unfinished providers would create configurations that fail at runtime. Always-on tracing would increase sensitive data and storage risk. **Roadmap alignment** This work supports governed Cloud and Sandbox agents and production diagnostics. ## Stack - Base PR: #12640. - Lower PRs: #12639 and #12638. - This PR contains only its 54-file administration and observability delta. - This is the final feature PR in the Codex production stack. ## What Changed - Added Codex-only Paperclip Runner permission and lifecycle controls. - Added bounded warm idle configuration. - Kept the provider field fixed to Codex. - Added administrator-only one-run raw trace requests. - Added a persistent future-run raw trace toggle. - Added trace status, metadata, ledger, and canonical runner inspection. - Added JSON-RPC request-origin grouping and finalization lineage. - Restored the stateful PRP transcript parser and focused projection tests required by trace inspection. - Added default-off Paperclip Developer Mode. - Added Honeycomb run links for authorized developer mode. - Disabled the legacy operational skill for `paperclip_runner`. - Did not expose OpenCode, ACPX, Pi, Claude Managed, or AWS runner choices. - Did not change migrations, workflows, dependencies, or `pnpm-lock.yaml`. ## Verification - GitHub Actions will run UI tests, server tests, repository typecheck, build, browser tests, security, and policy gates. - Tests cover Codex configuration defaults and bounds, administrator trace actions, persistent settings, ledger inspection, trace lineage, and Honeycomb links. - Existing server trace authorization and retention tests remain the backend authority. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check runner/task-workspace-experience...HEAD` passes. - The delta contains 54 files. ## Risks - Raw provider traces can contain sensitive provider data. - Existing server authorization controls access, reveal, download, retention, and deletion. - The UI gates trace actions by administrator access and developer mode. - All new instance settings remain off by default. - Fresh Paperclip Runner configuration remains Codex-only. - Direct adapters and legacy task behavior do not change in this PR. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
2e5a24e177 |
feat(runner): add qualified OpenCode runtime (#12588)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner provides a durable execution boundary for supported providers. > - The current production runtime supports Codex but cannot execute OpenCode sessions. > - OpenCode needs a qualified transport, strict input mapping, and normalized events. > - This pull request adds the OpenCode runtime as one isolated provider unit. > - The benefit is a reviewable provider expansion that does not weaken the existing Codex path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: packages/paperclip-runner and the Codex-local adapter configuration contract. **Problem or motivation** Paperclip Runner has provider-neutral contracts, but the production backend factory cannot start a qualified OpenCode session. This blocks OpenCode from using the durable runner path. **Proposed solution** Add the qualified OpenCode app-server proxy, driver, MCP bridge, backend, fixtures, and factory wiring. Keep existing Codex behavior unchanged. **Alternatives considered** Keeping OpenCode only on the direct adapter path would avoid this runtime work, but it would not provide durable runner recovery or normalized provider events. **Roadmap alignment** ROADMAP.md does not list a conflicting provider-runtime project. This change extends the existing Paperclip Runner architecture. ## What Changed - Added the qualified OpenCode app-server proxy and input queue. - Added collaboration-mode and provider-event normalization. - Added the OpenCode MCP bridge and native session backend. - Added strict fixtures and focused unit coverage. - Added only the package exports and adapter configuration required by this runtime. - Kept deferred SDK, lab, eval, and public package surfaces out of this change. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - Run the package type checks and focused OpenCode tests in CI. - Run repository typecheck, test, build, security, and policy gates through the stack-aware workflow. - Local tests were not run because this checkout is resource constrained. ## Risks - OpenCode protocol changes could affect event normalization or recovery. - The driver fails closed on malformed input and unsupported runtime behavior. - Existing Codex selection remains unchanged unless the stored provider is OpenCode. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack - Position: 1 of 4 - Base: master - Next: additional qualified provider runtimes |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9d0927fe8 |
fix(adapters): restore Paperclip skill for legacy runners (#12225)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy local adapters run agents that use the Paperclip skill for the control-plane workflow. > - PR #7029 removed the required-skill fallback and made runtime skill selection depend only on stored preferences. > - No migration or runtime fallback replaced that behavior for existing agents or non-CEO agents. > - PR #12138 added core skills to new CEOs, and PR #12147 added Claude skill discovery. These changes did not mount the operational skill for all legacy agents. > - This pull request makes the operational skill a legacy adapter runtime invariant. It keeps all other skills configurable. > - The native runner stays unchanged because its protocol supplies the control-plane contract. > - The benefit is that new and existing legacy agents can always operate through Paperclip. ## Linked Issues or Issue Description Refs #7029 Refs #12138 Refs #12147 **What happened?** A skill-capable legacy local agent could start without `paperclipai/paperclip/paperclip`. This happened when the agent had no stored skill preference. An explicit empty preference also removed the skill. The agent then reported that the Paperclip skill was not available. **Expected behavior** Every skill-capable legacy local adapter must mount the Paperclip operational skill when the runtime inventory contains it. Optional skills must remain configurable. The native runner must keep its current protocol-based behavior. **Steps to reproduce** 1. Create a non-CEO `codex_local` agent without `paperclipSkillSync` preferences. 2. Start a legacy heartbeat. 3. Inspect the managed `CODEX_HOME/skills` directory. 4. Observe that the Paperclip skill is absent before this change. **Paperclip version or commit** The problem reproduces on `master` before this pull request. PR #7029 introduced the configured-only selection behavior. **Deployment mode** Local development and self-hosted legacy local adapters. ## What Changed - Added a shared legacy skill resolver that always selects the canonical Paperclip operational skill when it is available. - Applied the resolver to direct adapter execution, ACPX execution, skill snapshots, and persistent skill sync. - Added Hermes skill materialization at sync and run boundaries. - Aligned Cursor, Gemini, and OpenCode execution-time injection with the configured child `HOME`. - Made Hermes stop execution when another installation blocks the required operational skill. - Kept optional skills controlled by `paperclipSkillSync.desiredSkills`. - Kept `paperclip_runner` on the configurable-only resolver. - Added regression coverage for missing preferences, empty preferences, each skill-capable legacy adapter, ACPX, Hermes, and native runner isolation. - Documented the legacy runtime invariant. ## Verification - `pnpm -r typecheck` passed on the pushed commit. - `pnpm build` passed on the pushed commit. - The adapter utility regression suites passed: 236 tests. - The changed server adapter suites passed: 48 tests across 12 files. - The OpenCode adapter suite passed: 8 tests. - The Hermes adapter suite passed: 7 tests. - `git diff --check` passed. - `pnpm test:run` is not clean on this macOS host. The command reported failures in unchanged workspace and filesystem suites. An isolated rerun of `company-skills.test.ts` and `company-skills-service.test.ts` reproduced 11 failures because macOS resolved `/var/...` paths as `/private/var/...`. The changed adapter suites pass independently. ## Risks - This change deliberately makes the operational skill non-removable for skill-capable legacy local adapters. - Existing agents receive the skill on their next list, sync, or run boundary. No database migration is required. - The resolver does not create a skill when the runtime inventory does not contain the canonical entry. - Hermes aborts a run if another installation occupies the required operational skill target. - Hermes removes only an undesired Paperclip-owned symlink that still points to the known Paperclip source. - The native runner does not receive the legacy default. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact serving model ID and context window were not exposed. The agent used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
821573ede8 |
refactor(adapter-utils): extract the shared workspace-restore teardown factory (#12196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run workspace restore steps when an ACP run ends. > - Claude, Codex, and Gemini each kept a near-identical teardown closure. > - Duplicate closures require the same defect fix in three files. > - This pull request adds one shared workspace-restore teardown factory and keeps each adapter's message strings. > - The benefit is one tested restore-failure path with the same output and outcome for all three adapters. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Claude, Codex, and Gemini ACP adapters restore the workspace during teardown and report restore failures with an allowlisted message. **Subsystem affected** `packages/adapters/` and `packages/adapter-utils/`. **Current behavior** Each adapter keeps a near-identical closure. The closure logs a start line, restores the workspace, classifies errors, and logs a fixed failure line. **Proposed behavior** A shared `createWorkspaceRestoreTeardown` factory owns the common steps. Each adapter passes its staged runtime, log sink, start line, and failure prefix. **Reason and benefit** The shared factory removes duplicate error handling. One tested implementation now preserves the existing output and outcome for all three adapters. **Breaking changes** None. The refactor preserves the emitted lines and returned outcomes. **Additional context** This pull request contains no public issue reference because no related public issue was found. ## What Changed - Add `createWorkspaceRestoreTeardown` to `packages/adapter-utils`. - Move the shared restore, classify, and allowlisted log flow into the factory. - Update the Claude, Codex, and Gemini ACP adapters to call the factory. - Add a table-driven test for all three message pairs. - Keep one end-to-end restore-failure regression test per adapter. ## Verification - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `npx vitest run packages/adapter-utils/src/workspace-restore-teardown.test.ts` - `npx vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts` - `npx vitest run packages/adapters/claude-local/src/server/acp.test.ts` - `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts` - `npx vitest run packages/adapters/gemini-local/src/server/acp.test.ts` - Continuous integration must pass before merge, except for the known pre-existing failures listed in the handoff. ## Risks Low risk. This change moves shared code without changing behavior. The adapter-specific message strings remain unchanged. ## Model Used OpenAI GPT-5, exact model ID `gpt-5`, tool use and code review assistance. The context window size was not provided by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
822e0aed93 |
fix(adapter-utils): move the workspace-restore merge lock to an instance-scoped root and surface restore failures on the run (#12187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters restore sandbox work into project workspaces after a run > - The restore lock used the target workspace parent, which can reject writes > - The teardown then hid restore errors, so a run could report success with lost work > - This pull request moves the lock into an instance-scoped root and reports safe restore failure codes > - The benefit is reliable restore coordination and visible failure evidence without changing run success semantics ## Linked Issues or Issue Description Refs: #10914 ## What Changed - Move the workspace-restore merge lock into a private, instance-scoped root. - Derive the lock key from the canonical target path with SHA-256. - Resolve the lock root from the caller environment and reject unsafe root types. - Classify restore failures with three allowlisted codes. - Add the failure code to run result JSON without exposing a host path or process identifier. - Keep restore failure fail-open for the run exit code and run status. ## Verification - Run `npx vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts`. - Run `npx vitest run packages/adapter-utils/src/acpx-engine/run-fault-matrix.test.ts`. - Run the four Codex credential suites. - Confirm the branch includes the current `master` commit and no manual lockfile edit. - Confirm all pull request checks and the Greptile review reach a terminal green state. ## Risks - The lock path changes for workspace restore and removes the sibling-directory fallback. - A misconfigured or inaccessible instance home can still stop lock setup. - Restore remains fail-open, so callers must inspect the result evidence when a restore fails. ## Model Used OpenAI GPT-5. The model used tool calls and code execution to validate and route an author-provided change. The implementing engineer authored the code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
397de98193 |
feat(runner): add flagged Codex execution adapter (#12188)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner now has protocol, provider, tool, package, persistence, and hidden server boundaries. > - The server still cannot select that path for a real agent heartbeat. > - A new runtime must not change any existing direct adapter. > - An experimental runtime must fail closed when its rollout flag is off. > - This pull request adds one guarded Codex vertical slice through runnerd. > - The benefit is a production-built runner path that users cannot start by default. ## Linked Issues or Issue Description Refs #11962 Refs #12111 Refs #12169 Refs #12176 **Subsystem affected** Cross-cutting. The change affects the runner package, server orchestration, shared settings, and adapter configuration UI. **Problem or motivation** The hidden PRP coordinator cannot execute a real heartbeat. The application also needs an explicit rollout boundary before it can expose the experimental runner. Existing direct adapters must keep their current execution and finalization behavior. **Proposed solution** Add `paperclip_runner` as a Codex-only adapter behind the default-off `enableNativeRunner` instance flag. Select the native runtime only for that adapter. Persist the run binding before runnerd starts. Wait for the durable PRP result and terminal event. Resume the real Codex provider thread on later heartbeats. Keep persisted native runs readable and recoverable after the flag changes. **Alternatives considered** The server could route `codex_local` through runnerd. That option would change an existing adapter and weaken rollback safety. The server could expose all providers now. That option would add unreviewed provider behavior. The build could depend on a prebuilt runner binary. That option would make source builds architecture-dependent and difficult to verify. **Roadmap alignment** This work supports the shipped enforced-outcomes, governed-tool, and self-healing-run milestones. It does not add a new roadmap surface. It is the guarded execution step after the merged hidden runner boundaries. **Additional context** This is the next replacement for the closed large runner pull request. Task-thread presentation remains a separate follow-up so this change can preserve the current direct-adapter UI. ## What Changed - Add `paperclip_runner` as an explicit Codex-only adapter. - Add the default-off `enableNativeRunner` instance flag. - Reject fresh create, hire, import, switch, and execution requests while the flag is off. - Allow edits to persisted runner agents while the flag is off. - Recover an already persisted native run even after the flag is disabled. - Keep every built-in direct adapter on its existing runtime path. - Persist an immutable native run binding and revisioned completion contract before runnerd starts. - Execute server to PRP to runnerd to Codex to server through the hidden coordinator. - Validate the durable result against the terminal event and exact completion criteria before finalization. - Preserve the Codex provider thread ID and use `thread/resume` on the next heartbeat. - Strip unsupported Codex configuration fields from the experimental adapter. - Build a target-native release runner binary from source and vendor it into the server distribution. - Install Rust only in the Docker build stage. Do not add a workflow or lockfile change. - Stop the runner process group on completion, cancellation, and forced shutdown. ## Verification - Run `pnpm --filter @paperclipai/paperclip-runner check:all`. All 69 TypeScript tests and 58 Rust tests pass. Protocol, conformance, replay, formatting, and generated-file checks pass. - Run the 12 focused adapter, settings, runtime-selection, coordinator, direct-isolation, and real Codex integration test files. All 186 tests pass. - The real integration test uses PostgreSQL, HTTP, WebSocket, runnerd, and a fake Codex app server. It proves one `thread/start` followed by one `thread/resume`. - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm check:token-gates`. - Build the Docker `build` target from a clean context. Confirm that the server distribution contains an executable `paperclip-runnerd` built with Debian Rust 1.85. - Start the server through the source-mode tsx entry point with the package `dist` directory absent. Confirm the vendor shim resolves source exports and the server boots. - Run `pnpm test:run` twice. On this macOS host, 405 files pass and 1 file skips. Eight untouched workspace and loopback tests fail because macOS resolves `/tmp` and `/var` through `/private` and because PID-derived test ports exceed 65535. Linux CI must pass the full suite. - Confirm that the diff contains 52 files. Confirm that it contains no `.github` or `pnpm-lock.yaml` change. ## Risks - The feature flag is off by default. A fresh native start fails with a stable error while the flag is off. - A persisted native run remains recoverable after the flag changes. This prevents rollout changes from corrupting recorded work. - Only local Codex execution is accepted. Other providers and remote work modes fail closed. - Existing direct adapters do not start runnerd, create native rows, use native status arbitration, or enter native finalization. - The runner receives its one-use bootstrap ticket through the child environment. The server does not put the ticket in command arguments or logs. - The server validates the company, task, agent, run, runner, session, completion contract, result, and terminal binding before it accepts completion. - The build compiles a target-native Rust binary. Cross-platform release packaging remains a later concern. Source builds and Docker builds compile for their current target. - Docker needs enough build memory for the existing server TypeScript compile. The Docker build stage sets a 4 GB V8 heap limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and applicable tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
79b464bf9d |
fix(server): surface skill materialization failures instead of dropping the skill (#12146)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Runtime skill listing materializes each company skill's files before
handing them to the agent's adapter
> - A materialization failure was swallowed with catch-to-null, and the
skill silently vanished from the runtime while the library still showed
it installed
> - Operators saw "installed", agents saw nothing, and nobody saw the
cause; on claude-local a missing desired skill could even crash the
prompt-bundle hasher
> - This pull request turns both failure paths into structured "missing"
entries with the real error and makes every adapter skip unmountable
entries explicitly
> - The benefit is that a broken skill shows up as broken, with its
cause, instead of not existing
## Linked Issues or Issue Description
**What happened?**
A company skill whose runtime files fail to materialize (deleted source,
missing stored SKILL.md copy, failed version snapshot) disappears from
`listRuntimeSkillEntries` with no trace. Agent skill snapshots report a
generic "not available" with no cause. On claude-local, a desired skill
whose source path does not exist reaches the prompt-bundle hasher, whose
`fs.lstat` throws and can fail the whole run.
**Expected behavior**
The skill appears with `sourceStatus: "missing"` and a `missingDetail`
carrying the underlying error, snapshots and the UI show it as broken,
and adapters skip it at mount time with a logged warning instead of
crashing or dangling-symlinking.
**Steps to reproduce**
Install a local-path skill referenced by an agent, delete its source
directory contents so the stored SKILL.md copy cannot be recovered, and
start a run: before this change the skill vanishes from the runtime set
silently; on claude-local a pinned-but-unmaterializable version can fail
bundle preparation.
## What Changed
- `server/src/services/company-skills.ts` `resolveRuntimeSkillSource`:
both `.catch(() => null)` sites (version snapshot, runtime
materialization) now return the structured `{status: "missing", source,
detail}` shape the deliberate missing branch already used, with the
underlying error message in `detail`.
- `packages/adapter-utils/src/server-utils.ts`:
`isPaperclipSkillSourceMissing` is exported with a doc comment.
- `packages/adapters/claude-local/src/server/execute.ts`: missing
desired skills are filtered out of the prompt bundle and each one logs a
`[paperclip] Warning` with its detail to the run output.
- `cursor-local`, `gemini-local`, `kimi-local`, `opencode-local`,
`pi-local` `execute.ts`: mount loops (and the cursor/gemini injection
calls) skip missing entries instead of symlinking a nonexistent path.
## Verification
- `cd server && npx vitest run
src/__tests__/company-skills-service.test.ts` — new test pins the
missing-with-cause entry for a failed materialization. Nine pre-existing
project-workspace tests in this file fail on my machine at clean
`master` too (environment-specific); their count is unchanged by this
PR.
- `cd server && npx vitest run
src/__tests__/heartbeat-runtime-skills.test.ts
src/__tests__/claude-local-skill-sync.test.ts
src/__tests__/cursor-local-skill-sync.test.ts
src/__tests__/cursor-local-skill-injection.test.ts
src/__tests__/gemini-local-skill-sync.test.ts` — 12 tests pass.
- `cd packages/adapters/claude-local && npx vitest run` — 244 passed, 1
skipped.
- `pnpm run typecheck` clean in server, adapter-utils, and all six
touched adapters.
## Risks
- Runtime skill entry lists grow by the previously dropped entries (now
flagged missing). All shipped consumers either intersect with desired
sets, already handle `sourceStatus: "missing"`, or now skip missing
entries at mount time. The snapshot layer already understood the missing
shape via the `materializeMissing: false` path, so downstream contracts
are unchanged.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking
and tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
63df7ad2b3 |
feat(login): use the login pseudo-terminal for Codex device login and de-Claude the shared channel (#12020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use provider-specific login flows > - Codex device login needs a live pseudo-terminal (PTY), while the shared channel still uses Claude-specific names > - The old streamed-exec path does not provide the prompt transport that Codex needs > - This pull request moves Codex device login to the shared login PTY and removes the dead streamed-exec path > - The benefit is one controlled login transport with fail-closed capability checks and safer credential reads ## Linked Issues or Issue Description **Problem or motivation** Codex device login used a streamed-exec path that did not provide the required prompt transport. The shared login channel also exposed Claude-specific names outside Claude code. **Expected behavior** The host selects a fixed login command from trusted adapter data. Codex login uses the provider login PTY. Providers without that capability fail closed. **Proposed solution** Use a server-controlled session home, create and validate it as a fresh 0700 directory, read credentials from one validated descriptor, and rename shared channel names to the neutral login PTY family. **Alternatives considered** Keep the shared login PTY as the single transport. Do not keep the removed streamed-exec path because it cannot provide the required prompt transport. **Roadmap alignment** This change supports the planned login transport work. It does not add a separate roadmap item. ## What Changed - Route Codex device login through the shared login PTY transport. - Select the login command from a closed internal command key. - Carry a server-controlled session home through the launch contract. - Create and validate the session home as a fresh 0700 directory owned by the login user. - Read the credential file with descriptor-relative, no-follow path walking and final descriptor checks. - Gate the login route and run lease on the provider login PTY capability. - Rename shared channel names to the neutral login PTY family. - Remove the streamed-exec transport value, selector field, driver branch, and related tests. - Hide Codex login in the user interface when the provider lacks the login PTY capability. ## Verification - Server unit suites pass: 89/89. - Adapter-utils suites pass: 262/262. - Codex-local suites pass: 326/326. - Credential-read reader suite passes: 20/20. - Daytona login PTY suite passes: 30/30. - Device-login suites pass: 56/56. - TypeScript checks pass for server, adapter-utils, and UI. - GitHub Actions must pass after pull request creation. - Greptile review must reach 5/5 with no open P2 findings, recommendations, or follow-ups. ## Risks - Providers without a login PTY capability lose Codex login support by design. - The credential read rejects invalid ownership, mode, type, path, and size. - The launch-time sandbox directory race remains outside the threat model because the login runs inside the sandbox and a hostile sandbox already controls its credential. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact context window and reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc42a67e7e |
fix(adapter-utils): extend the duplex fail-closed run disposition to the CLI lane (#11966)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through adapter execution lanes > - Duplex adapters can lose their control channel before a process completes > - The ACP lane already fails closed, but the CLI lane can report false success > - This pull request applies the same completion rule to the CLI lane and shares the loss code > - The benefit is consistent failure reporting when a duplex channel closes during a run ## Linked Issues or Issue Description **What happened?** A CLI-lane duplex run can lose its control channel before clean process completion. The run can then report `succeeded` with exit code 0 and no error code. **Expected behavior** The execution target must fail closed when the channel dies before clean completion. It must return exit code 1, the typed `duplex_channel_lost` error code, and a short stderr note. **Steps to reproduce** 1. Start a duplex adapter run through the CLI execution lane. 2. Close the duplex control channel before the process completes cleanly. 3. Inspect the run result and error code. **Paperclip version or commit** Commit `5e01523d4eb6df4a20a0bddd05374c9c42225203`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Claude Code, Codex, Cursor, Gemini, Kimi, OpenCode, and Pi local adapters. **Database mode** Not database-related. ## What Changed - Add an optional `errorCode` field to `RunProcessResult`. - Add a one-read completion seam to the execution target process options. - Fail closed when a duplex channel dies before clean process completion. - Add `settleRunDisposition()` to atomically read and mark orderly completion. - Share the typed duplex loss error code across the ACP and CLI lanes. - Mark non-success terminal results as orderly completion before teardown. - Wire the seam through the seven duplex adapters. - Add regression tests for channel loss, clean completion, and non-clean terminal results. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` — 118 passed. - `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "sandbox duplex run-disposition seam"` — 4 passed. - The author confirmed a clean type-check for `@paperclipai/adapter-utils` and the seven duplex adapter packages. - Pre-existing environment failures remain outside this change. They include `EACCES mkdir '/srv/paperclip'` and remote file-size setup failures. ## Risks The change alters terminal status for CLI duplex runs that lose control before clean completion. The typed error code and stderr note keep the failure visible. The broker marks failed, cancelled, and timed-out results as orderly completion to prevent false loss events during teardown. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, with the standard GPT-5 context window. The model assisted with the implementation and test work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
10d2781a29 |
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox adapters provide controlled execution for untrusted provider environments. > - The sandbox channel needs one persistent duplex transport with strict host control. > - The transport must remain off unless the instance setting and provider capability both allow it. > - The host must detect loss, bound resource use, and expose only safe telemetry. > - This pull request adds the broker, gated selection, kill-switch wiring, fixed observability, and real-process proof. > - The benefit is safer sandbox execution with bounded failure behavior and inspectable transport results. ## Linked Issues or Issue Description No public issue exists for this change. The related pull requests are #11738 and #11750. **Problem or motivation** The sandbox duplex channel needs a host-controlled broker, strict transport gates, bounded provider input, and safe loss telemetry. Without these controls, a provider can cause replay, resource growth, unsafe endpoint selection, or data exposure through telemetry. **Proposed solution** Add a host broker with nested time limits, request limits, one-shot loss, and per-id deduplication. Select duplex transport only when the instance setting and provider capability both equal true. Assign the endpoint and nonce on the host. Reject invalid readiness data and use the file bridge on failure. Add fixed redacted telemetry and a real-process end-to-end test harness. **Alternatives considered** Keep the file bridge as the only transport. This avoids new channel behavior but does not provide persistent duplex operation for supported sandbox providers. **Roadmap alignment** This change supports the Cloud / Sandbox agents section in ROADMAP.md. ## What Changed - Add the duplex bridge broker with bounded forward, response, and gateway wait budgets. - Bound concurrent requests, lifetime requests, and request-id bytes before retention or forwarding. - Select duplex transport only when both required gates are true. - Assign the loopback port and nonce on the host and enforce a liveness-only READY frame. - Fall back to the file bridge after invalid readiness, contamination, bind failure, or timeout. - Carry the kill switch through the server, acpx engine, and six local adapters. - Add fixed, redacted duplex telemetry with a provider allowlist. - Add a real-process end-to-end harness for readiness, round trips, loss, and teardown. - Add regression coverage for limits, loss, UTF-8 splits, concurrency, and telemetry dimensions. ## Verification - Adapter-utils, server, and Daytona typechecks pass locally. - Adapter-utils tests pass, including the codec, broker, execution-target sandbox, and real-process harness. - Server kill-switch tests pass. - Live Daytona tests pass with the required provider key and skip without that key. - The root pnpm-lock.yaml file has no diff. - The branch contains ten commits after origin/master. ## Risks - Duplex transport remains disabled unless both gates equal true. - A provider remains an untrusted boundary and needs least-privilege credentials and quotas. - The server telemetry recorder stays deferred; the default recorder does nothing. - A provider that pre-binds the host port causes a fail-closed fallback to the file bridge. - The change adds no database migration and changes no root lockfile. ## Model Used OpenAI GPT-5, exact model family GPT-5, large context window, reasoning, and tool use. The model assisted with Git handoff validation and PR preparation. The implementation commits came from the engineering worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
38d8f37172 |
fix(build): enforce Node 24 across Paperclip (#11792)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip runs across the CLI, server, adapters, plugins, CI, and container images. > - These surfaces declared different Node.js versions from 20 through 24. > - A newer `@types/node` major can expose APIs that the supported runtime does not provide. > - Node.js 20 is no longer a suitable project baseline, and Node.js 24 is the current LTS line. > - This pull request sets Node.js 24.11.0 as one repository-wide baseline, adds a drift check, and gives users actionable startup guidance when their runtime is too old. > - The benefit is one clear runtime contract for development, release, installation, and published packages. ## Linked Issues or Issue Description Refs #2734 Refs #11727 Refs #739 ## What Changed - Require Node.js 24.11.0 or newer in all 42 package manifests and runtime checks. - Use Node.js 24 in GitHub Actions, Docker images, smoke images, sandbox setup, portable installs, and esbuild targets. - Align every direct `@types/node` declaration on `^24.0.0`. - Prevent Dependabot from opening major `@types/node` upgrades without a matching runtime decision. - Add `.nvmrc` and a CI policy check for Node version drift. - Update ACP version gates, tests, and user documentation for the new minimum. - Print a non-blocking warning on CLI and server startup when Node is unsupported, with remediation through a version manager or the documented downloaded `install.sh` workflow. - Deduplicate that warning when `paperclipai run` boots the CLI and server in the same process. ## Verification - `node scripts/check-node-version-policy.mjs` - `node --check scripts/check-node-version-policy.mjs` - `node --check cli/esbuild.config.mjs` - `node --check scripts/generate-npm-package-json.mjs` - `bash -n scripts/install.sh scripts/test-install-sh-docker.sh scripts/e2e-install-lifecycle.sh` - Parsed all 42 package manifests and confirmed `engines.node` is `>=24.11.0`. - `git diff --check` - `vitest run packages/adapter-utils/src/sandbox-install-command.test.ts` passed with 3 tests. - `vitest run cli/src/node-version.test.ts` passed with 4 tests. - Directly exercised the shared warning helper for unsupported-version messaging and same-process deduplication. - The focused exe.dev suite could not resolve the locally unbuilt plugin SDK from this isolated worktree. A full offline workspace install was also blocked because the package-manager signature verifier requires registry access. The full suite was not run locally; draft CI performs a clean install and evaluates the wider impact. ## Risks - This is a breaking runtime change for users, plugins, and deployments that still use Node.js 20 or 22. - Published workspace packages will now produce an engine warning or failure in strict package managers on older Node.js releases. - Node.js 24 can reveal dependency, native module, Playwright, or agent CLI compatibility issues in CI. - The bootstrap installer now installs Node.js 24 when the current runtime is older than 24.11.0. - The portable sandbox fallback is pinned to Node.js 24.11.0 and depends on that upstream tarball remaining available. - Unsupported runtimes continue booting after a warning, so a later incompatibility can still fail at its point of use. - The CLI and server share the warning policy through the published `@paperclipai/shared` package; packaging checks must keep that subpath export available. - This PR does not commit `pnpm-lock.yaml` because repository policy assigns lockfile generation to CI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5. The exact deployment ID and context window are not exposed in this session. Reasoning, repository tools, shell execution, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e0b64529b3 |
feat(auth): normalize agent login in the sandbox onto one session table and a capability contract (#11730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandbox agents need a safe login path for each supported adapter > - Codex device login and Claude setup-token login used separate session stores and route logic > - Separate stores made session lookup, expiry, and login capability checks harder to keep consistent > - This pull request unifies both flows on one session table and one capability contract > - The benefit is one company-scoped login model with public session identifiers and shared lifecycle rules ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Codex and Claude sandbox login used separate session stores and different route paths. This split increased the risk of inconsistent company scoping, session lookup, and cleanup. **Proposed solution** Use `adapter_auth_sessions` for both login flows. Use public session identifiers for API access. Select login behavior from projected adapter capability data. Share the route spine, lease arguments, runner lifecycle, and reaper rules. **Alternatives considered** Keep two session tables and add matching fixes to both routes. This keeps duplicate logic and does not provide one capability contract, so this pull request uses shared infrastructure. **Roadmap alignment** This change supports the shipped Cloud / Sandbox agents milestone in `ROADMAP.md`. ## What Changed - Unify Codex device login and Claude setup-token login on `adapter_auth_sessions`. - Return and look up sessions with company-scoped public session identifiers. - Enforce one active session for each company, owner, and adapter. - Share the login route spine, sandbox lease arguments, runner lifecycle, and missing-auth check. - Add a standalone setup-token reaper with adapter-specific row selection. - Add optional login capability projection for adapters and drive route and UI selection from that data. - Rename the provider flag to `supportsLoginPty` and validate its deprecated alias. - Remove the old Claude setup-token session table and add the required migrations. ## Verification - Server typecheck passed with `tsc`. - Database typecheck passed. - UI typecheck passed with `tsc -b`. - Codex login service and route suites passed. - Setup-token session, route, and reaper suites passed. - Adapter session schema, plugin validator, capability projection, UI render, and Daytona suites passed. - GitHub Actions must confirm the complete CI gate after pull request creation. ## Risks - The migrations remove short-lived in-flight login rows during deployment. A login that spans the migration can continue until its provider lease expires. - The Codex credential store remains company-scoped. A cross-owner credential race remains a documented, board-accepted risk. - API clients that use internal session row identifiers no longer work. The API accepts only public session identifiers. ## Model Used Codex, GPT-5, exact runtime model ID not exposed in this handoff, large context window, reasoning, and repository tool use. The implementing engineer produced the code with AI assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
48f4ae16ac |
fix(codex-local): keep a promoted device-login credential when re-seeding the managed home (#11578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The codex_local adapter supports a device login that runs in a trusted sandbox and promotes the credential into the per-company managed Codex home > - The managed-home re-seeding step treats every regular-file `auth.json` as apikey-mode residue and removes it so the shared-home symlink can be restored > - The promotion writes the company credential as a regular file, so the first environment Test or run after a successful login deletes it > - On a server with no shared Codex login — any containerized deployment — nothing replaces the file, and the UI reports that the sandbox has no ready authentication right after it reported a successful login > - This pull request makes the cleanup identity-anchored: a subscription credential whose identity the shared source does not hold survives re-seeding > - The benefit is that a device login stays usable after Test and runs, on hosts with and without a shared Codex login ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the bug template. Related public PRs: [#11237](https://github.com/paperclipai/paperclip/pull/11237) added the sandbox device login and the credential promotion, [#11097](https://github.com/paperclipai/paperclip/pull/11097) added its building blocks, and the `ensureSymlink` heal for stale copies came from the fix for #5028. [#9621](https://github.com/paperclipai/paperclip/pull/9621) touches the adjacent sandbox auth sync-back lane but not this defect. **Subsystem affected** packages/adapters/codex-local — managed `CODEX_HOME` seeding (`codex-home.ts`). **Current behavior** A successful device login promotes the subscription `auth.json` into the company Codex home as a regular file, and the UI reports the login as authenticated. The next `seedManagedCodexHome` call — the environment Test probe and every execute both run it — removes any regular-file `auth.json` when no API key is configured, because the cleanup assumes such a file is apikey-mode residue left by a previous run. It then symlinks `auth.json` from the shared source home. On a server whose shared home has no Codex login (a container image, for example), there is no source to symlink, so the home ends with no credential at all. The Test probe then reports "The sandbox has no ready authentication for this adapter" immediately after a successful login, and a fresh login repeats the same cycle. On a server whose shared home does hold a login, the symlink silently replaces the promoted account with the host account. **Expected behavior** The credential a device login promoted stays in the company home across Test probes and runs. The #5028 heal (a stale regular-file copy of the shared credential becomes a symlink to the live source) and the apikey-residue cleanup keep working. **Steps to reproduce** 1. Run the server in an environment whose shared Codex home (`$CODEX_HOME` or `~/.codex`) has no `auth.json`. 2. Complete a Codex device login for a company; the promotion writes the company home `auth.json` and the UI reports authenticated. 3. Click Test on a codex_local agent (or start a run). The probe reports no ready authentication, and the promoted `auth.json` is gone from the company home. **Proposed solution** Make the cleanup identity-anchored, the same rule the promotion and the cache vend already use. A regular-file `auth.json` survives re-seeding when it holds a usable subscription identity that the shared source does not also hold, and the shared symlink does not replace it. A same-identity regular file is still the #5028 stale copy and is still healed into the symlink, because the symlink serves the same account with live, rotating tokens. An apikey-mode or unreadable file is still removed. ## What Changed - `seedManagedCodexHome` reads the target `auth.json` before the cleanup and keeps it when `readSubscriptionAccountId` yields an identity the shared source `auth.json` does not hold. The kept file is excluded from the shared symlink pass, and the function logs a fixed line when it keeps the file. - The function doc comment states the kept-promoted-credential rule. - Four new `seedManagedCodexHome` test cases: a promoted credential with no shared auth, a promoted credential with a different shared identity, the same-identity #5028 heal, and apikey-mode residue removal. ## Verification ```sh cd packages/adapters/codex-local npx tsc --noEmit # clean npx vitest run # 28 files, 321 passed, 1 skipped ``` The four new cases fail on the previous code: the first two observed the promoted file deleted (and, with a shared login present, replaced by the shared symlink). ## Risks Low risk. The change narrows one deletion path. Deployments that never use the device login see no difference: without a promoted subscription file, the cleanup and the symlink behave exactly as before, and the #5028 heal is pinned by an existing test plus a new same-identity test. The one deliberate behavioral shift: after a device login, the promoted company credential now stays authoritative over the shared host login for that company — which is the promotion's documented contract ("the company credential slot"). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
1f7959bc69 |
fix(codex-local): skip benign stderr warnings when deriving the fallback run error (#10003)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents execute through adapters; the codex_local adapter runs the Codex CLI and reports each run's outcome, including an error message when the CLI exits nonzero > - When no error can be parsed from the CLI's JSONL output, `toResult` in `packages/adapters/codex-local/src/server/execute.ts` falls back to the first non-empty stderr line as the run error > - The adapter itself passes the approvals-bypass flag, so the CLI's first stderr line is always the benign startup warning "YOLO mode is enabled. All tool calls will be automatically approved." > - Failed runs therefore record that warning as their error, hiding the real cause (for example an OpenAI API error further down in stderr) and making failures hard to diagnose from the run record > - This pull request derives the fallback error from the first meaningful stderr line, skipping a conservative set of known benign lines, and keeps the existing behavior when every line is benign > - The benefit is that failed Codex runs surface the actual failure reason instead of a harmless startup warning, without ever producing an emptier message than before ## Linked Issues or Issue Description No public issue exists for the codex_local case. The same bug class was fixed for gemini-local in Refs #5099 and Refs #3476; this PR applies the equivalent fix to codex_local. **What happened?** On a multi-tenant cloud deployment of Paperclip, several codex_local runs failed and their run records showed `error_code=adapter_failed` with the error text "YOLO mode is enabled. All tool calls will be automatically approved." That is a benign Codex CLI startup warning, printed on every run because the adapter passes the approvals-bypass flag itself. The real failure (an OpenAI API error printed later in stderr) was never surfaced. **Expected behavior** When the Codex CLI exits nonzero and no error was parsed from its JSONL output, the run error should be the first stderr line that actually explains the failure, not a startup warning the adapter itself provoked. **Steps to reproduce** 1. Configure a codex_local agent and make the underlying Codex CLI invocation fail after startup (for example, configure a model id the active credentials cannot use). 2. Run the agent so the CLI exits nonzero with no parsed JSONL error. 3. Inspect the run's error message: it shows the YOLO approvals warning (the first stderr line) instead of the real error printed further down in stderr. ## What Changed - Added `firstMeaningfulStderrLine` next to `firstNonEmptyLine` in `packages/adapters/codex-local/src/server/execute.ts`, with a conservative benign-line predicate covering the YOLO approvals warning and `[paperclip] ...` diagnostic lines the adapter injected (for example ACP fallback notes). - Used it only in the `toResult` fallback error derivation. If every stderr line is benign, the existing chain still applies (first non-empty line, then `Codex exited with code N`), so the message never gets emptier than today. Logging is unchanged. - Added `packages/adapters/codex-local/src/server/execute.stderr-error.test.ts`: four end-to-end cases through `execute()` with a mocked CLI process, plus unit coverage for the new helper. Tests were written first and confirmed failing before the fix. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/server/execute.stderr-error.test.ts` (7 tests pass; 5 failed before the fix as expected) - `pnpm exec vitest run packages/adapters/codex-local` (21 files, 188 tests pass) - `pnpm run typecheck` in `packages/adapters/codex-local` (clean) ## Risks Low risk. Only the derived fallback `errorMessage` changes, and only when a benign line would otherwise have been picked; parsed JSONL errors, logging, retry/quota/auth classification inputs, and the empty-stderr exit-code fallback are untouched. The benign-line list is deliberately conservative (exact prefixes) so real errors are never skipped. ## Model Used Claude Fable 5 (claude-fable-5), extended thinking, via Claude Code ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e5a7fd7038 |
Add sandbox device-login for the Codex adapter (#11237)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Agent adapters connect Paperclip to tools such as the Codex command line tool. > - A sandboxed Codex agent may start without a credential. > - The operator needs a safe sign-in flow that does not expose credentials to the shared package or the sandbox. > - This pull request adds a company-scoped device-login flow with a temporary Daytona sandbox. > - The flow promotes the credential only after readiness checks pass and removes the temporary sandbox after use. > - The result lets an operator sign in to a sandboxed Codex agent from the agent form. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** A Codex adapter that runs in a sandbox cannot authenticate when the company has no pre-provisioned Codex credential. **Proposed solution** Add a company-scoped device-login session. Start a temporary sandbox, run `codex login --device-auth`, stream the code and URL, verify readiness, promote the credential, and delete the sandbox. **Alternatives considered** Pre-provisioning a credential does not support first-time sandbox login. Keeping the credential in the login sandbox does not provide a durable company credential. **Roadmap alignment** This supports the roadmap item for cloud and sandbox agents. **Additional context** The flow uses a five-minute cleanup reaper, compare-and-set status changes, and a PostgreSQL advisory lock to protect promotion and cleanup. ## What Changed - Add the adapter login-session contract, database table, and migration. - Add company-scoped server routes and a service for sandbox device login. - Add credential promotion, readiness checks, and cleanup after login. - Add restart-safe cleanup for abandoned login sandboxes. - Add sandbox login controls to the agent creation and edit forms. - Keep device-login and vendor identifiers out of public shared and adapter UI symbols. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local exec vitest run` passed with 310 tests at the submitted commit. - The server login route, service, and reaper tests passed with 45 tests at the submitted commit. - The agent form render tests passed with 26 tests at the submitted commit. - The public-symbol leak check passed at the submitted commit. - A live Daytona sign-in flow still requires confirmation by a user with a live sandbox. ## Risks The migration adds a new company-scoped table. A promotion or cleanup race could remove a credential or leave a sandbox active, so the service uses claims, compare-and-set transitions, and an advisory lock. The live Daytona flow needs operator confirmation because local tests do not provide a real browser sign-in. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, extended reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
459469638e |
feat(adapter-codex-local): add secure device-login building blocks (#11097)
## Thinking Path > - Paperclip connects AI agents to local and remote runtimes. > - The Codex local adapter needs a safe device-login flow. > - A future sandbox integration needs strict prompt validation, secret protection, cleanup, and private credential storage. > - This pull request adds tested building blocks for that flow. > - The result gives a later Daytona integration a clear security boundary. ## Linked Issues or Issue Description No public issue covers this change. **Problem or motivation** The Codex local adapter has no safe, reusable flow to prove device login inside an isolated sandbox. **Proposed solution** Add parser, runner, credential export, and proof helpers. Validate the prompt, protect login data, store credentials in a private run-scoped home, and dispose all sandbox resources. **Alternatives considered** Do not connect a production Daytona driver in this change. Use an injected sandbox driver and focused tests first. This keeps the security controls testable before live provider integration. **Roadmap alignment** The change extends the Codex local adapter. It does not add a core Paperclip route or duplicate a planned core feature. **Additional context** The flow keeps the login URL, code, and token out of logs, results, and errors. The proof home uses a company-scoped root and a run-scoped private directory. ## What Changed - Add a pure parser for the exact Codex device-login URL and one-time code shape. - Add a sandbox runner with prompt handling, timeout, cancellation, and disposal. - Add a credential export step with company scoping, path checks, payload checks, private modes, locking, and cleanup. - Add redacted device-login fixtures and focused tests for parsing, secret redaction, runner outcomes, credential export, and cleanup. ## Verification - Run `pnpm --filter @paperclipai/adapter-codex-local exec vitest run`. - Run `pnpm --filter @paperclipai/adapter-codex-local exec tsc --noEmit`. - Review tests for strict URL and code validation, timeout, cancellation, disposal, secret redaction, path safety, payload safety, file modes, and cleanup. ## Risks - This change provides building blocks, not a live Daytona proof. - A later integration must connect the runner to a concrete sandbox driver. - Credential export depends on existing Codex authentication cache helpers. - Incorrect path or payload assumptions can reject valid credentials. ## Model Used Codex, GPT-5, tool use, code execution, and repository review. The Paperclip runtime controls the exact context window and reasoning mode. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |