mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 05:41:56 +02:00
codex/plugin-task-execution
518
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f12b647ae8 |
fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can collect more messages while its agent works. > - Legacy runners must stop the active process before they can receive those messages. > - The old Interrupt action cancelled the run but could leave the queue idle and hidden. > - Codex could also classify a cancelled run as successful or start a fresh process after cancellation. > - This pull request joins cancellation, preserves the provider session, and dispatches the current queue after cleanup. > - The benefit is reliable interruption with the saved message order, edits, and deletions. ## Linked Issues or Issue Description **What happened?** Interrupt could strand a legacy message queue. The UI could hide pending messages after the run stopped. A Codex signal exit could race the cancellation write. A stale session warning could also trigger a fresh process after an interrupted resume. **Expected behavior** Interrupt stops the active turn and sends the remaining messages once, in their saved order. Deleted messages stay deleted. An interrupted Codex turn keeps its session and does not restart itself. **Steps to reproduce** 1. Assign a task to a legacy Codex agent that runs a long command. 2. Queue three messages. Edit one, discard another, and move the last message first. 3. Click Interrupt in the queue. 4. Repeat the interruption while the resumed session runs another command. Related work: Refs #13160, which moves native queue steering into the wake-queue module. This change fixes legacy interruption and keeps native steering unchanged. ## What Changed - Add a revision-checked, company-scoped endpoint for legacy queue interruption. - Promote only the requested queue after the provider stops and releases its lease. Retry its persisted interrupt intent from the scheduler after a promotion error or server restart. - Keep pending legacy queues visible after a run stops. Use server state for the interrupt result. - Serialize owned process cancellation before classifying the adapter result. Preserve late session and log metadata. Acknowledge cancellation only when an actual process or process group was owned; scheduler placeholders retain their normal release policy. - Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the session has started. - Add cancellation race, multi-actor queue order, durable retry, resume fallback, and stale request regression tests. Document the behavior. ## Verification - Real browser tests passed with legacy Codex CLI and ACP engines, using Codex 0.153.4 and gpt-5.6-sol. - All three automated ACP browser scenarios passed locally: immediate Interrupt delivery, no replay of an unfinished write, and pause requiring Resume. Updated the old test expectation that required a separate “go” after Interrupt. - Browser tests covered queued edits, deletion, reordering, deleting the final message, and repeated interruption. - Two consecutive CLI interrupts kept one provider session. Both stopped processes exited. The final message arrived once. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed. - All 346 post-review scheduling, recovery, queue-route, archived-company, worktree-suppression, and stale-queue regression tests passed. - All 318 process-recovery and durable-chat tests passed after the final cancellation guard. - Codex adapter, queue UI, issue-page, and OpenAPI contract tests passed. - `pnpm build` passed. - Full local suite coverage completed with `PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI shards: 618 general server suites, all 145 serialized server suites, and all workspace groups. Every failing suite passed a targeted rerun after the fixes, rebuilding the native test fixture, correcting macOS temporary-path setup, or retrying setup/timing failures. Existing skips remain. - The original monolithic run reported failures before the final fixes; its failed suites were rerun rather than rerunning all 618 suites again. The final process-recovery/durable-chat regression run passed all 318 tests. - All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: [run 34654820774, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2), including typecheck, build, all test shards, E2E, and canary. The signoff and Cursor sandbox tests each hit a timeout in the initial attempt; both suites passed locally, and both failed shards passed their single CI rerun. All three corrected ACP browser scenarios passed in CI. - Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no open review threads. ## Risks Cancellation order affects local adapters. The tests cover signal exits, graceful exits, adapter exceptions, termination errors, and cancellation write errors. Embedded adapters keep their cancellation controls. Ordinary run cancellation and task pause keep their distinct queue policies. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code execution. The exact serving model ID and context-window size are not exposed in this session. The live test runner used OpenAI gpt-5.6-sol. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
87b3e5fc61 |
fix(adapter-utils): stage selected skills into the sandbox for a remote Claude ACP run (#13196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A user selects skills for an agent, and the host materializes those skills into a bundle the agent reads > - An agent can run in a remote sandbox, where the host must stage every file the agent needs > - On the Agent Client Protocol lane the host built that bundle and then named its host path in the prompt, but it never staged the bundle into the sandbox > - The agent therefore read a path that does not exist inside the sandbox, and the run failed on the missing skill file > - The command-line lane of the same adapter already stages a `skills` asset and reads the in-sandbox directory back from the staged runtime > - This pull request carries that proven pattern to the Agent Client Protocol lane, so a selected skill reaches the agent in a remote run ## Linked Issues or Issue Description No public issue exists for this change. The description follows. **What happened?** A remote run of the Claude adapter on the Agent Client Protocol lane could not read any selected skill. The host builds the skill bundle in its own state directory, then writes that host path into the prompt as `Skill root: <path>`. The remote seam of that lane staged one asset only, the configuration seed. It staged no skills asset, so no skill file crossed into the sandbox. The agent then tried to read the skill file at the host path, and the read failed with a missing-file error. **Expected behavior** A remote run receives the skills the user selected, and the prompt names the directory that holds those skills inside the sandbox. **Steps to reproduce** 1. Select one or more skills for an agent that uses the Claude adapter. 2. Start a run for that agent in a remote sandbox on the Agent Client Protocol lane. 3. Ask the agent to read the skill file at the path the prompt names. The file is not there. **Agent adapter(s) involved** The Claude local adapter, on its Agent Client Protocol lane. The shared engine in `packages/adapter-utils` carries the prompt rewrite. **Additional context** The command-line lane of the same adapter already stages a `skills` asset and remaps onto the staged directory. This change reuses that mechanism instead of adding a new transport. One other adapter shows the same host-path shape on its own Agent Client Protocol lane. That lane is tracked separately and this pull request does not change it. ## What Changed - Return the host skill bundle directory from the Claude skill runtime step, and carry it through the remote managed-home context to the staging seam. The value is null for a non-Claude agent, for a run that selects no skill, and for a run whose selected skills all fail to materialize. - Stage that bundle as a `skills` asset on the Claude Agent Client Protocol remote seam, and only when the run selected a skill. - **Stage that asset with `followSymlinks: false`.** The bundle holds an owned copy of each selected skill, and the copy step never copies a symbolic link at the root or at any depth. So the bundle contains no symbolic link, and staging has none to follow. Refusing to follow one also stops a link planted in the bundle directory after the copy from pulling an unrelated host file into the sandbox. A regression test walks the real adapter sources and pins the reviewed `followSymlinks` value at every skills staging site, so a new or changed site fails the test. - **Drop a skill whose staged copy has no usable `SKILL.md`** from the prompt, the skill identity, the command notes, and the bundle, and log which skill was dropped and why. Without this, a skill whose copy failed, or whose `SKILL.md` is a symbolic link the copy step skips, stayed advertised in the prompt while its file was absent — the same missing-file symptom this change exists to fix. - Rewrite the `Skill root:` prompt line, the skill identity, and the command notes onto the in-sandbox directory. The rewrite runs in the engine, after the workspace placement returns the staged runtime. A compatible session resume reuses the cached staged runtime, so the rewrite runs on that path too. - Keep the session fingerprint on the host-independent skill identity. A change to the selected skill set still invalidates a warm session, and the volatile sandbox path stays out of the hash. - A local run, and a run with no selected skill, keep their current behaviour. ## Verification - `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` — 28 of 28 pass. - `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` — the new engine tests and the staging-site tests pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`, and the same check on the adapter package — both exit 0. - The end-to-end test drives the lane against a local sandbox stand-in. It reads the skill root out of the prompt the runtime received, and then opens the skill file at that path. That is the reported symptom, proved closed. - The new tests carry a sensitivity control. Restoring only the production files to their previous content fails 7 of the 9 new tests. The other 2 do not depend on production code: one is a parser unit test for the source scanner. ## Risks Low risk, and the change is a two-way door. A revert restores the previous behaviour exactly. - **Scope.** The change touches one adapter lane. It does not change the local lane, and it does not change any other adapter. No existing staging site changes its `followSymlinks` value. - **The staged bundle and the workspace.** The staged skills land under the runtime directory inside the workspace. The workspace restore excludes that whole runtime directory, so the staged skills never return to the host worktree. A test proves the exclusion end to end. - **Session reuse.** The rewritten path never enters the session fingerprint, so it cannot invalidate a warm session, and a compatible resume applies the same staged path. - **Direction of data.** Files move from the host into the sandbox only. The change adds no path that writes sandbox content onto the host. - **A dropped skill.** A skill with no usable `SKILL.md` is now absent from the prompt instead of named but unreadable. The run logs the skill and the reason, so the cause is visible. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking and tool use, through Paperclip agents. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` passes — 28 tests. - [x] `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` passes. - [x] `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - [x] All continuous-integration gates are green. - [x] The automated review reports no open finding against the current head. ## Required Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant bug report template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect this change - [x] I have considered and documented the risks above - [x] All continuous-integration gates are green - [x] The automated review score is 5/5 with no open current-head findings - [x] I have addressed every reviewer comment that applies to the current head **Note on the branch history.** This branch first carried a different change: a filename-based admission filter that refused to stage files such as `.env` from a skill directory, together with a switch from symbolic-link bundles to copied bundles. That approach was rejected and **reverted** on this branch. It does not match the documented trust boundary, because the host already delivers credentials into the sandbox on purpose, and replacing the symbolic-link bundles broke live editing of a skill. The revert is in this branch's history. The file that work changed, `packages/adapter-utils/src/server-utils.ts`, is byte-for-byte identical to `master` here and is not part of this diff. Earlier review findings that name that file target the reverted code. All of them are resolved, and the automated review passes on the current head. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bce976d60d |
feat: bind an agent to a Codex login whose account differs from the company default (#13067)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter signs agents in to OpenAI, and a company keeps one default Codex identity in its shared company home > - A login with a DIFFERENT account than the company default is deliberately kept out of the shared home — one agent's sign-in must not switch every unbound agent's credentials — but that left the cross-account login inert: nothing connected the agent the operator was configuring to the credential the login stored > - The stored credential and its company secret already exist; only the last mile — an agent actually using them — was missing > - This pull request reports a non-secret binding claim on the authenticated login and lets the agent page bind that one agent's `CODEX_HOME` to the account's secret, exactly and only when the identities differ > - The benefit is that multi-account Codex becomes one click on the agent that needs it, with company-wide identity untouched ## Linked Issues or Issue Description **What happened?** On an agent's detail page, "Sign in with Codex" using a different OpenAI account than the company default succeeds but changes nothing for that agent. The credential lands in the per-identity store and a company secret names it, but the agent keeps using the company default. The Test keeps reporting that authentication is needed, and no repeat login helps. **Expected behavior** When the operator deliberately signs an agent's page in with a different account, that agent starts using that account. Agents that were not part of the action keep the company default. A same-account login keeps working through the shared company home with no per-agent pinning. **Steps to reproduce** 1. Configure a company whose Codex home holds account A. 2. Open a `codex_local` agent's detail page with a sandbox environment and complete "Sign in with Codex" using account B. 3. Press Test. Before this change the agent still resolves account A and the authentication-needed check returns. ## What Changed - `packages/adapters/codex-local` — the prerequisite shield: `isCodexAuthCachePath` recognizes per-identity credential-store entries, and `seedManagedCodexHome` refuses to symlink, heal, or API-key-overwrite an entry's `auth.json`. The seeding pass runs before every probe and execute; without the shield, an agent bound to an entry would have its stored login silently swapped for the host credential. Static shared config files still copy in. Rotation already survives binding: the sandbox copy-back writes rotated credentials into the identity-keyed store slot. - `server` — the promotion records whether the company default home ended on a different account than the login (any read failure degrades to `false`, so the client can never be told to bind wrongly). After the terminal commit, the routes layer remembers a non-secret claim — the opaque account-home secret id plus that verdict — in a bounded in-memory map, and merges it into the owner read of an `authenticated` `codex_local` session. A restart drops the claim; the panel then shows plain success. - `packages/shared` — `CodexAccountBindingClaim` on the owner session response. It carries no account identifier and no credential byte. - `ui` — the login panel reports the claim upward once. The edit-mode form binds the agent's `CODEX_HOME` to the secret and saves in one step, only when `companyIdentityDiffers` is true. Same-account logins bind nothing on purpose: the company-home refresh already carried them, and an unbound agent keeps following the company default across rotations. Create mode is unchanged. ## Verification - Adapter suite: 381 passed, 1 skipped (includes the new store-entry shield tests and the path-predicate cases). - Server suites (8 files): 130 passed, 15 skipped — including two new route tests that drive a login to `authenticated` and assert the claim with both identity verdicts. - UI render suite: 85 passed — including a panel test that the claim is reported upward exactly once. - `tsc --noEmit` clean in `packages/shared`, the adapter package, and `ui`; `server` clean for the touched file. ## Risks - The bind changes one agent's configuration through the normal agent-update patch, initiated by the operator's own login on that agent's page. The failure direction of every fallback is "offer nothing": a missing claim, a restart, or an unreadable company home all degrade to no bind. - The seed shield narrows what the seeding pass may touch; homes outside the credential store behave exactly as before, covered by the existing seed tests. - Builds on the sign-in credential-resolution fix (#13064), now merged; this branch is rebased onto master and the diff contains only the binding feature. Supersedes #13066, which GitHub auto-closed when its stacked base branch was deleted on merge. ## Model Used Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use in Claude Code (terminal). **Related PRs (searched; no duplicates found):** #12740, #12082, and #9621 touch adjacent Codex credential sync paths; #8495 is the standing hardening effort for probe auth seeding. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no standalone docs cover this flow; the behavioral contracts are documented in-line at each changed site) - [x] I have considered and documented any risks above |
||
|
|
2991a59b17 |
fix(adapters): prevent engine fallback and preserve usable runtime defaults (#13105)
## Thinking Path > - Paperclip manages agents that must write work and report task outcomes through its API. > - Local adapters select an execution engine and its permission settings. > - A higher ACP Node requirement can make an unchanged installation lose access to its default engine. > - The adapter then silently selects CLI, which can change permissions and block API access. > - This pull request keeps the engine choice fixed and reports missing prerequisites before work starts. > - It also gives explicit Codex CLI runs usable defaults and keeps managed services on a supported Node runtime. ## Linked Issues or Issue Description Refs #12215. Related changes: #11792 raised the Node requirement; #13094 addressed separate runner networking behavior. This change fixes the engine-selection and managed-launcher paths. **What happened?** An unchanged agent could switch from ACP to CLI after an upgrade. Codex CLI then used read-only permissions with networking disabled. The run could finish without updating its task. Repeated recovery attempts used the same unavailable setup. Managed updates also skipped the Node check and did not refresh old launchers. **Expected behavior** An unavailable engine must fail with a clear setup error. It must not silently select another engine. Explicit CLI runs must be able to write workspace files and call the API unless the operator configures stricter settings. Managed updates must validate Node and keep child tools on that runtime. **Steps to reproduce** 1. Run an ACP-default agent under Node 22 after the ACP minimum rises to 24.11. 2. Leave the engine unset and disable the approval/sandbox bypass. 3. Observe the old adapter select CLI and fail to write task disposition through the API. 4. Start a managed service with an old launcher and a supervisor PATH that selects a different Node for child tools. ## What Changed - Remove automatic engine fallback for Codex, Claude, Gemini, and Kimi. Check prerequisites for default and explicit ACP selections. - Return a configuration error with proof that provider work did not start. Stop automatic continuation retries for this error. - Enable Codex ACP workspace networking at the actual turn boundary. Upstream mode presets otherwise force it off even when config.toml enables it. Preserve explicit network denial and read-only mode. - Set workspace-write and network access defaults for explicit Codex CLI runs. Preserve explicit sandbox modes, profiles, and network restrictions. - Pin the validated Node directory in managed launcher PATH. Refresh legacy launchers during installs and npm/Git updates. - Reject updates on unsupported Node. Keep update checks, dry runs, and rollback available. - Synchronize the qualified Codex ACP executable identity across server, TypeScript runner, Rust runner, and provider-pack launch paths. - Add regression tests and update engine and installation documentation. ## Verification - [Full CI passed on the final head](https://github.com/paperclipai/paperclip/actions/runs/34387099695): typecheck, build/native runner verification, all general and serialized test shards, all browser shards, release registry, canary dry run, and policy checks. - Greptile: 5/5 on `2c1d6e2815830a5cd39e36c8a082cc0c4441b6c0`, with no unresolved review findings. Security gates are green. - Full workspace typecheck and build also passed locally. The final deployed Linux build passed. - Full Codex, Claude, Gemini, and Kimi source test suites: 804 passed, 2 skipped. Installer, updater, and launcher tests: 47 passed. Installed ACP turn-boundary tests: 3 passed. ACP packaging tests: 14 passed. Focused recovery classification tests also passed. - Real Linux Codex CLI runs, both fresh and resumed, wrote a workspace file and reached the control-plane health API with the new defaults. - Explicit read-only and network-disabled control probes retained those restrictions. - A real ACP run on the final deployed Linux build wrote a file and reached the control-plane API with HTTP 200, without engine fallback. The same probe failed DNS before the turn-policy patch. - Executable-identity and installed-policy contracts: 12 passed. Affected native server tests: 197 passed. Runner factory tests: 21 passed. Rust qualification and native provider integration tests: 11 passed. - Deployed the production changes to a Linux service on Node 24.20 after a verified database backup. Health, bootstrap readiness, static UI, executable/cwd identity, and guarded restart checks passed. The restart lost no runs. - Corrected stale Kimi skill-default and Gemini remote-archive fixtures; both suites pass. ## Risks - Default or legacy auto engine settings now fail when ACP is unavailable. Operators who intend to use CLI must select it explicitly. - Codex CLI now permits workspace writes and networking by default, and ACP workspace-write turns permit networking by default. Explicit operator sandbox settings remain authoritative. - Old managed launchers keep their pinned Node until they are reinstalled under a supported runtime. An old updater cannot repair itself; the documentation gives the current installer command. - Custom service wrappers and global/source installations must configure their runtime PATH. No database migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, shell execution, and test tools. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fe5e68d7a5 |
fix: make Codex sign-in and the environment test agree on the credential a run uses (#13064)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The `codex_local` adapter signs agents in to OpenAI with a device-code login, and the agent page has a Test button that probes the sandbox with the credentials a real run would use > - The login stored its credential where the Test never looked: the company Codex home kept an old shape-valid credential, so the Test failed with "authentication needed" right after a successful sign-in > - The Test also staged a different Codex home than a real run resolves, so the Test and real runs could disagree in both directions > - This pull request makes the login, the seeding pass, and the Test probe agree on one credential resolution > - The benefit is that a sign-in from the agent page fixes the Test on the next click, and a green Test means the same thing a real run experiences ## Linked Issues or Issue Description **What happened?** Sign in with Codex works during onboarding but not on the agent detail page. The operator completes the device-code login. The panel reports success. The Test button still reports that authentication is needed. No number of repeat logins changes the result. Three defects combine to cause this: 1. The device-login promotion only wrote the company default Codex home when that home held no shape-valid credential. A stale credential (for example a symlink to an old host login) blocked the write forever, so the fresh login stayed invisible to the Test. 2. The seeding pass that runs before every probe and execute replaced a same-identity regular-file `auth.json` with a symlink to the host credential, with no freshness comparison. Even a freshly promoted credential was deleted on the next Test. 3. The sandbox hello probe always staged the company default home. An agent with a configured `CODEX_HOME` was tested against one credential and ran with another. **Expected behavior** A completed sign-in updates the credential the Test probes. The Test stages the same Codex home a real run resolves. A stale credential never outranks a strictly newer one from an interactive login. **Steps to reproduce** 1. Configure a company whose Codex home holds a shape-valid credential that no longer authenticates (for example an old host login symlink). 2. Open a `codex_local` agent's detail page with a sandbox environment and press Test. The result shows the authentication-needed check. 3. Complete the "Sign in with Codex" device-code flow from the panel. 4. Press Test again. Before this change the result still shows authentication needed. ## What Changed - `packages/adapters/codex-local/src/server/adapter-auth-promotion.ts`: the promotion writes the company default home unconditionally. The shared `last_refresh` merge predicate scopes the write. It seeds an absent or unusable slot, refreshes a same-identity slot only with a strictly newer credential, and keeps a slot a different account or an API-key file holds. The atomic rename replaces a symlinked `auth.json` at the link itself. It never writes through into the host home. - `packages/adapters/codex-local/src/server/codex-home.ts`: the same-identity heal in `seedManagedCodexHome` is freshness-aware. A regular-file credential is swapped for the shared symlink only when the shared source is strictly fresher by `last_refresh`. Ties and unparseable timestamps keep the file, which matches the predicate's fail-closed direction. A genuine stale copy still heals as soon as the host credential rotates past it. - `packages/adapters/codex-local/src/server/test.ts`: the sandbox hello probe prepares and stages the same home a real run resolves. The identity-anchored cache vend runs first. A configured managed `CODEX_HOME` is seeded in place and staged. A genuine external override is staged as-is and never seeded or mutated. - Tests: new pins for the strictly-newer company-home refresh, the different-account keep, the symlink-replaced-without-writing-its-target property, the freshness-aware heal (newer kept, older healed, ties kept), the configured-home staging, and the external-home no-mutation proof. The remote-probe suite now pins `CODEX_HOME`/`PAPERCLIP_HOME` to scratch directories so no test can touch a real `~/.codex`. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src --root packages/adapters/codex-local` — 378 passed, 1 skipped. - Server suites for device login, reconciliation, and the codex adapter (8 files) — 128 passed, 15 skipped. - `tsc --noEmit` clean in the adapter package. ## Risks - Behavioral shift is scoped by the shared merge predicate: only a strictly newer same-identity login can displace a company-home credential, so a second account still never takes over the company slot, and API-key files are never displaced. - The heal keeps ties and unparseable timestamps instead of swapping. A kept file self-corrects on a later seed once the source is provably fresher; deleting a promoted credential is irreversible, so the failure direction is chosen deliberately. - The probe change makes the Test exercise the credential a run uses. A Test that previously passed against the company home while the agent's configured home was broken now fails honestly. ## Model Used Claude (Anthropic) — Claude Fable 5 (`claude-fable-5`), extended thinking, agentic tool use in Claude Code (terminal). Diagnosis traced through the live promotion locks, the on-disk Codex homes, and the adapter's credential-resolution code paths. **Related PRs (searched; no duplicates found):** #12740, #12082, and #9621 touch adjacent Codex credential sync paths; #8495 is the standing hardening effort for probe auth seeding. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no standalone docs cover this flow; the behavioral contracts are documented in-line at each changed site) - [x] I have considered and documented any risks above |
||
|
|
2043e0c735 |
fix: repair runner configuration, macOS execution, and artifact galleries (#13062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters select a provider, a model, and a runtime. > - Runner conversion rejected existing Claude agents. The model list mixed providers. > - The native Claude runner rejected custom models and could not launch on macOS. > - This pull request fixes conversion, model selection, and verified macOS execution. > - It also groups configuration fields consistently across adapters and opens artifact images in the task gallery. > - Operators can change an agent configuration and run the selected model on their Mac. ## Linked Issues or Issue Description **What happened?** Converting an existing Claude agent to Paperclip Runner failed with a Codex-only restriction. ACPX Claude showed unrelated models and required `claude-sonnet-5`. Its native runtime rejected macOS. Configuration mixed common model settings with process controls. Artifact cards labeled “Open gallery” navigated to attachment URLs instead of opening the task gallery. **Expected behavior** Conversion keeps agent identity and compatible settings. ACPX Claude uses the normal Claude catalog and accepts typed model IDs. Codex uses the native runner. The verified Claude runtime can launch on macOS ARM64 and x64. Common configuration sections place the same fields together across adapters. Artifact images open in the shared task gallery with navigation and downloads. **Steps to reproduce** 1. Open the configuration of an existing Claude agent. 2. Convert it to Paperclip Runner. 3. Select ACPX Claude and a different catalog model or a typed model ID. 4. Save the agent and run a disposable task on macOS. 5. Inspect configuration and advanced run-policy controls across adapters. **Paperclip version or commit** The bugs were reproduced on `165ca56a22adb60e5fda56045442d9c8498116a8`. This branch was rebased onto `7ed122911`. **Deployment mode** Built from source. Local test-drive instance on macOS ARM64 with an isolated database. Related work: #11798 addresses unsupported ACP session options in the existing adapter path. #13048 addresses working-folder preservation. This change fixes native runner configuration and launch behavior. ## What Changed - Remove the Codex-only conversion restriction. Preserve agent identity, instructions, directories, credentials, and compatible model settings. Reset incompatible sessions while retaining history. - Show ACPX Claude and native Codex as distinct provider choices. Remove ACPX Codex from advertised configuration. Normalize legacy configurations before fresh runs without rewriting historical run descriptors. - Select model catalogs and cache entries by provider. Support refresh and typed model IDs. Pass exact Claude IDs through session creation, model changes, and recovery. - Add verified macOS ARM64 and x64 Claude SDK snapshots. Bound executable allocation and total snapshot size. Preserve package checks, dependency isolation, process ownership, cancellation, and Linux descriptor loading. - Probe local runtime readiness. Report remote platform checks as incomplete until the remote runner verifies its runtime. - Surface actual model rejection and allow correction and retry. - Repair missing ACPX goal-capability helpers exposed by the post-rebase live test. Persist and restore the optional capability without breaking session startup. - Put Agent identity first and intentionally remove the Capabilities editor, as requested. This is removal of UI editing, not relocation: preserve existing capability metadata and API compatibility without adding another editor. Use the themed select for configurable permission modes, with normal text instead of monospace. - Put model and provider under Adapter. Give environment variables their own section. Fold command and arguments under Configuration. Fold lifecycle, timeout, and interrupt grace under Advanced Run Policy. Hide single-option permission controls. - Open image and video artifact cards in the existing task gallery, including cards in the artifacts panel. Chat attachment images use the same gallery. Preserve standalone media previews and download links. ## Verification - Rebased focused UI/API/database suites: 293 tests passed. - Rebased native runtime and ACPX suites: 242 passed, 7 skipped. - Repository typecheck, build, and token gates passed for the runner changes. Gallery follow-up UI typecheck, build, and token gates also passed. - Follow-up UI suites passed (86 tests), packaging checks passed (14 tests), and the final focused runtime suites passed (126 passed, 7 skipped). - Linux container isolation and lifecycle fixtures passed before rebase (57 passed, 2 skipped). Rust ACPX provider-session tests passed after rebase (8 tests). - Browser tests completed actual Claude and native Codex tasks on macOS ARM64. They covered conversion, catalog refresh, a non-default catalog model, a typed `haiku` ID, save/reload, cancel, follow-up session continuity, invalid-model errors, and recovery. - Final-revision live tests completed a typed Claude task, a follow-up with the same provider session, and a native Codex task on macOS ARM64. - Browser tests confirmed the moved interrupt-grace field saves and survives reload. Cross-adapter tests cover Claude, Codex, Gemini, process, gateway, and schema forms. - Full local run: 7,080 passed, 30 skipped, and two timeouts. Both timeout suites passed on isolated rerun (84 tests); the failures were the plugin login-worker exit diagnostic and the runner real-server vertical slice. - Final follow-up checks: 50 registry tests and 45 snapshot/installation tests passed (6 platform-specific skips). Oversized executable rejection is covered before allocation or reading; unsupported-platform tests invoke the real installation probe. - Runner head `ddb5101c483a297f74875ab96b3c66035b002d50`: all CI gates green, including full runner verification, repository build, typecheck, general/serialized server suites, browser tests, and canary dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34286178670). - Greptile: 5/5 on that runner head. All four review threads resolved. Superagent, Socket, and Snyk checks green. - After snapshot hardening, another real Claude task completed on this Mac using the rebuilt runtime. - Gallery follow-up: 148 focused tests passed, covering artifact selection, shared attachment collections, deduplication, image/video cards, standalone previews, downloads, and closing. Live browser verification completed on the settings follow-up: artifact selection, 6-image pagination with wrapping, download action, and closing all stayed on the same task URL. All checks passed on gallery head `96136da58ff195bf6ca00b281eb3022ad12d7bd8`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34287987536). Greptile returned 5/5 on that exact head with no unresolved threads. - Final settings polish: 96 focused tests, UI typecheck/build, and token gates passed. A real browser walkthrough verified readable permission options, identity placement, Capabilities removal, and permission save/reload. Original test-agent permission mode restored. All 31 checks passed on final head `e46540d6bf32bfb0566dca16b2f4a75ba437618c`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/34292797886). Greptile returned 5/5 with no unresolved threads. ## Risks - Capabilities intentionally has no editable UI field after this change. Existing values remain readable and API-compatible; removing the field does not erase stored metadata. - macOS launch now copies verified package files into private snapshots. The implementation must retain isolation and clean up snapshots on exit. - Runtime provider or model changes reset the current session. Historical runs remain available. - The macOS x64 SDK executable digest was verified, but a live Intel Mac run was not available. Linux verification used container fixtures, not a real Claude task. - Remote environment tests report a warning when only the platform has been checked. They do not claim package readiness from the server host. ## Model Used OpenAI Codex, based on GPT-6. The exact served model identifier and context-window limit are not exposed in this session. Used reasoning, repository inspection, code execution, Rust and TypeScript tests, and browser automation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites and both timeout suites on rerun; full-run counts above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5128b4f323 |
fix(claude): default unset models to Opus 5 (#13055)
Resolve unset Claude models to Opus 5 across CLI and ACP execution, preserve explicit and provider-specific overrides, and show the default in agent configuration. Verified 212 focused tests after merging master, UI typecheck and token gates, and all CI checks. Greptile reviewed the final head at 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ebaeba40ee |
feat: simplify agent onboarding and configuration (#13011)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators create agents and configure their runtimes in the board UI. > - The old creation flow presents several choices and a large form before an agent can start. > - The existing onboarding controls already provide clear provider connection steps. > - This pull request uses those controls in a new-agent wizard and organizes the full configuration pages. > - Operators can connect, test, save, and assign a first task while keeping the existing configuration tools. ## Linked Issues or Issue Description Related: #10974. That earlier open PR also reorganizes agent configuration. This PR follows the reviewed Storybook designs for agent creation and the current configuration tabs. **What existing behavior does this improve?** Agent creation, provider connection, runtime tests, and full agent configuration. **Current behavior** The creation dialog leads to a large manual configuration form. Provider login controls differ from onboarding. Environment variables and secret access appear in separate places. **Proposed behavior** Choose a name and adapter. Connect Claude or Codex through the existing onboarding controls. Configure and test the runtime, save the agent, and open a task dialog with that agent assigned. Use the same design on the existing configuration tabs. **Reason and benefit** The first setup asks for fewer decisions. The full editor keeps instructions, skills, runtime controls, secret access, permissions, keys, and revisions available in clear sections. **Breaking changes** The board creation and configuration layouts change. The test-environment API adds an optional, allowlisted `testCredentials` field for one-shot probes. Database contracts stay the same. Native ACPX tests now reject unsupported local platforms before a CLI login can mask the runtime restriction. ## What Changed - Added a new-agent wizard with numbered steps, adapter branding, provider connections, editable model choices, runtime tests, and confirmation. - Added Codex app-server, Claude ACPX, and OpenCode runner choices. - Stored API credentials through existing secret APIs and persisted references in agent configuration. New setup keys are isolated from credentials used by existing agents. - Preserved external-agent invitations beside the wizard, including optional messages, one-time prompts, and clipboard fallback. - Added OpenRouter provider and secret bindings for Pi and OpenCode. - Added adapter-specific prerequisite fields for Cursor, Gemini, Kimi, and Hermes. Cursor Cloud keys are saved as new organization secrets. - Fixed Cursor Cloud repository field mapping, omitted empty remote environment values, and added useful model and repository error messages. - Preserved complete MCP assignments when multiple valid profiles contain more than 250 tools in total. Generated profiles retain exact tool selectors. - Added service branding and deployment-aware adapter choices. Cloud setup offers Claude, Codex, and OpenCode; local native runners require the experimental setting. - Made the agent list responsive at intermediate widths. - Applied the reviewed design to the real agent configuration pages. Kept the instruction editor, skills, and existing mutations. - Combined secret access and environment variables under one Save and Discard action. - Added interactive Storybook screens for setup, configuration, confirmation, authentication, and test results. - Fixed Pi provider-error parsing and thinking-effort persistence. Native ACPX validates Linux x64 on the actual local, SSH, or sandbox target. - Redacted the complete transient probe-credential field from HTTP error logs, including rejected provider names. ## Verification - Current head `df0292fe6` has a fresh Greptile 5/5 review with no unresolved findings. All 31 executed CI checks passed, including the aggregate verification gate and all browser E2E shards. Storybook visual regression is skipped by its workflow; the local Storybook build passed. - Browser tests completed real assigned tasks with direct Codex, Claude, OpenCode, Pi, and native Codex. - Verified external-agent invitation generation and automatic prompt copying in the live browser. - Pi and OpenCode used an existing OpenRouter secret. Browser checks covered save and reload, instruction edits, skill selection, environment-variable Save and Discard, and assigned task creation. - Invalid Claude API credentials remained on the connection step with an error. A live Pi/OpenRouter invalid-key probe returned a provider failure and left the user-secret inventory unchanged (zero entries before and after). - Full workspace typecheck and build passed after rebasing onto current master. After review fixes, server and UI typechecks, token gates, and the full build passed again. Storybook built successfully. - All 5,542 local UI tests passed. The Cursor Cloud and Pi adapter regressions passed all 24 tests. Review regressions passed 69 server tests and all 18 agent-list tests. - The local full test command ran 6,971 general server tests successfully. Editing review fixes during that long run caused nine tests to use stale modules; fresh isolated runs passed. An unrelated embedded-Postgres fixture hit the host shared-memory limit; its 15 affected tests passed when the fixture groups ran separately. - Local workspace groups passed after rerunning 18 CLI tests sequentially to avoid host database limits and parallel-load timeouts. The local full command stopped at the general server phase, so serialized server verification comes from the five passing CI shards. - Browser testing at 390px confirmed that the agent action menu opens and the page has no horizontal overflow. CI browser E2E shards passed. - Review the `Onboarding / New agent` and `Agents / Configuration refresh` Storybook groups. In the real app, create an agent, run its connection test, save it, assign a task, and reload its configuration. ## Risks - This changes the main agent setup and configuration UI. Regression tests cover routing, persistence, secret bindings, and form actions. - Native Claude ACPX requires Linux x64. Direct Claude works on macOS. Remote checks execute a bounded platform probe and reject unsupported or unverified targets. - A native OpenCode task reached the provider context limit because of its tool payload. Its provider connection test passed. Direct OpenCode completed a task. This existing native execution limit is not fixed here. - Claude and Codex connection keys use the existing user-secret store. Other runtime setup keys use distinct organization secrets. Existing credentials are never rotated. Probes do not store entered keys. Failed agent creation removes newly staged credentials. - Cursor Cloud has not completed a live task. Its authenticated account still needs GitHub repository access. The live run passed MCP provisioning, remote environment validation, and explicit Auto model selection before the repository prerequisite blocked execution. - Generated runtime MCP profiles can exceed the public profile-edit request limit. They still contain exact catalog selectors and preserve permission boundaries. - No database migration, dependency, lockfile, or workflow changes are included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser automation. The runtime did not expose the exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dceee9a4e |
fix(runner): persist warm Daytona workspaces (#12901)
## Thinking Path > - Paperclip manages AI agent work and the execution state for each task. > - Remote agents run in sandbox environments such as Daytona. > - Daytona keeps files while a sandbox is stopped, but deletion removes those files. > - Runner Codex did not copy successful remote workspace changes back to the host workspace. > - A warm sandbox could therefore hide data loss until Daytona replaced or deleted the sandbox. > - This pull request makes the host workspace durable after every successful turn and keeps verified reusable sandboxes warm. > - The benefit is reliable multi-turn work across warm reuse, restart, stop, and sandbox replacement. ## Linked Issues or Issue Description **What happened?** A successful native Codex turn in Daytona could leave workspace changes only in the remote sandbox. A later warm turn appeared to work because it reused that filesystem. A replacement sandbox could start from stale host data and lose the successful changes. **Expected behavior** Paperclip must merge each successful remote turn into the authoritative host workspace before it completes the run. A verified warm lease may reuse its remote files. A replacement lease must reconstruct the exact durable workspace seed. **Steps to reproduce** 1. Run Codex in a reusable Daytona environment. 2. Write a file during one successful turn. 3. Replace the Daytona sandbox before the next turn. 4. Observe that the next turn can start without the prior file on the unpatched code. Related remote workspace foundation: #10070. ## What Changed - Added explicit `host_current`, `durable_seed`, and `adopt_remote` workspace preparation modes. - Added atomic, versioned native workspace descriptors and seed archives under `PAPERCLIP_HOME`. - Added real native sandbox export and three-way host merge before terminal result completion. - Added workspace-only recovery after a proposed result. Recovery does not submit another provider turn or consume the provider retry budget. - Added fail-closed handling when a sandbox with unexported changes is gone. - Kept healthy reusable Daytona sandboxes started for legacy Codex and Runner Codex. - Kept the Runner Codex process and provider session across verified warm turns. - Added the paid `daytona-warm-continuity` browser suite. It contains exactly the legacy Codex and Runner Codex cells. Each cell performs three measured turns. - Documented `pnpm test:e2e:runner -- --suite daytona-warm-continuity`. No package script was added. - Added no database migration. The metadata format is backward compatible and idempotent. ## Verification - `pnpm typecheck` - `pnpm test:e2e:runner:unit` — 114 passed - Native workspace, finalizer, session, and environment tests — 232 passed - Daytona provider tests — 150 passed - Workspace staging and merge tests — 98 passed - Runner transport tests — 63 passed - Legacy Codex restore tests — 5 passed - Rust format and compile checks pass through root typecheck - The paid Daytona suite was not run locally because the required Daytona, OpenAI, and immutable image credentials are not present. ## Risks - The main risk is an incorrect workspace identity or merge after a crash. Durable descriptors bind the run, workspace, lease, provider lease, local root, remote root, and baseline digest. Ambiguous evidence fails closed. - The host merge may conflict with concurrent host edits. The existing three-way merge and exclusion rules handle this case and surface failures. - A deleted sandbox cannot recover unexported bytes. Paperclip now blocks with `workspace_sync_out_unrecoverable` instead of reporting success or rerunning the provider. - There is no database migration. Descriptor writes and recovery are atomic and idempotent. ## Model Used OpenAI Codex with GPT-5. The run used agentic reasoning, repository inspection, code execution, test execution, Git, and GitHub CLI tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
8f0c1d4548 |
feat(cli): add isolated test-drive command (#12894)
Add a foreground-only test-drive workflow with isolated data, provider-backed CEO bootstrap, OpenCode/OpenRouter support, worktree execution setup, reuse safeguards, and delayed browser opening. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
77312ee2d9 |
feat(codex): add GPT-6 Astra support (#12851)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The Codex local adapter supplies model metadata to the server and the user interface. > - OpenAI now lists `gpt-6-astra` as a supported Codex model. > - Paperclip did not list this model or its model-specific controls. > - This pull request adds the model through the existing adapter metadata path. > - The benefit is that agents and task overrides can use the exact model ID and supported controls. ## Linked Issues or Issue Description **Subsystem affected** `packages/adapters` and `ui` **Problem or motivation** Paperclip does not expose `gpt-6-astra` in Codex model selectors. Operators cannot select and save the model through the normal agent and task forms. **Proposed solution** Register the exact model ID in the Codex local adapter. Use the adapter as the source for the model-specific reasoning options. Preserve the current default model. Forward the saved model, reasoning effort, and fast-mode controls through both Codex execution lanes. **Alternatives considered** A user-interface-only model list would duplicate adapter metadata. A model alias would not match the official model ID. Both options were rejected. **Roadmap alignment** This is a small adapter compatibility update. It does not duplicate a planned item in `ROADMAP.md`. ## What Changed - Added `gpt-6-astra` to the Codex local adapter model registry and fast-mode support list. - Added the official Astra reasoning efforts: `low`, `medium`, `high`, `xhigh`, `max`, and `ultra`. - Used the adapter metadata in agent and task model selectors. - Preserved supported effort choices when the model changes. Cleared an effort only when the new model does not support it. - Added tests for registration, user-interface selection, configuration persistence, and CLI and ACP forwarding. ## Verification - `pnpm exec vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/acp.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/lib/codex-reasoning-effort.test.ts ui/src/components/AgentConfigForm.render.test.tsx ui/src/components/IssueProperties.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/issue-assignee-overrides.test.ts` passed 245 tests. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed all four gates across 939 files. - `pnpm --filter @paperclipai/ui build` passed and supplied isolated user-interface build proof. - `pnpm build` passed. - `pnpm test:run` passed 5,812 tests and failed 24 workspace-runtime tests in this isolated host. The failures use invalid generated ports above 65,535, incomplete nested-worktree fixture configuration, or `/tmp` path aliases. The focused tests for this change all pass. GitHub CI must pass before review handoff. - GitHub CI run `33918372718` passed all required checks and the aggregate verify gate on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - Independent engineering review approved the exact remediation head after 170/170 reviewer tests passed. - Greptile reported 5/5 with no open review threads on exact head `6ac6be2cee0a5996c82bdf674fcb7f46cb4c5fde`. - The model ID and capabilities were checked against the [official OpenAI Codex model list](https://developers.openai.com/codex/models). ## Risks - Low risk. The change adds one model and model-specific selector options. It does not change the default model. - OpenAI can change model capabilities later. The adapter metadata must stay aligned with the official Codex metadata. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with model ID `gpt-5.6-sol`, a 272,000-token context window, reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af3023f1e3 |
fix(runner): repair paid provider startup paths (#12769)
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
a0028d7e1b |
chore(deps-dev): bump vitest from 4.1.10 to 4.1.11 (#12262)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.10 to 4.1.11. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v4.1.11</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Revive global concurrency limit for test lifecycle [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> and <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10992">vitest-dev/vitest#10992</a> <a href="https://github.com/vitest-dev/vitest/commit/5146df80b"><!-- raw HTML omitted -->(5146d)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Encode iframeId in tester iframe URL [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a>, <strong>Pduhard</strong> and <strong>Claude Opus 4.8</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10955">vitest-dev/vitest#10955</a> <a href="https://github.com/vitest-dev/vitest/commit/10b2cd201"><!-- raw HTML omitted -->(10b2c)<!-- raw HTML omitted --></a></li> <li>Trigger playwright/chromium gc on lower disk availability [backport to v4] - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>OpenCode</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10951">vitest-dev/vitest#10951</a> <a href="https://github.com/vitest-dev/vitest/commit/9851dbc41"><!-- raw HTML omitted -->(9851d)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>mocker</strong>: <ul> <li>Restrict redirect mocks to the fs allowlist [backport to v4] - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/10974">vitest-dev/vitest#10974</a> <a href="https://github.com/vitest-dev/vitest/commit/fe5a11d3c"><!-- raw HTML omitted -->(fe5a1)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v4.1.10...v4.1.11">View changes on GitHub</a></h5> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/9bd8d464e6328c567c2dbcd8fdd977d57a9425c2"><code>9bd8d46</code></a> chore: release v4.1.11 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10995">#10995</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/9851dbc41c286a30abfb6b29cce65f3e5b7b40a1"><code>9851dbc</code></a> fix(browser): trigger playwright/chromium gc on lower disk availability [back...</li> <li>See full diff in <a href="https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
9064cfd09e |
feat(codex-local): give each Codex account its own home and path secret (#12709)
## Thinking Path > - Paperclip is the control plane for companies that use AI agents for work > - Local adapters connect Paperclip agents to provider command line tools > - The Codex adapter stores login data in a shared company home > - A shared home cannot keep credentials for more than one Codex account > - This pull request gives each account a safe home and a matching company secret > - The benefit is that one company can use multiple Codex accounts at the same time ## Linked Issues or Issue Description **Problem or motivation** A company can hold only one Codex subscription credential because device login uses one shared home. A second account cannot log in without replacing or conflicting with the first credential. **Proposed solution** This change validates the vendor account identifier, stores each credential in its own home, and creates a company secret that points to that home. Repeat login calls return success when the matching secret already exists. **Roadmap alignment** The change supports the roadmap goal for centrally managed secrets with scoped access and audited resolution. **Additional context** The security review returned approve with no blocking finding. The branch adds shared account-handle validation and tests for device login and the Codex local adapter. ## What Changed - Add strict allowlist validation for Codex account handles. - Store each Codex account credential in a separate home under the Codex cache root. - Verify that the resolved account home stays inside the cache root. - Create the `CODEX_HOME_<handle>` company secret for each account. - Keep repeat and concurrent login calls safe and idempotent. - Add shared helper and route, adapter, and validation tests. ## Verification - `pnpm --filter @paperclipai/adapter-codex-local test` passes with 343 tests. - `pnpm --filter @paperclipai/server test src/__tests__/agent-device-login-routes.test.ts` passes with 25 tests. - The adapter suite passes with 23 tests. - The shared package and Codex adapter typechecks pass. - Continuous integration must pass on every check before merge. ## Risks The account handle becomes part of a directory path and secret name. The strict allowlist and root containment check reduce path traversal risk. Existing single-account homes remain unchanged unless a new device login creates an account-specific home. ## Model Used OpenAI GPT-5 (exact runtime model ID: gpt-5), with tool use and code execution. The runtime context window is not exposed in this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dfdfc8664e |
feat(claude-local): add Claude Fable 5.1 support (#12730)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude local adapter lets operators select a Claude model for an agent. > - Claude Fable 5.1 was absent from the adapter model lists. > - The adapter runtime also used a Claude Code build that rejected Fable 5.1. > - This pull request adds the direct Anthropic ID and the AWS Bedrock inference profile ID. > - It also updates the Claude ACP runtime and keeps the Paperclip usage and isolation patches. > - The benefit is that operators can select and run Claude Fable 5.1 through the Claude adapter. ## Linked Issues or Issue Description Refs #8810. That issue covers related model ID handling. This change does not change provider-prefixed model IDs. **Agent or provider** Claude Code through the built-in `claude_local` adapter. The requested model is Claude Fable 5.1. **Why this adapter is useful** Operators can use Fable 5.1 without entering an undocumented model ID. The configured model also reaches both supported Claude execution lanes. **How the agent is invoked** The CLI lane sends `--model claude-fable-5-1`. The ACP lane sends `ANTHROPIC_MODEL=claude-fable-5-1` to `@agentclientprotocol/claude-agent-acp`. **Are you willing to implement it?** Yes. This pull request includes the implementation and tests. **Additional context** Claude Code 2.1.232 rejected Fable 5.1 and required version 2.1.251 or newer. ACP package 0.73.0 includes Claude Code 2.1.257. The update keeps Paperclip's usage metadata and isolated-context behavior. ## What Changed - Added `claude-fable-5-1` to the direct Claude fallback list. - Added `us.anthropic.claude-fable-5-1` to the AWS Bedrock list. - Kept the existing default model at the first position in each list. - Updated the Claude ACP dependency from 0.70 to 0.73. - Carried the Paperclip usage and isolated-context changes into the 0.73 patch. - Added a Claude Code 2.1.251 minimum-version preflight for Fable 5.1 when using the standard `claude` executable, surfaced in both adapter Test and execution. Explicit custom wrappers retain their existing compatibility contract. - Kept local adapter Tests from executing caller-selected binaries: when runtime `PATH` selects a different Claude executable than the trusted probe, the Test warns and defers the authoritative version check to execution instead of approving or rejecting the alternate installation. - Added tests for model listing, discovery deduplication, Bedrock filtering, model pass-through in both execution lanes, old-CLI rejection before launch, custom-wrapper compatibility, and local runtime-PATH mismatch handling. ## Verification - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm exec vitest run packages/adapters/claude-local/src/server/execute.remote.test.ts packages/adapters/claude-local/src/server/test.remote.test.ts packages/adapters/claude-local/src/server/test.probe.test.ts packages/adapters/claude-local/src/server/acp.test.ts server/src/__tests__/adapter-models.test.ts` (72 tests passed) - `node --test scripts/acpx-patch-packaging.test.mjs` (13 tests passed) - `pnpm -r typecheck` - `pnpm build` - A local Paperclip agent run completed with `usageJson.model` set to `claude-fable-5-1` through ACP 0.73.0 and its bundled Claude Code 2.1.257. - `pnpm test:run` completed 5,638 passing tests and 24 skipped tests. It also found 24 failures in unrelated workspace-runtime, path-canonicalization, and runtime-exposure tests on macOS with Node 26. These failures do not touch this diff. Clean pull request CI is the final full-suite gate. ## Risks - The ACP dependency update can change Claude runtime behavior outside model selection. Focused ACP tests, the full typecheck, the production build, and a real local Fable run reduce this risk. - The 0.73 patch must stay aligned with the installed ACP version. Dependency-resolution CI verifies the manifest and patch pair. - Fable 5.1 adds a short `claude --version` preflight to standard CLI-lane Tests and runs. The result is intentionally not cached so an in-place Claude Code upgrade takes effect without restarting Paperclip. Explicit custom wrappers are not version-probed because their output and compatibility contract can differ from the standard executable. - Local Tests preserve the existing deny-by-default probe boundary and do not execute a binary selected by caller-controlled `PATH`. A mismatched runtime binary produces an explicit warning without blocking an otherwise valid setup; execution independently validates the actual runtime-selected CLI before launch. - The AWS Bedrock identifier differs from earlier IDs because Fable 5.1 has no `-v1` suffix. The model-list test locks this exact value. - There is no schema change or migration. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Provider: OpenAI. Model: GPT-5 Codex. The host did not expose a more specific model ID or context-window size. Capabilities used: agentic reasoning, repository editing, shell execution, web research, and local runtime verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f94521017 |
fix(runner): restore local session and task integrity (#12721)
## Thinking Path > - Paperclip is the control plane for agents that perform work. > - Paperclip Runner connects durable provider sessions to individual task runs through PRP. > - Provider continuity and per-run authority are different lifetimes. > - The existing implementation mixed those lifetimes and lost event metadata between provider frames, runnerd, persistence, API sanitization, and the task thread. > - That caused failed continuation, missing progress and Plans, duplicate replies, hidden failures, and unsafe recovery. > - This repair gives every heartbeat fresh authority, preserves qualified provider-session continuity, and restores one lossless presentation path without changing direct adapters. ## Linked Issues or Issue Description **What happened?** A second native heartbeat could reuse tickets, leases, command receipts, sequence state, and run identity from the first heartbeat. Provider phase and item identity could be lost before the UI read them. Redaction could corrupt protocol discriminators while still missing malformed credential tails. The task thread could fold progress into the final response, hide failures, or show more than one final answer. Native Codex also exposed approval modes that do not yet have a durable approval bridge. **Expected behavior** Each heartbeat uses a new PRP authority epoch. Codex and OpenCode preserve exact qualified provider sessions; ACPX emits an explicit continuity event when its qualified process-replacement policy is used. Every accepted provider event is presented, classified as internal, or surfaced as unsupported. The task page shows chronological progress, reasoning summaries, activity, Plans, interactions, terminal failures, and exactly one final reply. Direct adapters retain their existing path. **Steps to reproduce** 1. Enable the unified experimental Paperclip Runner setting. 2. Create a local native Codex, OpenCode, ACPX Claude, or ACPX Codex agent. 3. Run response, Plan, structured-question/resume, restart, cancellation, and failure scenarios. 4. Reload the task while active, waiting, failed, and settled. 5. On the old implementation, observe stale run authority, missing classifications, incomplete output, or duplicated/folded replies. **Paperclip version or commit** The repair is based directly on `master` at `87d05e194b643810d16d20612115acd01d735d43`. **Deployment mode** Local development with the embedded database. Related work: Refs #12616, #12646, #12666, #12685, and #12700. ## What Changed - Rotates PRP control-plane, outbox, ticket, lease, command, receipt, and sequence authority for each heartbeat while carrying forward only a validated provider-session identity. - Reads `control-plane-state.json`, validates both durable schemas and lifecycle values, resumes coherent current runs, archives qualified settled authority, and quarantines malformed or mismatched scoped state without moving ambiguous live legacy state. - Preserves Codex provider phase and stable item identities so commentary remains progress and only `final_answer` becomes final. - Adds raw OpenCode HTTP/SSE boundary coverage and canonical reasoning lifecycle mapping. - Makes ACPX normalization lossless for visible reasoning, tool lifecycle metadata, stable bounded identities, Plan revisions, structured requests, failures, and qualified process replacement. Only the compatible terminal assistant message is promoted as final. - Applies schema-aware redaction before generic JWT-shaped detection and scans every diagnostic string leaf. Malformed raw/escaped quoted credential tails are redacted in both server and durable Rust state. - Restores snapshot-style chronological task presentation, expandable tool activity, inline Plan cards, visible waiting/resume/cancel/failure states, and exactly one final answer. - Makes `never` the only qualified native Codex permission mode and rejects unsupported persisted native modes with remediation. OpenCode and ACPX policies remain intact. - Keeps the unified experimental Runner setting as the only enablement flag. Onboarding and direct Codex, Claude, and OpenCode stay on their legacy execution/finalization paths. - Adds cross-language goldens, authority/recovery/fault coverage, exact response/count assertions, and native plus legacy acceptance scenarios. ## Verification - Pull-request GitHub Actions run Rust formatting/tests, TypeScript checks, server/UI tests, builds, protocol drift checks, browser E2E, and security scans. - A separate workflow-only validation ref is pinned directly on this PR head and runs the 35-cell paid local matrix: three core scenarios plus structured-question resume and restart/resume for native Codex, native OpenCode, ACPX Claude, ACPX Codex, and direct Codex/Claude/OpenCode. Run: https://github.com/paperclipai/paperclip/actions/runs/33682434315 - Acceptance requires exact single visible replies, monotonic sequences, matching envelope discriminators, one semantic terminal, one run terminal, no unresolved interaction, no duplicate mutation, no secret leakage, provider continuity, and zero native rows for direct adapters. - Per maintainer direction, tests are running in GitHub Actions rather than on the slower local host. Only formatters and static diff checks were run locally. ## Risks - Recovery from old or partial filesystem state is sensitive. The repair fails closed, preserves active or unverifiable authority, and quarantines only state whose scoped ownership is safe to move. - Provider event formats can change. Closed validators and boundary goldens turn new or malformed events into visible diagnostics instead of silent drops. - Shared task presentation could affect direct adapters. Runtime-fact gating plus the direct-adapter matrix protect the existing path. - Managed and remote providers are not qualified here. Shared code continues to compile and fail safely, but live qualification is deferred. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact deployed snapshot and context-window size are not exposed to this task. It used agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] The paid local-provider matrix is green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fdf8c8464d |
feat(runner): add managed provider backends (#12699)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner provides durable, provider-neutral agent execution. > - The current stack supports qualified local providers but omits the managed provider paths from the integration branch. > - Claude Managed Agents and AWS AgentCore need explicit profile qualification, durable recovery, usage accounting, and cleanup controls. > - This pull request adds those managed backends as the third part of the Runner parity stack. > - The benefit is managed execution without weakening the default-off Runner rollout gate. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: Runner, server orchestration, database profiles, CLI, and adapter configuration UI. **Problem or motivation** The current Runner stack cannot select or execute the managed Claude Agents API or AWS Bedrock AgentCore Harness backends. It also lacks qualified profile storage and recovery checks for those remote resources. **Proposed solution** Add qualified managed and remote profiles, API and CLI management, exact provider selection, durable lifecycle handling, cumulative usage accounting, bounded cleanup, and retention acknowledgement. Keep `enableNativeRunner` default-off. **Alternatives considered** A direct copy of the old integration branch was rejected because its provider contracts, model values, credential flow, and migration history no longer match the current base. A single large parity pull request was also rejected because stacked review keeps each subsystem bounded. **Roadmap alignment** This continues the existing Runner architecture and rollout work. It does not introduce a separate execution system. **Additional context** This pull request is based on the merged #12691 and #12685 stack. It also closes the delayed security-review findings reported on #12691 by binding qualified ACPX and OpenCode launch artifacts to the bytes actually executed. A GitHub search for managed agent, AgentCore, and Claude managed work found no duplicate public issue or pull request. ## What Changed - Add Claude Managed Agents and AWS AgentCore provider executors to runnerd. - Add qualified managed and remote profile storage, routes, OpenAPI contracts, CLI commands, and migration 0237. - Validate profile ownership, enabled state, exact qualified revision, model, agent version, and secret binding before persistence and recovery. - Persist durable provider session and owned skill state for restart-safe cleanup. - Reconcile uncertain create responses and delete remote sessions before owned skills. - Track cumulative provider usage and enforce positive session spend caps. - Recover interrupted AgentCore usage at the next turn boundary by charging the prior invocation ceiling exactly once; keep the session gated until an explicit monotonic budget raise. - Isolate AgentCore AWS configuration from host profiles and credential-process/SSO configuration while preserving workload identity. - Require OpenCode 1.18.17 and fixed build-owned provider-pack artifact paths; remove the ambient executable override. - Snapshot and content-verify ACPX and OpenCode commands, scripts, and provider executables before launch. Linux executes sealed inherited descriptors; macOS uses authenticated private snapshots with retry-safe rematerialization at the spawn boundary. - Persist canonical ACPX and OpenCode launch-profile digests, reject drift across fresh recovery, and make recovery failures sticky. - Close and journal unsafe ACPX active-turn recovery before any provider bootstrap or reconnect. - Add managed provider fields to the Runner configuration UI and permission projection. - Preserve the default-off `enableNativeRunner` experimental flag. ## Verification - `pnpm -r typecheck` - `pnpm build` - Focused managed server, database, CLI, Runner TypeScript, Rust, Claude, AgentCore, ACPX, OpenCode, process-supervisor, and durable-recovery tests passed. - `cargo test -p paperclip-runner-core --lib --locked` (160 tests) - `cargo check --workspace --all-targets --locked` - Native Codex integration tests passed (60 tests); native provider tests passed (7 tests); server native-runtime tests passed (87 tests). - Verified-launch replacement, nested-spawn retry, exact-version, profile-drift, sticky-failure, and no-bootstrap active-recovery tests passed. - `git diff --check` - The PR changes 91 files. `pnpm-lock.yaml` is unchanged. The Rust workspace lockfile adds the approved `rustix` dependency used for safe descriptor handling while `#![forbid(unsafe_code)]` remains enabled. ## Risks - The provider APIs can change while they are in beta. Exact qualification and fail-closed recovery checks limit drift. - Remote cleanup can fail after a partial create. Durable ownership inventories and retry-safe deletion preserve recovery state. - Migration 0237 adds profile tables. The generated migration and snapshot pass the repository migration checks. - Managed execution can incur provider cost. Positive default spend caps and explicit retention acknowledgement limit accidental use. - An interrupted AgentCore invocation without final metadata is conservatively charged to its active session ceiling. This can overstate cost, but cannot undercount it; later work requires an explicit budget increase. - Linux qualified launches use sealed memory descriptors. macOS lacks executable-descriptor APIs, so the runner uses owner-only private snapshots and minimizes linked-path lifetime; hostile same-UID processes remain outside the documented local-host trust boundary. - The global Runner feature remains default-off. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
84bedd4ca1 |
feat(runner): activate qualified OpenCode and ACPX providers (#12691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner is the experimental native runtime for governed agent work. > - The runtime contracts already describe Codex, OpenCode, and ACPX providers. > - The merged control plane still rejected OpenCode and ACPX for new runner agents. > - Runnerd also selected only the Codex provider implementation. > - This pull request activates the qualified OpenCode and ACPX paths from the form to runnerd. > - The benefit is one durable runner path with provider-specific permissions and recovery. ## Linked Issues or Issue Description Refs #12685 **Subsystem affected** This change affects the runner package, server orchestration, adapter configuration, and UI configuration. **Problem or motivation** Paperclip Runner stores provider contracts for OpenCode and ACPX. New agents cannot select those providers. Runnerd cannot execute those stored provider descriptors. The UI also shows only Codex. **Proposed solution** Accept the qualified OpenCode 1.18.17 profile and the fixed ACPX Claude and Codex profiles. Route them through runnerd. Keep provider selection, model selection, permissions, credentials, events, and recovery inside closed provider-specific boundaries. **Alternatives considered** One option was to keep the contracts dormant. That option leaves stored configuration and runtime behavior out of sync. Another option was to enable every ACPX agent. That option is not safe because Pi does not yet have the same verified launch path. **Roadmap alignment** This change supports the completed cloud and sandbox agent milestone. It also supports self-healing runs and governed agent execution. It does not add a new roadmap surface. ## What Changed - Add one server profile resolver for Codex, OpenCode, and qualified ACPX descriptors. - Keep `adapterConfig` as the provider and permission authority for fresh runs. - Add Paperclip Runner provider, ACPX agent, and provider-specific permission controls to the UI. - Reset the model to a compatible qualified value when the provider changes. - Route Codex, OpenCode, and ACPX through the durable runnerd provider selector. - Add a durable ACPX executor with bounded state, recovery, events, tool receipts, and identity checks. - Remove Codex labels from OpenCode events, results, evidence, and recovery diagnostics. - Pass only provider-specific credential names to child processes. - Keep ACPX Pi unavailable and reject it before process launch. - Keep the existing Paperclip Runner experimental flag unchanged. ## Verification - `pnpm exec vitest run packages/paperclip-runner/src/backends/native-backend-factory.test.ts packages/paperclip-runner/src/live/runnerd-codex-transport.test.ts packages/adapters/codex-local/src/ui/build-config.test.ts ui/src/adapters/codex-local/config-fields.test.tsx server/src/__tests__/adapter-registry.test.ts server/src/__tests__/adapter-routes.test.ts server/src/__tests__/agent-adapter-validation-routes.test.ts server/src/__tests__/company-portability.test.ts server/src/services/native-runtime/runtime-mode.test.ts server/src/services/native-runtime/native-session-executor.test.ts server/src/services/heartbeat-runner-provider-config.test.ts` - The focused TypeScript, server, and UI suites passed 274 tests. - `cargo test -p paperclip-runner-core --test native_provider_backend` - The executable native provider integration suite passed 4 tests. - `cargo test -p paperclip-runner-core --lib` - The Rust unit suite passed 91 tests. - `pnpm -r typecheck` - `pnpm check:token-gates` - `pnpm build` - `git diff --check codex/runner-parity-task-runtime...HEAD` ## Risks - This changes provider process selection and durable recovery. The experimental flag still gates every fresh Paperclip Runner run. - OpenCode requires a model in `provider/model` form and stays pinned to version 1.18.17. - ACPX accepts only exact Claude and Codex profile versions and models. Pi stays unavailable. - ACPX steering stays unavailable and reports that limit through the driver capabilities. - Child processes receive explicit environment allowlists. They do not inherit the full server environment. - This pull request has no database migration. ## Model Used OpenAI Codex, GPT-5, with tool use, code execution, and subagent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
b4f302d040 |
feat(grok-local): copy a refreshed sandbox credential back to the host (#12696)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through provider-specific adapters in local and remote environments > - A remote Grok run can refresh its credential inside its sandbox > - The host copy can become stale when teardown discards that refreshed credential > - This pull request copies the refreshed credential back through a locked, fail-closed teardown path > - The benefit is that later Grok runs can use the refreshed host credential without another login ## Linked Issues or Issue Description Refs: #12618 **Agent or provider** Grok local adapter. **Why this adapter is useful** A remote Grok run can refresh its access token during a run. Copying the refreshed credential back to the host keeps later runs ready to use. **How the agent is invoked** Paperclip invokes the Grok local adapter through its remote subscription run path. The adapter stages the company Grok home as a sandbox asset. The change adds a copy-out step on the teardown path. ## What Changed - `grok-auth-merge-decision.cjs` adds a host predicate in its own process. It compares the whole `<issuer>::<uuid>` identity key of the two files. It reads `expires_at` as an ISO-8601 string, an epoch-seconds number, or an epoch-milliseconds number. It exits 10 to use the source, 20 to keep the destination, 21 when the expiry shape is unreadable, and 22 when the source expiry sits more than 400 days after the host clock. It fails closed in every unclear case: an unusable side, a different identity, an absent expiry, a tie, an unreadable expiry, and an implausible expiry all keep the destination. - `grok-auth-merge-decision.ts` adds a wrapper that runs the predicate and maps the exit code to a typed result. - `grok-auth-copyback.ts` adds `copyBackGrokAuth({ hostHomeDir, readSandboxAuth, log, env })`. It locks on `hostHomeDir` with `withDirectoryMergeLock`, stages the sandbox bytes into a private `0600` temporary file, runs the predicate, and installs the file with an atomic rename in the same directory. It keeps no backup of the displaced credential. It leaves no temporary file on the success path, the keep path, or an error path. On an error it logs the `errno` code only, then re-throws. - `execute.ts` adds a `restore` callback to the Grok `home` asset. The callback takes the destination from `resolveManagedGrokHomeDir(process.env, agent.companyId)`, never from `env.GROK_HOME`. A copy-out failure does not fail the run. - `package.json` updates the `build` script to copy `grok-auth-merge-decision.cjs` into `dist/server/`, because `tsc` does not copy a `.cjs` file. **The credential shape this predicate reads** A redacted sample of a real vendor credential answered four structural questions. The answers hold no credential bytes, no account identifier, no file path, and no timestamp value. 1. `expires_at` is present. 2. `expires_at` sits inside the value object, under the `<issuer>::<uuid>` key. It is not a top-level field. 3. `expires_at` is an ISO-8601 string. It carries UTC time with a trailing `Z` and six fractional-second digits. 4. A normal run rewrites `auth.json`. The value object carries a `refresh_token` next to `expires_at`, so the client refreshes the access token and rewrites the file. ## Verification - [x] `pnpm vitest run packages/adapters/grok-local` — 114 tests in 12 files pass. - [x] `pnpm --filter @paperclipai/adapter-grok-local typecheck` — clean. - [x] `pnpm --filter @paperclipai/adapter-grok-local build` — succeeds, and `dist/server/grok-auth-merge-decision.cjs` exists after the build. - [x] Continuous integration is green on every check. ## Risks The predicate keeps the host credential when identity, expiry, file access, or freshness data is unclear. The copy-out path can log an error and leave the run successful when it cannot install the refreshed credential. The atomic rename and directory lock protect the host file from partial writes and concurrent copy-out actions. ## Model Used OpenAI GPT-5, current deployment. The exact runtime version and context window are not exposed to this agent. The model used tool calls and code inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b6de5327e |
Remove cheap model profiles (#12683)
## Thinking Path > - Paperclip manages agents that use different model providers and adapters. > - Paperclip must keep agent execution rules clear and predictable. > - The cheap-model profile added a second execution mode across adapters, task recovery, APIs, and the UI. > - That mode increased configuration and recovery complexity. > - This pull request removes the cheap-model profile as a product feature. > - The benefit is one model-selection path for normal work and recovery work. ## Linked Issues or Issue Description **What existing behavior does this improve?** This change simplifies model selection across agent configuration, task execution, recovery, and adapter capabilities. **Current behavior** Paperclip exposes cheap-model profiles in adapter metadata, agent runtime configuration, task overrides, recovery rules, APIs, and the board UI. Recovery work can select a different model profile from the agent's configured model. **Proposed behavior** Paperclip uses the agent's configured model for normal work and recovery work. Status-only recovery stays limited to coordination work. The API rejects legacy model-profile configuration. A migration removes stored model-profile values from existing agent, issue, and historical revision records. **Reason and benefit** One model path reduces configuration, API, UI, and recovery complexity. It also prevents status recovery from becoming a separate product-level model-routing feature. **Breaking changes** This change removes model-profile fields and adapter capability metadata. Existing stored model-profile values are removed by an idempotent migration. The validators reject new legacy profile values with clear errors. ## What Changed - Removed model-profile types, adapter capabilities, API fields, and model selection logic. - Removed cheap-model controls from agent and task UI surfaces. - Kept status-only recovery limited to coordination context while normal continuations use the configured agent model. - Added an idempotent migration that removes stored model-profile values from agents, issues, and configuration revisions without changing issue update timestamps. - Updated tests and product documentation for the single-model behavior. ## Verification - `pnpm check:token-gates` passes. - `pnpm -r typecheck` passes. - `pnpm build` passes. - `pnpm test:run` completed with 5,607 passing tests and 8 environment-sensitive failures in unrelated fixed-port and database-deadlock suites. The same failures repeated in an isolated rerun. CI is the final clean-room result. ## Risks - This is an intentional breaking change for clients that send model-profile fields. - The migration changes legacy agent, issue, and configuration-revision JSON. It is idempotent and preserves unrelated fields and issue update timestamps. - The change is cross-cutting because the removed feature existed in adapters, shared contracts, the server, plugins, and the UI. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. Reasoning and tool use were enabled. The runtime did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ccf3355b2e |
feat(grok-local): stage a curated Grok home into remote subscription runs (#12618)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters run agents in local or remote sandboxes. > - A remote Grok subscription run needs its credential file inside the sandbox. > - The adapter did not stage the Grok home, so the sandbox had no credential file. > - This pull request stages only the allowed Grok credential file and sets the reported home path. > - The adapter removes the temporary staged home before teardown completes. > - The benefit is reliable Grok subscription authentication with limited credential exposure. ## Linked Issues or Issue Description **What happened?** A remote Grok subscription run had no credential file in its sandbox. The adapter sent no Grok home asset. **Expected behavior** The adapter should stage the allowed Grok credential file and set `GROK_HOME` to the reported sandbox path. **Steps to reproduce** 1. Start a remote Grok run in subscription mode. 2. Inspect the sandbox environment and home asset. 3. Confirm that the run has `GROK_HOME` and `auth.json`. **Paperclip version or commit** `3df33b5b8f49063a5d1ab608f8ce372572ef09d1` **Deployment mode** Remote sandbox run. ## What Changed - Stage a private temporary Grok home for remote subscription runs. - Copy only the allowed `auth.json` file and set its mode to `0600`. - Pass the staged directory as the remote `home` asset. - Set `GROK_HOME` to the path that the remote runtime reports. - Remove the staged directory before awaited teardown calls. - Keep the API-key lane free of credential staging. - Add tests for the allowlist, file mode, empty source home, run lanes, and teardown cleanup. ## Verification - `pnpm vitest run packages/adapters/grok-local` passes. - `pnpm --filter @paperclipai/adapter-grok-local typecheck` passes. - `grok-home.test.ts` covers the allowlist, mode `0600`, and empty source home. - `execute.test.ts` covers the remote subscription lane, the API-key lane, and cleanup after restore failure. ## Risks - The change affects only remote Grok subscription runs that use a credential file. - The allowlist limits the staged content to `auth.json`. - The API-key lane does not stage a home or set `GROK_HOME`. - CI must confirm adapter behavior across the supported runtime matrix. ## Model Used - Codex, GPT-5, current 2026 model version, large context window, reasoning mode, and tool use assisted the repository handoff and pull request management. The implementation author supplied the code and local verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
131f5c4065 |
feat(runner): add administration and observability (#12641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Administrators need bounded controls for experimental native execution. > - The lower stack adds remote Codex execution and the task workspace. > - Operators need to configure Codex safely and inspect provider traces. > - Unsupported providers must not appear as runnable choices. > - This pull request adds Codex-only administration and observability. > - The benefit is a default-off operational surface for production diagnosis. ## Linked Issues or Issue Description Refs #12640. Refs #12616. Refs #12352. **Subsystem affected** Agent configuration, instance experimental settings, run ledger, provider trace inspector, and administrator actions. **Problem or motivation** The native runner lacks one safe operator surface for Codex permissions, lifecycle, raw trace capture, and run inspection. The integration branch also contains provider choices that the production backend cannot execute yet. **Proposed solution** Expose only the qualified Codex controls. Keep Paperclip Developer Mode and runner preview ingress off by default. Gate raw trace actions by administrator access and existing trace authorization. **Alternatives considered** Exposing unfinished providers would create configurations that fail at runtime. Always-on tracing would increase sensitive data and storage risk. **Roadmap alignment** This work supports governed Cloud and Sandbox agents and production diagnostics. ## Stack - Base PR: #12640. - Lower PRs: #12639 and #12638. - This PR contains only its 54-file administration and observability delta. - This is the final feature PR in the Codex production stack. ## What Changed - Added Codex-only Paperclip Runner permission and lifecycle controls. - Added bounded warm idle configuration. - Kept the provider field fixed to Codex. - Added administrator-only one-run raw trace requests. - Added a persistent future-run raw trace toggle. - Added trace status, metadata, ledger, and canonical runner inspection. - Added JSON-RPC request-origin grouping and finalization lineage. - Restored the stateful PRP transcript parser and focused projection tests required by trace inspection. - Added default-off Paperclip Developer Mode. - Added Honeycomb run links for authorized developer mode. - Disabled the legacy operational skill for `paperclip_runner`. - Did not expose OpenCode, ACPX, Pi, Claude Managed, or AWS runner choices. - Did not change migrations, workflows, dependencies, or `pnpm-lock.yaml`. ## Verification - GitHub Actions will run UI tests, server tests, repository typecheck, build, browser tests, security, and policy gates. - Tests cover Codex configuration defaults and bounds, administrator trace actions, persistent settings, ledger inspection, trace lineage, and Honeycomb links. - Existing server trace authorization and retention tests remain the backend authority. - Local tests were not run. The requested verification policy uses GitHub Actions for this series. - `git diff --check runner/task-workspace-experience...HEAD` passes. - The delta contains 54 files. ## Risks - Raw provider traces can contain sensitive provider data. - Existing server authorization controls access, reveal, download, retention, and deletion. - The UI gates trace actions by administrator access and developer mode. - All new instance settings remain off by default. - Fresh Paperclip Runner configuration remains Codex-only. - Direct adapters and legacy task behavior do not change in this PR. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with GPT-5.6. The work used high-reasoning agent mode, repository tools, GitHub tools, and parallel code-audit agents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
300a89ec13 |
Detect the qualifier-less Claude usage-limit message in quota classification (#12475)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The heartbeat runtime classifies adapter run failures, and the
recovery service uses that classification to decide between automatic
retry, a timed provider-quota wait, and a board escalation
> - The Claude CLI changed its subscription-limit stop message to
"You've hit your limit · resets 2:30am (UTC)", and no quota matcher
knows this qualifier-less wording
> - A limit-hit run therefore classifies as `adapter_failed` (or
`claude_auth_required`), recovery burns its continuation retries against
a hard limit, and the issue blocks with the opaque "No live execution
path" notice instead of waiting for the reset and retrying automatically
> - This pull request teaches the adapter and the recovery service the
new wording, and titles stranded-escalation notices from the classified
run error code so operators see the cause at a glance
> - The benefit is that usage-limit stops self-heal at the provider
reset time, and the notices that do post say "Error: usage limit
reached" or "Error: not logged in to Claude" instead of a generic title
## Linked Issues or Issue Description
No public issue exists; the underlying problem follows the bug template:
**What happened?**
On a staging deployment, an assigned `in_progress` issue hit the Claude
subscription usage limit. The run recorded the error `Claude run failed:
subtype=success: You've hit your limit · resets 2:30am (UTC)`. The
automatic continuation retry failed the same way in 34 seconds with
`errorCode: adapter_failed`. Terminal-run recovery then escalated: the
issue moved to `blocked` with the notice "No live execution path" and a
board-owned recovery action. The notice gave the operator no indication
that the cause was a usage limit with a known reset time.
**Expected behavior**
A usage-limit stop classifies as `provider_quota` with the reset clock
parsed into `retryNotBefore`. The recovery service takes its
provider-quota wait path: a system-owned recovery action that waits for
the reset time and retries the original assignee automatically. If an
escalation notice does post, its title names the classified cause.
**Steps to reproduce**
1. Run a `claude_local` agent on an issue until the Claude subscription
limit is hit, so the CLI result is "You've hit your limit · resets
\<time\> (UTC)".
2. Let terminal-run recovery retry the continuation.
3. Observe the issue block with the "No live execution path" notice
instead of a timed quota wait. `classifyAdapterFailureForRecovery`
returns `null` for the recorded error text; `CLAUDE_PROVIDER_QUOTA_RE`
and `PROVIDER_QUOTA_ERROR_RE` both fail to match it.
## What Changed
- `CLAUDE_PROVIDER_QUOTA_RE` and `CLAUDE_EXTRA_USAGE_RESET_RE`
(claude-local adapter) accept "you've hit your limit" with no qualifier,
alongside the existing "session"/"usage" wordings, so the run classifies
as `provider_quota` and the reset clock lands in `retryNotBefore`.
- `PROVIDER_QUOTA_ERROR_RE` and `isProviderQuotaRecovery` (recovery
service) accept the same wording, so runs recorded before the adapter
fix (errorCode `adapter_failed` with the limit text in the error) also
route to the quota wait.
- `parseProviderQuotaClockReset` parses the "resets 2:30am (UTC)" clock
shape alongside the existing "try again at" shape.
- `buildStrandedRecoveryEscalationNotice` titles the notice from the
source run's classified error code when one is mapped: `provider_quota`
→ "Error: usage limit reached", `claude_auth_required` → "Error: not
logged in to Claude", `acpx_auth_required` → "Error: agent login
required". The raw failure text stays withheld from the issue thread;
only the server-classified code is surfaced. Unmapped codes keep the
existing seed/cause titles.
## Verification
- `pnpm vitest run
packages/adapters/claude-local/src/server/parse.test.ts
server/src/services/recovery/provider-failure-classification.test.ts
server/src/services/recovery/stranded-notice.test.ts` — 72 tests pass,
including 5 new cases that use the exact new CLI message.
- `pnpm vitest run server/src/__tests__/issue-recovery-actions.test.ts
server/src/__tests__/heartbeat-retry-scheduling.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts` — 189 tests
pass (no reroute regressions from the widened matchers).
- `tsc --noEmit` clean for `@paperclipai/adapter-claude-local` and
`@paperclipai/server`.
## Risks
- Low risk. The regex widenings are additive; every previously matched
wording still matches, and the existing negative test ("Workspace
storage capacity limit reached." stays unclassified) still passes.
- Behavioral shift, intended: an `adapter_failed` run whose error text
is the new limit wording now routes to the silent system-owned quota
wait instead of a board escalation. This matches how the older limit
wordings already behave.
- The notice title change only affects escalations whose source run
carries one of the three mapped error codes; all other notices render
exactly as before.
## Model Used
Claude Fable 5 (`claude-fable-5`, Claude Code CLI, extended thinking
with tool use).
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
2e5a24e177 |
feat(runner): add qualified OpenCode runtime (#12588)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Runner provides a durable execution boundary for supported providers. > - The current production runtime supports Codex but cannot execute OpenCode sessions. > - OpenCode needs a qualified transport, strict input mapping, and normalized events. > - This pull request adds the OpenCode runtime as one isolated provider unit. > - The benefit is a reviewable provider expansion that does not weaken the existing Codex path. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: packages/paperclip-runner and the Codex-local adapter configuration contract. **Problem or motivation** Paperclip Runner has provider-neutral contracts, but the production backend factory cannot start a qualified OpenCode session. This blocks OpenCode from using the durable runner path. **Proposed solution** Add the qualified OpenCode app-server proxy, driver, MCP bridge, backend, fixtures, and factory wiring. Keep existing Codex behavior unchanged. **Alternatives considered** Keeping OpenCode only on the direct adapter path would avoid this runtime work, but it would not provide durable runner recovery or normalized provider events. **Roadmap alignment** ROADMAP.md does not list a conflicting provider-runtime project. This change extends the existing Paperclip Runner architecture. ## What Changed - Added the qualified OpenCode app-server proxy and input queue. - Added collaboration-mode and provider-event normalization. - Added the OpenCode MCP bridge and native session backend. - Added strict fixtures and focused unit coverage. - Added only the package exports and adapter configuration required by this runtime. - Kept deferred SDK, lab, eval, and public package surfaces out of this change. ## Verification - GitHub Actions is the authoritative verification environment for this PR. - Run the package type checks and focused OpenCode tests in CI. - Run repository typecheck, test, build, security, and policy gates through the stack-aware workflow. - Local tests were not run because this checkout is resource constrained. ## Risks - OpenCode protocol changes could affect event normalization or recovery. - The driver fails closed on malformed input and unsupported runtime behavior. - Existing Codex selection remains unchanged unless the stored provider is OpenCode. > For core feature work, check [ROADMAP.md](ROADMAP.md) first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected. ## Model Used OpenAI Codex, GPT-5.6, with repository tools, code execution, and parallel agent review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass — GitHub Actions is authoritative for this resource-constrained checkout - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack - Position: 1 of 4 - Base: master - Next: additional qualified provider runtimes |
||
|
|
b3343dbd64 |
feat(connections): add self-serve intent runtime (#12345)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a governed way to request app connections during issue work. > - The catalog now describes the available providers and setup methods. > - A request must become a durable, company-scoped intent before an operator acts on it. > - This pull request adds that intent runtime across server, agent, CLI, and shared contracts. > - The benefit is a safe bridge from agent need to operator-approved setup. ## Linked Issues or Issue Description Refs #11965 This is stack 7 of 11. It depends on stack 6 and replaces another reviewable part of #11965. ## What Changed - Add connection intent types, validation, service logic, and routes. - Add agent runtime tools and CLI support for connection requests. - Add issue-thread interaction support for connection intents. - Add runtime, route, adapter, and contract tests. - Hold the final resolved-continuation row lock through asynchronous adapter preparation until an actual process spawn, so parking or reassignment cannot cross that boundary. - Report Hermes Gateway's first remote run request through the shared dispatch hook so the resolved-intent lock is released at the true dispatch boundary. - Revalidate the addressed user's live non-viewer membership and connection-management authority for every intent mutation, including OAuth completion. ## Verification - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/tool-access-service.test.ts` - Result: 176 tests passed. - `pnpm build` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` (32 passed; includes non-process dispatch lock-release coverage) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/connection-intents-service.test.ts -t "addressed-user mutation"` (1 passed) - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/tool-access-service.test.ts -t "binds OAuth callback completion to the initiating board session"` (1 passed) - `pnpm --filter @paperclipai/hermes-paperclip-adapter test -- src/gateway/server/execute.test.ts` (23 passed; includes dispatch-hook ordering and exactly-once coverage) - `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck` ## Risks - A malformed intent could create an unusable operator request. - Validators and company checks reject invalid or cross-company requests. - The final continuation gate holds the issue row lock through adapter preparation until process or remote dispatch; later operator changes use the normal active-run interruption path. - The change does not add a database migration. > I checked `ROADMAP.md`. This stack continues the existing app connection work from #11965 and does not duplicate another planned item. ## Model Used OpenAI Codex, GPT-5. The runtime model ID and context window were not exposed. The model used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have linked the public source pull request with `Refs #` - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a20a4944ec |
feat: add Grok device login to the sandbox login panel (#12469)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip uses adapters to connect agents and model providers to its control plane > - The sandbox login panel supports displayed-code login for selected adapters > - Grok users need the same login path and a private credential home for later runs > - This pull request adds Grok support to the shared device-login path and preserves the existing Codex path > - The benefit is one secure login flow for both adapters with company-scoped credential storage ## Linked Issues or Issue Description **Agent or provider** Grok Local needs displayed-code login support in the sandbox login panel. **Why this adapter is useful** This change lets users sign in to Grok from the sandbox login panel. It also gives later Grok runs access to the stored credential. **How the agent is invoked** The Grok local adapter uses its login command through the shared displayed-code login flow. Later runs receive the managed home through `GROK_HOME`. **Additional context** The change uses adapter-scoped login lifecycle handling. It stores the credential in a company-scoped directory with mode `0700`, and it stores the credential file with mode `0600`. ## What Changed - Rename the shared device-login modules to adapter-neutral names. - Scope the shared login lifecycle to a closed adapter set. - Return the device-login URL that the provider prints. - Add the Grok prompt parser, login command, capability, and login panel entry. - Store the Grok credential in a private, company-scoped home directory. - Pass `GROK_HOME` to later Grok runs. - Add tests for the Grok adapter, the Daytona sandbox provider, the server login path, and the user interface. ## Verification - Run `pnpm vitest run packages/adapters/grok-local/src/server/adapter-auth-promotion.test.ts`. - Run the Grok adapter package suite. - Run the Daytona sandbox provider suite. - Run the server device-login suites. - Run the user interface suite. - Confirm the full CI suite passes. ## Risks The change extends shared login lifecycle code to another adapter. A regression could affect Codex login. The credential path uses explicit `chmod` calls to keep the directory at mode `0700` and the file at mode `0600`. ## Model Used OpenAI Codex, GPT-5. The runtime used tool calls and code review support. The runtime did not provide a context-window value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
06cd21ed0f |
fix(observability): declare the optional OpenTelemetry peer dependencies (#12249)
## Thinking Path > - Paperclip manages AI agents for work. > - Paperclip includes an observability path that operators can enable for tracing. > - The server loads several OpenTelemetry packages only when tracing is enabled. > - The documentation calls these packages optional peer dependencies, but the server manifest does not declare them. > - This gap hides supported versions and stops Dependabot from maintaining the packages. > - This pull request aligns package metadata, runtime checks, and documentation with the opt-in tracing design. > - The change gives operators clear installation behavior and keeps the no-op default. ## Linked Issues or Issue Description This pull request fixes a package metadata and installation defect. Related observability work appears in [#8476](https://github.com/paperclipai/paperclip/pull/8476) and [#9672](https://github.com/paperclipai/paperclip/pull/9672). The server documentation described optional OpenTelemetry peer dependencies, but `server/package.json` did not declare them. Package managers and Dependabot could not see the supported version ranges. The UI and Claude local adapter also relied on automatic peer installation for `yjs` and `@anthropic-ai/sdk`. The package manifests now declare the optional runtime packages. A default install does not install optional tracing peers. The server keeps its no-op behavior when tracing is disabled or a peer is absent. ## What Changed - Add seven optional OpenTelemetry packages to `server/package.json` and mark each package as optional. - Keep `@opentelemetry/api` as a normal dependency for the no-op interface. - Disable automatic peer installation in `.npmrc`. - Declare `yjs` for the UI package and `@anthropic-ai/sdk` for the Claude local adapter. - Check declared peer versions before the server loads a dynamic OpenTelemetry import. - Keep the endpoint gate, dynamic imports, and fail-open behavior unchanged. - Update the observability and README documentation. - Tell Dependabot that its npm parser does not read `peerDependencies`. ## Verification - Targeted server tests pass: 34 passed and 2 skipped. - The skipped tests require the real OpenTelemetry SDK and remain pre-existing. - The pull request workflow regenerates the lockfile because manifest files and `.npmrc` changed. - The policy job confirms that the pull request does not include `pnpm-lock.yaml`. - GitHub checks pass except `security/snyk (cryppadotta)`, which remains pending after its authorized wait cap. - Greptile Review reports 5/5 with no open findings. - Server typecheck passes. ## Risks - Optional peers can produce a diagnostic when the installed version does not match the declared range. - A missing optional peer does not stop the server. - Disabling automatic peer installation can expose undeclared package use in other workspaces. - This pull request declares the affected packages and adds tests for the changed behavior. - This pull request makes no database or API changes. ## Model Used OpenAI Codex, GPT-5, with repository inspection and pull request preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes / Closes / Refs OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a9d0927fe8 |
fix(adapters): restore Paperclip skill for legacy runners (#12225)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy local adapters run agents that use the Paperclip skill for the control-plane workflow. > - PR #7029 removed the required-skill fallback and made runtime skill selection depend only on stored preferences. > - No migration or runtime fallback replaced that behavior for existing agents or non-CEO agents. > - PR #12138 added core skills to new CEOs, and PR #12147 added Claude skill discovery. These changes did not mount the operational skill for all legacy agents. > - This pull request makes the operational skill a legacy adapter runtime invariant. It keeps all other skills configurable. > - The native runner stays unchanged because its protocol supplies the control-plane contract. > - The benefit is that new and existing legacy agents can always operate through Paperclip. ## Linked Issues or Issue Description Refs #7029 Refs #12138 Refs #12147 **What happened?** A skill-capable legacy local agent could start without `paperclipai/paperclip/paperclip`. This happened when the agent had no stored skill preference. An explicit empty preference also removed the skill. The agent then reported that the Paperclip skill was not available. **Expected behavior** Every skill-capable legacy local adapter must mount the Paperclip operational skill when the runtime inventory contains it. Optional skills must remain configurable. The native runner must keep its current protocol-based behavior. **Steps to reproduce** 1. Create a non-CEO `codex_local` agent without `paperclipSkillSync` preferences. 2. Start a legacy heartbeat. 3. Inspect the managed `CODEX_HOME/skills` directory. 4. Observe that the Paperclip skill is absent before this change. **Paperclip version or commit** The problem reproduces on `master` before this pull request. PR #7029 introduced the configured-only selection behavior. **Deployment mode** Local development and self-hosted legacy local adapters. ## What Changed - Added a shared legacy skill resolver that always selects the canonical Paperclip operational skill when it is available. - Applied the resolver to direct adapter execution, ACPX execution, skill snapshots, and persistent skill sync. - Added Hermes skill materialization at sync and run boundaries. - Aligned Cursor, Gemini, and OpenCode execution-time injection with the configured child `HOME`. - Made Hermes stop execution when another installation blocks the required operational skill. - Kept optional skills controlled by `paperclipSkillSync.desiredSkills`. - Kept `paperclip_runner` on the configurable-only resolver. - Added regression coverage for missing preferences, empty preferences, each skill-capable legacy adapter, ACPX, Hermes, and native runner isolation. - Documented the legacy runtime invariant. ## Verification - `pnpm -r typecheck` passed on the pushed commit. - `pnpm build` passed on the pushed commit. - The adapter utility regression suites passed: 236 tests. - The changed server adapter suites passed: 48 tests across 12 files. - The OpenCode adapter suite passed: 8 tests. - The Hermes adapter suite passed: 7 tests. - `git diff --check` passed. - `pnpm test:run` is not clean on this macOS host. The command reported failures in unchanged workspace and filesystem suites. An isolated rerun of `company-skills.test.ts` and `company-skills-service.test.ts` reproduced 11 failures because macOS resolved `/var/...` paths as `/private/var/...`. The changed adapter suites pass independently. ## Risks - This change deliberately makes the operational skill non-removable for skill-capable legacy local adapters. - Existing agents receive the skill on their next list, sync, or run boundary. No database migration is required. - The resolver does not create a skill when the runtime inventory does not contain the canonical entry. - Hermes aborts a run if another installation occupies the required operational skill target. - Hermes removes only an undesired Paperclip-owned symlink that still points to the known Paperclip source. - The native runner does not receive the legacy default. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5. The exact serving model ID and context window were not exposed. The agent used reasoning, tool use, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
821573ede8 |
refactor(adapter-utils): extract the shared workspace-restore teardown factory (#12196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters run workspace restore steps when an ACP run ends. > - Claude, Codex, and Gemini each kept a near-identical teardown closure. > - Duplicate closures require the same defect fix in three files. > - This pull request adds one shared workspace-restore teardown factory and keeps each adapter's message strings. > - The benefit is one tested restore-failure path with the same output and outcome for all three adapters. ## Linked Issues or Issue Description **What existing behavior does this improve?** The Claude, Codex, and Gemini ACP adapters restore the workspace during teardown and report restore failures with an allowlisted message. **Subsystem affected** `packages/adapters/` and `packages/adapter-utils/`. **Current behavior** Each adapter keeps a near-identical closure. The closure logs a start line, restores the workspace, classifies errors, and logs a fixed failure line. **Proposed behavior** A shared `createWorkspaceRestoreTeardown` factory owns the common steps. Each adapter passes its staged runtime, log sink, start line, and failure prefix. **Reason and benefit** The shared factory removes duplicate error handling. One tested implementation now preserves the existing output and outcome for all three adapters. **Breaking changes** None. The refactor preserves the emitted lines and returned outcomes. **Additional context** This pull request contains no public issue reference because no related public issue was found. ## What Changed - Add `createWorkspaceRestoreTeardown` to `packages/adapter-utils`. - Move the shared restore, classify, and allowlisted log flow into the factory. - Update the Claude, Codex, and Gemini ACP adapters to call the factory. - Add a table-driven test for all three message pairs. - Keep one end-to-end restore-failure regression test per adapter. ## Verification - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/adapter-gemini-local typecheck` - `npx vitest run packages/adapter-utils/src/workspace-restore-teardown.test.ts` - `npx vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts` - `npx vitest run packages/adapters/claude-local/src/server/acp.test.ts` - `npx vitest run packages/adapters/codex-local/src/server/acp.test.ts` - `npx vitest run packages/adapters/gemini-local/src/server/acp.test.ts` - Continuous integration must pass before merge, except for the known pre-existing failures listed in the handoff. ## Risks Low risk. This change moves shared code without changing behavior. The adapter-specific message strings remain unchanged. ## Model Used OpenAI GPT-5, exact model ID `gpt-5`, tool use and code review assistance. The context window size was not provided by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0cedb45df3 |
build(deps-dev): bump typescript from 5.9.3 to 7.0.2 (#11880)
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 7.0.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/microsoft/TypeScript/releases">typescript's releases</a>.</em></p> <blockquote> <h2>TypeScript 7.0.2</h2> <p><a href="https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/">https://devblogs.microsoft.com/typescript/announcing-typescript-7-0/</a></p> <p>This tag was originally released at: <a href="https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2">https://github.com/microsoft/typescript-go/releases/tag/typescript%2Fv7.0.2</a></p> <h2>TypeScript 6.0.3</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.3%22">fixed issues query for TypeScript 6.0.3 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.2%22">fixed issues query for TypeScript 6.0.2 (Stable)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0.1 RC</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-rc/">release announcement blog post</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22">fixed issues query for TypeScript 6.0.0 (Beta)</a>.</li> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.1%22">fixed issues query for TypeScript 6.0.1 (RC)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> <h2>TypeScript 6.0 Beta</h2> <p>For release notes, check out the <a href="https://devblogs.microsoft.com/typescript/announcing-typescript-6-0-beta/">release announcement</a>.</p> <ul> <li><a href="https://github.com/Microsoft/TypeScript/issues?utf8=%E2%9C%93&q=milestone%3A%22TypeScript+6.0.0%22+is%3Aclosed+">fixed issues query for Typescript 6.0.0 (Beta)</a>.</li> </ul> <p>Downloads are available on:</p> <ul> <li><a href="https://www.npmjs.com/package/typescript">npm</a></li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/microsoft/TypeScript/commit/1e4744d68260a7cb91b62b12edc3f6a2187faaf1"><code>1e4744d</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/a5a219c3b5da0db4fa0ecf6c0b1f588c9af9c669"><code>a5a219c</code></a><code>microsoft/typescript-go#4558</code></li> <li><a href="https://github.com/microsoft/TypeScript/commit/ecfe30dce91368d52c9a49b6095bb0b673a238f8"><code>ecfe30d</code></a> Update status localization</li> <li><a href="https://github.com/microsoft/TypeScript/commit/5de25b5f8fec2ca35eadaed041f1f06d2e214895"><code>5de25b5</code></a> Hide executable name in TypeScript status</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d7ce74a75da2b80e8201506a1599c06549432b93"><code>d7ce74a</code></a> Show bundled TypeScript version for packaged servers</li> <li><a href="https://github.com/microsoft/TypeScript/commit/29be66a607707f90d7a53103a4469bb3015a4d54"><code>29be66a</code></a> Correct TS 7 release version to 7.0.2</li> <li><a href="https://github.com/microsoft/TypeScript/commit/ed2bd1bfa4aac5211ce4bc58fcd1313c7eddc8ff"><code>ed2bd1b</code></a> Merge branch 'main' into ts7-release</li> <li><a href="https://github.com/microsoft/TypeScript/commit/887307575c58ea640dbeba3b4e8fdb6347cd3044"><code>8873075</code></a> Bump the github-actions group across 1 directory with 3 updates (microsoft/ty...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/9427131ae2d4e230a90ee8a09daac4e75da3e311"><code>9427131</code></a> Set up stable / nightly extension split, other prep (microsoft/typescript-go#...</li> <li><a href="https://github.com/microsoft/TypeScript/commit/d4eaca5460a1f5f02a829e62706794b0a6fb903e"><code>d4eaca5</code></a><code>microsoft/typescript-go#4549</code></li> <li>Additional commits viewable in <a href="https://github.com/microsoft/TypeScript/compare/v5.9.3...v7.0.2">compare view</a></li> </ul> </details> <details> <summary>Maintainer changes</summary> <p>This version was pushed to npm by <a href="https://www.npmjs.com/~microsoft1es">microsoft1es</a>, a new releaser for typescript since your current version.</p> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
822e0aed93 |
fix(adapter-utils): move the workspace-restore merge lock to an instance-scoped root and surface restore failures on the run (#12187)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters restore sandbox work into project workspaces after a run > - The restore lock used the target workspace parent, which can reject writes > - The teardown then hid restore errors, so a run could report success with lost work > - This pull request moves the lock into an instance-scoped root and reports safe restore failure codes > - The benefit is reliable restore coordination and visible failure evidence without changing run success semantics ## Linked Issues or Issue Description Refs: #10914 ## What Changed - Move the workspace-restore merge lock into a private, instance-scoped root. - Derive the lock key from the canonical target path with SHA-256. - Resolve the lock root from the caller environment and reject unsafe root types. - Classify restore failures with three allowlisted codes. - Add the failure code to run result JSON without exposing a host path or process identifier. - Keep restore failure fail-open for the run exit code and run status. ## Verification - Run `npx vitest run packages/adapter-utils/src/workspace-restore-merge.test.ts`. - Run `npx vitest run packages/adapter-utils/src/acpx-engine/run-fault-matrix.test.ts`. - Run the four Codex credential suites. - Confirm the branch includes the current `master` commit and no manual lockfile edit. - Confirm all pull request checks and the Greptile review reach a terminal green state. ## Risks - The lock path changes for workspace restore and removes the sibling-directory fallback. - A misconfigured or inaccessible instance home can still stop lock setup. - Restore remains fail-open, so callers must inspect the result evidence when a restore fails. ## Model Used OpenAI GPT-5. The model used tool calls and code execution to validate and route an author-provided change. The implementing engineer authored the code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
397de98193 |
feat(runner): add flagged Codex execution adapter (#12188)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Paperclip Runner now has protocol, provider, tool, package, persistence, and hidden server boundaries. > - The server still cannot select that path for a real agent heartbeat. > - A new runtime must not change any existing direct adapter. > - An experimental runtime must fail closed when its rollout flag is off. > - This pull request adds one guarded Codex vertical slice through runnerd. > - The benefit is a production-built runner path that users cannot start by default. ## Linked Issues or Issue Description Refs #11962 Refs #12111 Refs #12169 Refs #12176 **Subsystem affected** Cross-cutting. The change affects the runner package, server orchestration, shared settings, and adapter configuration UI. **Problem or motivation** The hidden PRP coordinator cannot execute a real heartbeat. The application also needs an explicit rollout boundary before it can expose the experimental runner. Existing direct adapters must keep their current execution and finalization behavior. **Proposed solution** Add `paperclip_runner` as a Codex-only adapter behind the default-off `enableNativeRunner` instance flag. Select the native runtime only for that adapter. Persist the run binding before runnerd starts. Wait for the durable PRP result and terminal event. Resume the real Codex provider thread on later heartbeats. Keep persisted native runs readable and recoverable after the flag changes. **Alternatives considered** The server could route `codex_local` through runnerd. That option would change an existing adapter and weaken rollback safety. The server could expose all providers now. That option would add unreviewed provider behavior. The build could depend on a prebuilt runner binary. That option would make source builds architecture-dependent and difficult to verify. **Roadmap alignment** This work supports the shipped enforced-outcomes, governed-tool, and self-healing-run milestones. It does not add a new roadmap surface. It is the guarded execution step after the merged hidden runner boundaries. **Additional context** This is the next replacement for the closed large runner pull request. Task-thread presentation remains a separate follow-up so this change can preserve the current direct-adapter UI. ## What Changed - Add `paperclip_runner` as an explicit Codex-only adapter. - Add the default-off `enableNativeRunner` instance flag. - Reject fresh create, hire, import, switch, and execution requests while the flag is off. - Allow edits to persisted runner agents while the flag is off. - Recover an already persisted native run even after the flag is disabled. - Keep every built-in direct adapter on its existing runtime path. - Persist an immutable native run binding and revisioned completion contract before runnerd starts. - Execute server to PRP to runnerd to Codex to server through the hidden coordinator. - Validate the durable result against the terminal event and exact completion criteria before finalization. - Preserve the Codex provider thread ID and use `thread/resume` on the next heartbeat. - Strip unsupported Codex configuration fields from the experimental adapter. - Build a target-native release runner binary from source and vendor it into the server distribution. - Install Rust only in the Docker build stage. Do not add a workflow or lockfile change. - Stop the runner process group on completion, cancellation, and forced shutdown. ## Verification - Run `pnpm --filter @paperclipai/paperclip-runner check:all`. All 69 TypeScript tests and 58 Rust tests pass. Protocol, conformance, replay, formatting, and generated-file checks pass. - Run the 12 focused adapter, settings, runtime-selection, coordinator, direct-isolation, and real Codex integration test files. All 186 tests pass. - The real integration test uses PostgreSQL, HTTP, WebSocket, runnerd, and a fake Codex app server. It proves one `thread/start` followed by one `thread/resume`. - Run `pnpm -r typecheck`. - Run `pnpm build`. - Run `pnpm check:token-gates`. - Build the Docker `build` target from a clean context. Confirm that the server distribution contains an executable `paperclip-runnerd` built with Debian Rust 1.85. - Start the server through the source-mode tsx entry point with the package `dist` directory absent. Confirm the vendor shim resolves source exports and the server boots. - Run `pnpm test:run` twice. On this macOS host, 405 files pass and 1 file skips. Eight untouched workspace and loopback tests fail because macOS resolves `/tmp` and `/var` through `/private` and because PID-derived test ports exceed 65535. Linux CI must pass the full suite. - Confirm that the diff contains 52 files. Confirm that it contains no `.github` or `pnpm-lock.yaml` change. ## Risks - The feature flag is off by default. A fresh native start fails with a stable error while the flag is off. - A persisted native run remains recoverable after the flag changes. This prevents rollout changes from corrupting recorded work. - Only local Codex execution is accepted. Other providers and remote work modes fail closed. - Existing direct adapters do not start runnerd, create native rows, use native status arbitration, or enter native finalization. - The runner receives its one-use bootstrap ticket through the child environment. The server does not put the ticket in command arguments or logs. - The server validates the company, task, agent, run, runner, session, completion contract, result, and terminal binding before it accepts completion. - The build compiles a target-native Rust binary. Cross-platform release packaging remains a later concern. Source builds and Docker builds compile for their current target. - Docker needs enough build memory for the existing server TypeScript compile. The Docker build stage sets a 4 GB V8 heap limit. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with GPT-5. The exact deployment ID and context-window size are not exposed. The model used agentic reasoning, repository tools, code execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and applicable tests pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
243430f76e |
feat: agents see the company skill library at runtime (#12147)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent's runtime mounts only its own enabled skills; nothing tells
the model what else the company skill library holds
> - From inside a sandbox, "installed but not enabled for me" and "does
not exist" look identical, so agents tell users freshly installed skills
are not installed
> - This pull request renders the library as a deterministic markdown
section appended to claude-local agent instructions, and adds a
paperclipListSkills MCP tool
> - The benefit is that agents report the true state ("installed, not
enabled for me — ask an operator to enable it") instead of a false
negative
## Linked Issues or Issue Description
**What existing behavior does this improve?**
How agents reason about the company skill library at runtime.
**Subsystem affected**
`packages/adapter-utils` (new pure builder),
`packages/adapters/claude-local` (instructions append),
`packages/mcp-server` (new tool).
**Current behavior**
The runtime hands adapters the full library list, but only the agent's
enabled skills are mounted, and no prompt content or MCP tool describes
the rest. Agents inspect their sandbox, find nothing, and report
installed skills as not installed.
**Proposed behavior**
A "Company skill library" markdown section lists every skill as
`enabled`, `installed, not enabled for you`, or `enabled but
unavailable: <cause>`, with instructions to report the not-enabled state
accurately and ask an operator to enable it. claude-local appends it to
the agent instructions text. A `paperclipListSkills` MCP tool exposes
the same list on demand.
**Breaking changes**
None. Other adapters are untouched (they can adopt the builder later);
the manifest is deterministic, so the claude-local prompt-bundle cache
only busts when the library actually changes.
## What Changed
- New `packages/adapter-utils/src/skill-library-manifest.ts` with
`buildSkillLibraryManifestMarkdown` (pure, key-sorted, deterministic;
renders the missing-cause detail from #12146).
- `packages/adapters/claude-local/src/server/execute.ts` appends the
manifest to `combinedInstructionsContents` (creating it when no
instructions file is configured).
- `packages/mcp-server/src/tools.ts` adds `paperclipListSkills` hitting
`GET /companies/:companyId/skills`.
## Verification
- `npx vitest run
packages/adapter-utils/src/skill-library-manifest.test.ts` (from repo
root) — 3 tests: byte-identical output for shuffled input, state
rendering incl. the unavailable cause, change detection.
- `cd packages/mcp-server && npx vitest run` — new tool routing test
passes (13 passed; 1 pre-existing failure on my machine reproduces
unchanged at the branch base).
- `cd packages/adapters/claude-local && npx vitest run` — 244 passed, 1
skipped.
- `pnpm run typecheck` clean in adapter-utils, mcp-server, claude-local.
## Risks
- Prompt growth is one line per installed skill plus a five-line header
— bounded and only present when the library is non-empty. Stacked on
#12146 so the manifest's "enabled but unavailable" state reflects real
materialization failures.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking
and tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
79b464bf9d |
fix(server): surface skill materialization failures instead of dropping the skill (#12146)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Runtime skill listing materializes each company skill's files before
handing them to the agent's adapter
> - A materialization failure was swallowed with catch-to-null, and the
skill silently vanished from the runtime while the library still showed
it installed
> - Operators saw "installed", agents saw nothing, and nobody saw the
cause; on claude-local a missing desired skill could even crash the
prompt-bundle hasher
> - This pull request turns both failure paths into structured "missing"
entries with the real error and makes every adapter skip unmountable
entries explicitly
> - The benefit is that a broken skill shows up as broken, with its
cause, instead of not existing
## Linked Issues or Issue Description
**What happened?**
A company skill whose runtime files fail to materialize (deleted source,
missing stored SKILL.md copy, failed version snapshot) disappears from
`listRuntimeSkillEntries` with no trace. Agent skill snapshots report a
generic "not available" with no cause. On claude-local, a desired skill
whose source path does not exist reaches the prompt-bundle hasher, whose
`fs.lstat` throws and can fail the whole run.
**Expected behavior**
The skill appears with `sourceStatus: "missing"` and a `missingDetail`
carrying the underlying error, snapshots and the UI show it as broken,
and adapters skip it at mount time with a logged warning instead of
crashing or dangling-symlinking.
**Steps to reproduce**
Install a local-path skill referenced by an agent, delete its source
directory contents so the stored SKILL.md copy cannot be recovered, and
start a run: before this change the skill vanishes from the runtime set
silently; on claude-local a pinned-but-unmaterializable version can fail
bundle preparation.
## What Changed
- `server/src/services/company-skills.ts` `resolveRuntimeSkillSource`:
both `.catch(() => null)` sites (version snapshot, runtime
materialization) now return the structured `{status: "missing", source,
detail}` shape the deliberate missing branch already used, with the
underlying error message in `detail`.
- `packages/adapter-utils/src/server-utils.ts`:
`isPaperclipSkillSourceMissing` is exported with a doc comment.
- `packages/adapters/claude-local/src/server/execute.ts`: missing
desired skills are filtered out of the prompt bundle and each one logs a
`[paperclip] Warning` with its detail to the run output.
- `cursor-local`, `gemini-local`, `kimi-local`, `opencode-local`,
`pi-local` `execute.ts`: mount loops (and the cursor/gemini injection
calls) skip missing entries instead of symlinking a nonexistent path.
## Verification
- `cd server && npx vitest run
src/__tests__/company-skills-service.test.ts` — new test pins the
missing-with-cause entry for a failed materialization. Nine pre-existing
project-workspace tests in this file fail on my machine at clean
`master` too (environment-specific); their count is unchanged by this
PR.
- `cd server && npx vitest run
src/__tests__/heartbeat-runtime-skills.test.ts
src/__tests__/claude-local-skill-sync.test.ts
src/__tests__/cursor-local-skill-sync.test.ts
src/__tests__/cursor-local-skill-injection.test.ts
src/__tests__/gemini-local-skill-sync.test.ts` — 12 tests pass.
- `cd packages/adapters/claude-local && npx vitest run` — 244 passed, 1
skipped.
- `pnpm run typecheck` clean in server, adapter-utils, and all six
touched adapters.
## Risks
- Runtime skill entry lists grow by the previously dropped entries (now
flagged missing). All shipped consumers either intersect with desired
sets, already handle `sourceStatus: "missing"`, or now skip missing
entries at mount time. The snapshot layer already understood the missing
shape via the `materializeMissing: false` path, so downstream contracts
are unchanged.
## Model Used
- Claude Fable 5 (`claude-fable-5`, Anthropic) with extended thinking
and tool use, via Claude Code.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
c5382b36ba |
build(deps-dev): bump vite from 6.4.3 to 8.2.2 (#11887)
Bumps [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) from 6.4.3 to 8.2.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitejs/vite/releases">vite's releases</a>.</em></p> <blockquote> <h2>plugin-legacy@8.2.2</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/plugin-legacy@8.2.2/packages/plugin-legacy/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.2.2</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.2.2/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>plugin-legacy@8.2.1</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/plugin-legacy@8.2.1/packages/plugin-legacy/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.2.1</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.2.1/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>create-vite@8.2.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/create-vite@8.2.0/packages/create-vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>plugin-legacy@8.2.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/plugin-legacy@8.2.0/packages/plugin-legacy/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.2.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.2.0/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.2.0-beta.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.2.0-beta.0/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.5</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.5/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.4</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.4/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.3</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.3/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.2</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.2/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.1</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.1/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>create-vite@8.1.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/create-vite@8.1.0/packages/create-vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>plugin-legacy@8.1.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/plugin-legacy@8.1.0/packages/plugin-legacy/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>v8.1.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/v8.1.0/packages/vite/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <h2>plugin-legacy@8.1.0-beta.0</h2> <p>Please refer to <a href="https://github.com/vitejs/vite/blob/plugin-legacy@8.1.0-beta.0/packages/plugin-legacy/CHANGELOG.md">CHANGELOG.md</a> for details.</p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md">vite's changelog</a>.</em></p> <blockquote> <h2><!-- raw HTML omitted --><a href="https://github.com/vitejs/vite/compare/v8.2.1...v8.2.2">8.2.2</a> (2026-08-20)<!-- raw HTML omitted --></h2> <h3>Features</h3> <ul> <li><strong>deps:</strong> widen <code>@vitejs/devtools</code> peer range to v0.5.0 (<a href="https://redirect.github.com/vitejs/vite/issues/23302">#23302</a>) (<a href="https://github.com/vitejs/vite/commit/495d9ff5a7d843ca876a9e49799947a5deb704c7">495d9ff</a>)</li> </ul> <h3>Bug Fixes</h3> <ul> <li><strong>bundled-dev:</strong> handle lazy request error (<a href="https://redirect.github.com/vitejs/vite/issues/23291">#23291</a>) (<a href="https://github.com/vitejs/vite/commit/3ba026dade4af56df08815310d3458fa110f5c5c">3ba026d</a>)</li> <li><strong>bundled-dev:</strong> hot update through circular imports instead of reloading (<a href="https://redirect.github.com/vitejs/vite/issues/23259">#23259</a>) (<a href="https://github.com/vitejs/vite/commit/3dbddefaafc091a879b06f9279296f776691e455">3dbddef</a>)</li> <li><strong>config:</strong> resolve sourcemap paths against sourcemap location (<a href="https://redirect.github.com/vitejs/vite/issues/23239">#23239</a>) (<a href="https://github.com/vitejs/vite/commit/05a003e6a17a84d75f907ea0f1598bc39b8dce6c">05a003e</a>)</li> <li><strong>css:</strong> don't pass empty targets to lightningcss (<a href="https://redirect.github.com/vitejs/vite/issues/23295">#23295</a>) (<a href="https://github.com/vitejs/vite/commit/2804636ff608d105928009d274ffba7cfbe55340">2804636</a>)</li> <li><strong>define:</strong> fix match escaped dots to support $-prefixed define keys (<a href="https://redirect.github.com/vitejs/vite/issues/23249">#23249</a>) (<a href="https://github.com/vitejs/vite/commit/dcf88bd2ad2b1a8845f9029587cc8c825e382d42">dcf88bd</a>)</li> <li><strong>deps:</strong> update all non-major dependencies (<a href="https://redirect.github.com/vitejs/vite/issues/23217">#23217</a>) (<a href="https://github.com/vitejs/vite/commit/ba958bddfc9cabe302c6b34269dcf5c9634531e0">ba958bd</a>)</li> <li><strong>deps:</strong> update rolldown-related dependencies (<a href="https://redirect.github.com/vitejs/vite/issues/23218">#23218</a>) (<a href="https://github.com/vitejs/vite/commit/83ecb2c8059e8ce946a7cc835d4c14ef78aef4fd">83ecb2c</a>)</li> <li><strong>module-runner:</strong> exclude completed modules from in-flight cycle detection (fix <a href="https://redirect.github.com/vitejs/vite/issues/22999">#22999</a>) (<a href="https://redirect.github.com/vitejs/vite/issues/23009">#23009</a>) (<a href="https://github.com/vitejs/vite/commit/d9b10a98db1c293ee64300bd75d568b44c8ae931">d9b10a9</a>)</li> <li><strong>optimizer:</strong> close custom extension analysis bundles (<a href="https://redirect.github.com/vitejs/vite/issues/23207">#23207</a>) (<a href="https://github.com/vitejs/vite/commit/8fb76752836f61224d3095b502fa237b478a06b2">8fb7675</a>)</li> <li>reduce Windows 8.3-short-name detection false-positives (<a href="https://redirect.github.com/vitejs/vite/issues/23066">#23066</a>) (<a href="https://github.com/vitejs/vite/commit/02cffa9e2d38d5d8f12e4043ee9d0f7abb1471e2">02cffa9</a>)</li> <li>respect <code>resolve.preserveSymlinks</code> when resolving root (fix <a href="https://redirect.github.com/vitejs/vite/issues/23197">#23197</a>) (<a href="https://redirect.github.com/vitejs/vite/issues/23198">#23198</a>) (<a href="https://github.com/vitejs/vite/commit/8413052731836d4aaf3eb94a0f25788dd35d2888">8413052</a>)</li> <li><strong>ssr:</strong> rewrite computed key of destructing parameter (<a href="https://redirect.github.com/vitejs/vite/issues/23307">#23307</a>) (<a href="https://github.com/vitejs/vite/commit/9db0b61d4c9c7caad7ea1d9670b637faf2bb6c93">9db0b61</a>)</li> <li><strong>vite:</strong> update outdated upstream file links in license comments (<a href="https://redirect.github.com/vitejs/vite/issues/23285">#23285</a>) (<a href="https://github.com/vitejs/vite/commit/c0f2fc607ee97ee4499337b04826420c00654065">c0f2fc6</a>)</li> </ul> <h3>Documentation</h3> <ul> <li><strong>build:</strong> note cssTarget precedence (<a href="https://redirect.github.com/vitejs/vite/issues/23200">#23200</a>) (<a href="https://github.com/vitejs/vite/commit/a20a35ec0685e374519864d0f41dd5f6e9ba0271">a20a35e</a>)</li> </ul> <h3>Miscellaneous Chores</h3> <ul> <li>fix ts errors in build test cases (<a href="https://redirect.github.com/vitejs/vite/issues/23209">#23209</a>) (<a href="https://github.com/vitejs/vite/commit/a0cfcf72f8ef8bf0f2f11d553333b9bb31f1d316">a0cfcf7</a>)</li> </ul> <h3>Code Refactoring</h3> <ul> <li>use JSON import attributes instead of readFileSync in constants (<a href="https://redirect.github.com/vitejs/vite/issues/23258">#23258</a>) (<a href="https://github.com/vitejs/vite/commit/1d9fa392a43229241f80630236f8552ce8f7cd0f">1d9fa39</a>)</li> <li>use named regex constants over inline literals (<a href="https://redirect.github.com/vitejs/vite/issues/22964">#22964</a>) (<a href="https://github.com/vitejs/vite/commit/5c1c6c609718303202832f706884192e1f1e9223">5c1c6c6</a>)</li> </ul> <h3>Tests</h3> <ul> <li><strong>define:</strong> close rolldown bundler after generate (<a href="https://redirect.github.com/vitejs/vite/issues/23231">#23231</a>) (<a href="https://github.com/vitejs/vite/commit/b4d66fee14d970f45b8a6f3d7d6aee73ca9b88ab">b4d66fe</a>)</li> <li><strong>module-runner:</strong> add TLA circular import case (<a href="https://redirect.github.com/vitejs/vite/issues/23299">#23299</a>) (<a href="https://github.com/vitejs/vite/commit/4a261f242831bef92afd2f1aacfb81eab9dec371">4a261f2</a>)</li> <li><strong>module-runner:</strong> simplify server-hmr tests (<a href="https://redirect.github.com/vitejs/vite/issues/23300">#23300</a>) (<a href="https://github.com/vitejs/vite/commit/599b44b6600ec426e10cd556908d53b027b0c4fb">599b44b</a>)</li> <li><strong>ssr:</strong> add destructing assignment case for moduleRunnerTransform (<a href="https://redirect.github.com/vitejs/vite/issues/23308">#23308</a>) (<a href="https://github.com/vitejs/vite/commit/cb77e2a93bad2a8ece00b4aa0ef507c092582c45">cb77e2a</a>)</li> </ul> <h3>Build System</h3> <ul> <li>use JSON import attributes instead of readFIleSync in rolldown configs (<a href="https://redirect.github.com/vitejs/vite/issues/23251">#23251</a>) (<a href="https://github.com/vitejs/vite/commit/d615bcdb23d96c1ca5ce1ee45e21d8d87381106f">d615bcd</a>)</li> </ul> <h2><!-- raw HTML omitted --><a href="https://github.com/vitejs/vite/compare/v8.2.0...v8.2.1">8.2.1</a> (2026-08-06)<!-- raw HTML omitted --></h2> <h3>Bug Fixes</h3> <ul> <li><strong>build:</strong> make client chunkImportMap work with <code>sharedPlugins: true</code> (<a href="https://redirect.github.com/vitejs/vite/issues/23184">#23184</a>) (<a href="https://github.com/vitejs/vite/commit/15f03073c915d6ffb9a1fda447ef66b02bf5cde8">15f0307</a>)</li> <li><strong>bundled-dev:</strong> inject client script tag before chunk scripts (<a href="https://redirect.github.com/vitejs/vite/issues/23161">#23161</a>) (<a href="https://github.com/vitejs/vite/commit/eac0cc84aa2472a85a19ee84561c1ba71e381a55">eac0cc8</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitejs/vite/commit/de1111ab0be00879b404e7ed3b2a80e264edddc1"><code>de1111a</code></a> release: v8.2.2</li> <li><a href="https://github.com/vitejs/vite/commit/cb77e2a93bad2a8ece00b4aa0ef507c092582c45"><code>cb77e2a</code></a> test(ssr): add destructing assignment case for moduleRunnerTransform (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23308">#23308</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/9db0b61d4c9c7caad7ea1d9670b637faf2bb6c93"><code>9db0b61</code></a> fix(ssr): rewrite computed key of destructing parameter (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23307">#23307</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/8413052731836d4aaf3eb94a0f25788dd35d2888"><code>8413052</code></a> fix: respect <code>resolve.preserveSymlinks</code> when resolving root (fix <a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23197">#23197</a>) (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23">#23</a>...</li> <li><a href="https://github.com/vitejs/vite/commit/05a003e6a17a84d75f907ea0f1598bc39b8dce6c"><code>05a003e</code></a> fix(config): resolve sourcemap paths against sourcemap location (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23239">#23239</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/495d9ff5a7d843ca876a9e49799947a5deb704c7"><code>495d9ff</code></a> feat(deps): widen <code>@vitejs/devtools</code> peer range to v0.5.0 (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23302">#23302</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/1d9fa392a43229241f80630236f8552ce8f7cd0f"><code>1d9fa39</code></a> refactor: use JSON import attributes instead of readFileSync in constants (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/2">#2</a>...</li> <li><a href="https://github.com/vitejs/vite/commit/2804636ff608d105928009d274ffba7cfbe55340"><code>2804636</code></a> fix(css): don't pass empty targets to lightningcss (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23295">#23295</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/599b44b6600ec426e10cd556908d53b027b0c4fb"><code>599b44b</code></a> test(module-runner): simplify server-hmr tests (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23300">#23300</a>)</li> <li><a href="https://github.com/vitejs/vite/commit/4a261f242831bef92afd2f1aacfb81eab9dec371"><code>4a261f2</code></a> test(module-runner): add TLA circular import case (<a href="https://github.com/vitejs/vite/tree/HEAD/packages/vite/issues/23299">#23299</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitejs/vite/commits/v8.2.2/packages/vite">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bef9288669 |
build(deps): bump @agentclientprotocol/claude-agent-acp from 0.69.0 to 0.70.0 (#11873)
Bumps [@agentclientprotocol/claude-agent-acp](https://github.com/agentclientprotocol/claude-agent-acp) from 0.69.0 to 0.70.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/releases">@agentclientprotocol/claude-agent-acp's releases</a>.</em></p> <blockquote> <h2>v0.70.0</h2> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.69.0...v0.70.0">0.70.0</a> (2026-08-17)</h2> <h3>Features</h3> <ul> <li>switch providers for loaded Claude sessions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1002">#1002</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/50a95434e94318456f2d07c3d21aaf3595c3407d">50a9543</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/claude-agent-acp/blob/main/CHANGELOG.md">@agentclientprotocol/claude-agent-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.69.0...v0.70.0">0.70.0</a> (2026-08-17)</h2> <h3>Features</h3> <ul> <li>switch providers for loaded Claude sessions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1002">#1002</a>) (<a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/50a95434e94318456f2d07c3d21aaf3595c3407d">50a9543</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/d0aafb1ca26427285ffaeac8d8a4452fff28e9c3"><code>d0aafb1</code></a> chore(main): release 0.70.0 (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1010">#1010</a>)</li> <li><a href="https://github.com/agentclientprotocol/claude-agent-acp/commit/50a95434e94318456f2d07c3d21aaf3595c3407d"><code>50a9543</code></a> feat: switch providers for loaded Claude sessions (<a href="https://redirect.github.com/agentclientprotocol/claude-agent-acp/issues/1002">#1002</a>)</li> <li>See full diff in <a href="https://github.com/agentclientprotocol/claude-agent-acp/compare/v0.69.0...v0.70.0">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
b67dced1bf |
build(deps): bump @agentclientprotocol/codex-acp from 1.2.0 to 1.6.2 (#11883)
Bumps [@agentclientprotocol/codex-acp](https://github.com/agentclientprotocol/codex-acp) from 1.2.0 to 1.6.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/codex-acp/releases">@agentclientprotocol/codex-acp's releases</a>.</em></p> <blockquote> <h2>v1.6.2</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.6.1...v1.6.2">1.6.2</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>right-size the apt timeouts so a slow mirror still finishes (<a href="https://github.com/agentclientprotocol/codex-acp/commit/86e0772204a07d6fc4a8853c523ceb5006431f88">86e0772</a>)</li> </ul> <h2>v1.6.1</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.6.0...v1.6.1">1.6.1</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>kill stalled apt from outside and serialize the unit suite (<a href="https://github.com/agentclientprotocol/codex-acp/commit/51e011fef27b812b238bf29c2a815f8ad149fa87">51e011f</a>)</li> </ul> <h2>v1.6.0</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.5.1...v1.6.0">1.6.0</a> (2026-08-19)</h2> <h3>Features</h3> <ul> <li>harden release pipeline against hangs and e2e flakes (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/413">#413</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/39af81c29b79a85f878db096f9cb593b6d1c7429">39af81c</a>)</li> </ul> <h2>v1.5.1</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.5.0...v1.5.1">1.5.1</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>update codex to 0.148.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/410">#410</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/3616954dc0e24af83b512adb618d7acbc5b98de5">3616954</a>)</li> </ul> <h2>v1.5.0</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.4.0...v1.5.0">1.5.0</a> (2026-08-17)</h2> <h3>Features</h3> <ul> <li>switch providers for loaded Codex sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/404">#404</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/47b57da5641a04df9aeeedc254a3aef53a9497da">47b57da</a>)</li> </ul> <h2>v1.4.0</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.3.0...v1.4.0">1.4.0</a> (2026-08-16)</h2> <h3>Features</h3> <ul> <li>report changed files to AIR (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/403">#403</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/e305394d3f001f21e600597f41a3bee3d4530762">e305394</a>)</li> </ul> <h2>v1.3.0</h2> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.2.0...v1.3.0">1.3.0</a> (2026-08-14)</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/agentclientprotocol/codex-acp/blob/main/CHANGELOG.md">@agentclientprotocol/codex-acp's changelog</a>.</em></p> <blockquote> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.6.1...v1.6.2">1.6.2</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>right-size the apt timeouts so a slow mirror still finishes (<a href="https://github.com/agentclientprotocol/codex-acp/commit/86e0772204a07d6fc4a8853c523ceb5006431f88">86e0772</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.6.0...v1.6.1">1.6.1</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>kill stalled apt from outside and serialize the unit suite (<a href="https://github.com/agentclientprotocol/codex-acp/commit/51e011fef27b812b238bf29c2a815f8ad149fa87">51e011f</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.5.1...v1.6.0">1.6.0</a> (2026-08-19)</h2> <h3>Features</h3> <ul> <li>harden release pipeline against hangs and e2e flakes (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/413">#413</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/39af81c29b79a85f878db096f9cb593b6d1c7429">39af81c</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.5.0...v1.5.1">1.5.1</a> (2026-08-19)</h2> <h3>Bug Fixes</h3> <ul> <li>update codex to 0.148.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/410">#410</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/3616954dc0e24af83b512adb618d7acbc5b98de5">3616954</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.4.0...v1.5.0">1.5.0</a> (2026-08-17)</h2> <h3>Features</h3> <ul> <li>switch providers for loaded Codex sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/404">#404</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/47b57da5641a04df9aeeedc254a3aef53a9497da">47b57da</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.3.0...v1.4.0">1.4.0</a> (2026-08-16)</h2> <h3>Features</h3> <ul> <li>report changed files to AIR (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/403">#403</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/e305394d3f001f21e600597f41a3bee3d4530762">e305394</a>)</li> </ul> <h2><a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.2.0...v1.3.0">1.3.0</a> (2026-08-14)</h2> <h3>Features</h3> <ul> <li>add versioned context compaction metadata (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/396">#396</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/c4a9311f60a638e3a4b03a475afff1d7678e594f">c4a9311</a>)</li> <li>align typed session failures with AIR protocol (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/393">#393</a>) (<a href="https://github.com/agentclientprotocol/codex-acp/commit/e4fb92fffd8b8b9db9b40591ccbdb375c9f3f525">e4fb92f</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/9780d314d34616b476b1ae451ad31089b3dce49a"><code>9780d31</code></a> chore(main): release 1.6.2 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/417">#417</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/86e0772204a07d6fc4a8853c523ceb5006431f88"><code>86e0772</code></a> fix: right-size the apt timeouts so a slow mirror still finishes</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/096f5a88501db50c4420726e84c39f60f08c457f"><code>096f5a8</code></a> chore(main): release 1.6.1 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/416">#416</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/51e011fef27b812b238bf29c2a815f8ad149fa87"><code>51e011f</code></a> fix: kill stalled apt from outside and serialize the unit suite</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/50bd611451c02868cc2b50bd6a7fc61ae5ef9b41"><code>50bd611</code></a> chore(main): release 1.6.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/414">#414</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/39af81c29b79a85f878db096f9cb593b6d1c7429"><code>39af81c</code></a> feat: harden release pipeline against hangs and e2e flakes (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/413">#413</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/ad658e6ec64e8b70c455b10457ccc34f77173c9b"><code>ad658e6</code></a> chore(main): release 1.5.1 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/412">#412</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/3616954dc0e24af83b512adb618d7acbc5b98de5"><code>3616954</code></a> fix: update codex to 0.148.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/410">#410</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/3d5682722545a4b2d7cfcf8bdabbbfadbdaa37ea"><code>3d56827</code></a> chore(main): release 1.5.0 (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/409">#409</a>)</li> <li><a href="https://github.com/agentclientprotocol/codex-acp/commit/47b57da5641a04df9aeeedc254a3aef53a9497da"><code>47b57da</code></a> feat: switch providers for loaded Codex sessions (<a href="https://redirect.github.com/agentclientprotocol/codex-acp/issues/404">#404</a>)</li> <li>Additional commits viewable in <a href="https://github.com/agentclientprotocol/codex-acp/compare/v1.2.0...v1.6.2">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d1573244b5 |
refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip records first-party events, OpenTelemetry data, and local run-log events > - The code and documents used one term for these three data paths > - This naming made the required review level unclear > - This pull request names each data path in the module names, documents, and code comments > - The benefit is a clear review rule without a runtime change ## Linked Issues or Issue Description **Issue type** Unclear or confusing. **Where is the issue?** `packages/shared/src/telemetry/README.md`, `doc/observability.md`, `doc/run-log-events.md`, and the duplex instrumentation modules. **What's wrong?** The repository used Telemetry for first-party events, OpenTelemetry data, and local run-log events. This usage made the data path and review level unclear. **Suggested fix** Use Telemetry only for Paperclip first-party events. Use Observability for OpenTelemetry data. Use the run log for rows in `heartbeat_run_events`. Related public pull requests: #8476 and #9672. ## What Changed - Rename the duplex instrumentation modules and identifiers from `Telemetry` to `Observability`. - Move the Observability and run-log contracts out of the Telemetry README. - Add `doc/observability.md` and `doc/run-log-events.md` as the canonical documents. - Add a file-path review rule to `AGENTS.md`. - Correct the remaining code comments that name the wrong data path. - Keep all event names, payloads, database records, spans, configuration keys, environment variables, and runtime paths unchanged. ## Verification - `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts` passes. - `npx vitest run packages/adapter-utils/src/published-exports.test.ts` passes. - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes with 42 tests. - `pnpm --filter @paperclipai/adapter-utils typecheck` passes. - `pnpm --filter server typecheck` passes. - The old module name does not remain in TypeScript or JSON files, except for the intentional publication guard. - CI and Greptile checks remain pending after PR creation. ## Risks - The old duplex module subpath no longer has a compatibility shim. The board accepted this intentional hard break. - The new duplex module subpath stays blocked from package publication. - The change has no runtime effect. The main risk is an incorrect document or module reference. ## Model Used OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code review support. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR with the documentation issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
83fefaadd1 |
fix(grok_local): do not warn when the default model sentinel is unavailable (#12062)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each agent runs under an adapter. The `grok_local` adapter runs the Grok Build CLI. > - The adapter has an environment test. It probes the CLI and reports checks to the operator. > - `DEFAULT_GROK_LOCAL_MODEL` is `"grok-build"`. This value is a sentinel. It means "use the Grok CLI's own default model". > - `execute.ts` only passes `--model` when the configured model differs from the sentinel. So the sentinel is never sent to grok. > - The environment test still compared the sentinel to the models that `grok models` lists. Real grok never lists `grok-build`. > - So every probe emitted a false "Configured model not found" warning, even on a correctly configured agent. > - This pull request stops the false warning and keeps the real check for user-set models. > - The benefit is an accurate environment test: operators see a warning only when it is real. ## Linked Issues or Issue Description No public issue exists. The problem, in bug-report form: **What happened?** The `grok_local` environment test always warns `Configured model "grok-build" not found in available models`, even when the agent works. `grok-build` is the default sentinel, not a real model id, and it is never sent to the CLI. **Expected behavior** When the model is left at the default, the test reports the CLI's own default model as info and does not warn. It warns only when a user sets a real model that `grok models` does not list. **Steps to reproduce** 1. Create a `grok_local` agent and leave the model at its default. 2. Run the adapter environment test. 3. See the `grok_model_not_found` warning, although `grok models` and the hello probe succeed. **Agent adapter(s) involved** grok_local (Grok Build CLI). ## What Changed - `packages/adapters/grok-local/src/server/test.ts`: the model check now treats the default sentinel as valid and reports it as info (`Using the Grok CLI's default model (<default>)`). It still warns when an explicitly configured, non-sentinel model is absent from the discovered list. This matches `execute.ts`, which never sends the sentinel to grok. - `packages/adapters/grok-local/src/server/test.test.ts`: adds a test that the default sentinel does not warn when it is absent from the real model list, and a test that a real, unavailable model still warns. ## Verification - `pnpm exec vitest run packages/adapters/grok-local/src/server/test.test.ts` — 5 passed. ## Risks Low risk. The change only affects one adapter's environment-test reporting. It does not change how runs pass `--model`. No schema, no runtime behavior change. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (n/a — no doc change) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
fc9e9b704f |
fix: stop teaching agents to curl literal {id} route templates (#12061)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Adapters inject prompt text that teaches agents how to call the
Paperclip API, including copy-pasteable curl examples
> - Some of those URLs contained brace placeholders like
`/api/issues/{id}/checkout`
> - Agents paste such lines verbatim; the placeholder reaches the server
as `/api/issues/%7Bid%7D` and 404s, and request logs show agents doing
exactly that
> - The acpx engine's API note already avoids this by using
`$PAPERCLIP_TASK_ID`, and its test pins `/api/issues/{id}` out of the
prompt
> - This pull request applies the same standard to the gemini adapter,
the shared prompt template, and the openclaw gateway workflow
> - The benefit is that agents stop burning turns on placeholder 404s
and doc examples stay safe to execute as written
## Linked Issues or Issue Description
No public issue exists for this defect. The description below follows
the bug report template.
**What happened?**
Server request logs show agents issuing `GET /api/issues/%7Bid%7D` — the
literal, percent-encoded text `{id}` — which 404s. The source is adapter
prompt text: the gemini adapter's API note embeds a curl example with
`/api/issues/{id}/checkout` in the URL, the shared agent prompt template
mentions `/api/issues/{issueId}` endpoints, and the harness checkout
notice names `/api/issues/{id}/checkout`. Models copy these strings into
real requests.
**Expected behavior**
URL paths in prompt text must carry environment variables or real ids,
never brace placeholders, in every string an agent might execute
verbatim. Where a placeholder is unavoidable, the prompt must state
explicitly that the literal text must never be sent.
**Steps to reproduce**
1. Give an agent the gemini adapter's API access note.
2. Watch it call `curl ...
"$PAPERCLIP_API_URL/api/issues/{id}/checkout"` as written.
3. The server logs `POST /api/issues/%7Bid%7D/checkout 404`.
## What Changed
- gemini-local's API note curl example now uses `$PAPERCLIP_TASK_ID` and
tells the agent to substitute a real issue id when that variable is
absent — the same convention as the acpx engine's API note.
- The shared agent prompt template (`server-utils.ts`) uses
`$PAPERCLIP_TASK_ID` in its interaction-creation and resume-endpoint
mentions, and the harness checkout notice names `POST
/api/issues/$PAPERCLIP_TASK_ID/checkout`.
- openclaw-gateway's endpoint workflow keeps its `{issueId}`
placeholders — they are defined by its "determine issueId" step — but
now states explicitly that the literal text must never be sent in a URL.
- `server-utils.test.ts` pins the new form and adds negative pins that
keep `/api/issues/{id}` and `/api/issues/{issueId}` out of the shared
prompt template, mirroring the existing acpx-engine negative pin.
- The `confirmation:{issueId}:plan:{revisionId}` idempotency-key
template is untouched: it is a value-construction pattern, not a URL.
## Verification
- `npx vitest run packages/adapter-utils/src/server-utils.test.ts
packages/adapters/gemini-local packages/adapters/openclaw-gateway` — 152
passed. The single failure (`pre-selects gemini-api-key auth in the
managed HOME for sandbox execution`) is a pre-existing
environment-specific failure on the development machine, unrelated to
prompt text; CI is authoritative for it.
- `pnpm --filter @paperclipai/adapter-utils --filter
@paperclipai/adapter-gemini-local --filter
@paperclipai/adapter-openclaw-gateway typecheck`.
## Risks
- Low risk: prompt-text and test changes only; no runtime logic changes.
- Agents that memorized the old example strings keep working — the
routes are unchanged, only the placeholder text in prompts is.
## Model Used
- Claude Fable 5 (Anthropic), model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code (CLI harness), 200k
context window.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
63df7ad2b3 |
feat(login): use the login pseudo-terminal for Codex device login and de-Claude the shared channel (#12020)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters use provider-specific login flows > - Codex device login needs a live pseudo-terminal (PTY), while the shared channel still uses Claude-specific names > - The old streamed-exec path does not provide the prompt transport that Codex needs > - This pull request moves Codex device login to the shared login PTY and removes the dead streamed-exec path > - The benefit is one controlled login transport with fail-closed capability checks and safer credential reads ## Linked Issues or Issue Description **Problem or motivation** Codex device login used a streamed-exec path that did not provide the required prompt transport. The shared login channel also exposed Claude-specific names outside Claude code. **Expected behavior** The host selects a fixed login command from trusted adapter data. Codex login uses the provider login PTY. Providers without that capability fail closed. **Proposed solution** Use a server-controlled session home, create and validate it as a fresh 0700 directory, read credentials from one validated descriptor, and rename shared channel names to the neutral login PTY family. **Alternatives considered** Keep the shared login PTY as the single transport. Do not keep the removed streamed-exec path because it cannot provide the required prompt transport. **Roadmap alignment** This change supports the planned login transport work. It does not add a separate roadmap item. ## What Changed - Route Codex device login through the shared login PTY transport. - Select the login command from a closed internal command key. - Carry a server-controlled session home through the launch contract. - Create and validate the session home as a fresh 0700 directory owned by the login user. - Read the credential file with descriptor-relative, no-follow path walking and final descriptor checks. - Gate the login route and run lease on the provider login PTY capability. - Rename shared channel names to the neutral login PTY family. - Remove the streamed-exec transport value, selector field, driver branch, and related tests. - Hide Codex login in the user interface when the provider lacks the login PTY capability. ## Verification - Server unit suites pass: 89/89. - Adapter-utils suites pass: 262/262. - Codex-local suites pass: 326/326. - Credential-read reader suite passes: 20/20. - Daytona login PTY suite passes: 30/30. - Device-login suites pass: 56/56. - TypeScript checks pass for server, adapter-utils, and UI. - GitHub Actions must pass after pull request creation. - Greptile review must reach 5/5 with no open P2 findings, recommendations, or follow-ups. ## Risks - Providers without a login PTY capability lose Codex login support by design. - The credential read rejects invalid ownership, mode, type, path, and size. - The launch-time sandbox directory race remains outside the threat model because the login runs inside the sandbox and a hostile sandbox already controls its credential. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact context window and reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cc42a67e7e |
fix(adapter-utils): extend the duplex fail-closed run disposition to the CLI lane (#11966)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs agents through adapter execution lanes > - Duplex adapters can lose their control channel before a process completes > - The ACP lane already fails closed, but the CLI lane can report false success > - This pull request applies the same completion rule to the CLI lane and shares the loss code > - The benefit is consistent failure reporting when a duplex channel closes during a run ## Linked Issues or Issue Description **What happened?** A CLI-lane duplex run can lose its control channel before clean process completion. The run can then report `succeeded` with exit code 0 and no error code. **Expected behavior** The execution target must fail closed when the channel dies before clean completion. It must return exit code 1, the typed `duplex_channel_lost` error code, and a short stderr note. **Steps to reproduce** 1. Start a duplex adapter run through the CLI execution lane. 2. Close the duplex control channel before the process completes cleanly. 3. Inspect the run result and error code. **Paperclip version or commit** Commit `5e01523d4eb6df4a20a0bddd05374c9c42225203`. **Deployment mode** Built from source. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Claude Code, Codex, Cursor, Gemini, Kimi, OpenCode, and Pi local adapters. **Database mode** Not database-related. ## What Changed - Add an optional `errorCode` field to `RunProcessResult`. - Add a one-read completion seam to the execution target process options. - Fail closed when a duplex channel dies before clean process completion. - Add `settleRunDisposition()` to atomically read and mark orderly completion. - Share the typed duplex loss error code across the ACP and CLI lanes. - Mark non-success terminal results as orderly completion before teardown. - Wire the seam through the seven duplex adapters. - Add regression tests for channel loss, clean completion, and non-clean terminal results. ## Verification - `npx vitest run packages/adapter-utils/src/execution-target-sandbox.test.ts` — 118 passed. - `npx vitest run packages/adapter-utils/src/acpx-engine/execute.test.ts -t "sandbox duplex run-disposition seam"` — 4 passed. - The author confirmed a clean type-check for `@paperclipai/adapter-utils` and the seven duplex adapter packages. - Pre-existing environment failures remain outside this change. They include `EACCES mkdir '/srv/paperclip'` and remote file-size setup failures. ## Risks The change alters terminal status for CLI duplex runs that lose control before clean completion. The typed error code and stderr note keep the failure visible. The broker marks failed, cancelled, and timed-out results as orderly completion to prevent false loss events during teardown. ## Model Used OpenAI Codex, GPT-5, tool use and code execution, with the standard GPT-5 context window. The model assisted with the implementation and test work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
10d2781a29 |
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox adapters provide controlled execution for untrusted provider environments. > - The sandbox channel needs one persistent duplex transport with strict host control. > - The transport must remain off unless the instance setting and provider capability both allow it. > - The host must detect loss, bound resource use, and expose only safe telemetry. > - This pull request adds the broker, gated selection, kill-switch wiring, fixed observability, and real-process proof. > - The benefit is safer sandbox execution with bounded failure behavior and inspectable transport results. ## Linked Issues or Issue Description No public issue exists for this change. The related pull requests are #11738 and #11750. **Problem or motivation** The sandbox duplex channel needs a host-controlled broker, strict transport gates, bounded provider input, and safe loss telemetry. Without these controls, a provider can cause replay, resource growth, unsafe endpoint selection, or data exposure through telemetry. **Proposed solution** Add a host broker with nested time limits, request limits, one-shot loss, and per-id deduplication. Select duplex transport only when the instance setting and provider capability both equal true. Assign the endpoint and nonce on the host. Reject invalid readiness data and use the file bridge on failure. Add fixed redacted telemetry and a real-process end-to-end test harness. **Alternatives considered** Keep the file bridge as the only transport. This avoids new channel behavior but does not provide persistent duplex operation for supported sandbox providers. **Roadmap alignment** This change supports the Cloud / Sandbox agents section in ROADMAP.md. ## What Changed - Add the duplex bridge broker with bounded forward, response, and gateway wait budgets. - Bound concurrent requests, lifetime requests, and request-id bytes before retention or forwarding. - Select duplex transport only when both required gates are true. - Assign the loopback port and nonce on the host and enforce a liveness-only READY frame. - Fall back to the file bridge after invalid readiness, contamination, bind failure, or timeout. - Carry the kill switch through the server, acpx engine, and six local adapters. - Add fixed, redacted duplex telemetry with a provider allowlist. - Add a real-process end-to-end harness for readiness, round trips, loss, and teardown. - Add regression coverage for limits, loss, UTF-8 splits, concurrency, and telemetry dimensions. ## Verification - Adapter-utils, server, and Daytona typechecks pass locally. - Adapter-utils tests pass, including the codec, broker, execution-target sandbox, and real-process harness. - Server kill-switch tests pass. - Live Daytona tests pass with the required provider key and skip without that key. - The root pnpm-lock.yaml file has no diff. - The branch contains ten commits after origin/master. ## Risks - Duplex transport remains disabled unless both gates equal true. - A provider remains an untrusted boundary and needs least-privilege credentials and quotas. - The server telemetry recorder stays deferred; the default recorder does nothing. - A provider that pre-binds the host port causes a fail-closed fallback to the file bridge. - The change adds no database migration and changes no root lockfile. ## Model Used OpenAI GPT-5, exact model family GPT-5, large context window, reasoning, and tool use. The model assisted with Git handoff validation and PR preparation. The implementation commits came from the engineering worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fbd20b28d3 |
fix(grok-local): stop defaulting --permission-mode to dontAsk (#11898)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The `grok_local` adapter runs the native Grok Build CLI in headless
mode for unattended agent heartbeats
> - Grok CLI 1.0 started to enforce the `dontAsk` permission mode as
deny-by-default, and it takes precedence over `--always-approve`
> - The adapter passes both flags on every run, so each run dies on its
first tool call and is still recorded as a success
> - This pull request removes the `dontAsk` default so unattended runs
rely on `--always-approve` alone
> - The benefit is that `grok_local` agents can execute tools again on
current Grok CLI releases
## Linked Issues or Issue Description
No public issue exists. Description per the bug template:
**What happened?**
Every `grok_local` run on Grok CLI 1.0.x stops on its first tool call.
The stream shows the tool call move from `pending` to `failed` with
"User cancelled the execution for tool `run_terminal_command`", and the
session ends with `stopReason: "cancelled"` after one turn. The CLI
exits 0, so Paperclip records the run as succeeded with no work done,
and the issue lands in missing-disposition recovery.
**Expected behavior**
Unattended runs must auto-approve tool executions. The adapter already
passes `--always-approve` for this.
**Steps to reproduce**
In a clean Linux environment with Grok CLI 1.0.3 and `XAI_API_KEY` set,
run the adapter's exact invocation shape:
`grok --output-format streaming-json --permission-mode dontAsk
--always-approve --disable-web-search --single "Run the shell command:
echo ok"`
The tool call is denied. Drop `--permission-mode dontAsk` (or use
`--permission-mode bypassPermissions`) and the same command executes the
tool. On Grok 0.2.x the original combination worked because the CLI
accepted `dontAsk` without enforcing it; the 0.2.39 embedded docs state
the flag takes effect only for `bypassPermissions` / always-approve.
**Paperclip version or commit**
master (
|
||
|
|
adfbe2d4b9 |
feat(environments): refer to the managed default environment by name, not the sandbox driver key (#11838)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed deployments provision a platform-managed default environment
for agent runs; the UI shows this environment in selectors, the agent
form, run details, and the environments page
> - Those surfaces append the raw driver key to the environment name, so
users see labels like "Paperclip Computer (sandbox)", "Paperclip
Computer · sandbox", and fallback copy such as "Managed sandbox" and
"The sandbox has no ready authentication"
> - "sandbox" is infrastructure vocabulary, not the product name of the
environment; showing it next to the managed environment's name is
confusing and off-brand
> - This pull request renders platform-managed environments by name
alone and rewords the sandbox-phrased copy, while user-created
environments keep the driver suffix so mixed lists stay distinguishable
> - The benefit is that the default environment reads as one clear
product name everywhere, and self-hosted users lose nothing: their own
environments still show the driver
## Linked Issues or Issue Description
**What existing behavior does this improve?**
Display of the platform-managed default environment across the UI.
**Subsystem affected**
UI (environment selectors, agent config form, environments page, agents
page, run details) and the claude-local/codex-local adapter auth checks.
**Current behavior**
The agent form labels the inherited default environment as "Name
(sandbox)". Environment selectors and the environments list render "Name
· sandbox". The agents page describes the environment as "<provider>
sandbox provider". The agent form's fallback label is "Managed sandbox".
Adapter auth checks say "The sandbox has no ready authentication for
this adapter."
**Proposed behavior**
Platform-managed environment rows (`metadata.managedByPaperclip`) render
their name alone. The fallback label is "Paperclip Computer". The agents
page describes managed environments as "Managed by Paperclip". Run
details omit the driver suffix for sandbox-driver environments (the
adjacent Provider entry already identifies the mechanism). Adapter auth
checks say "This environment has no ready authentication for this
adapter."
**Reason and benefit**
The managed environment carries a product name. Appending the raw driver
key ("sandbox") to it is noise and contradicts the product naming.
User-created environments keep the driver suffix, so mixed lists stay
distinguishable.
**Breaking changes**
None. Message text of the auth check is not read programmatically; the
UI keys off `ADAPTER_AUTH_MISSING_CHECK_CODE`. Rows without the managed
marker render exactly as before.
## What Changed
- New `environmentDisplayLabel` helper in
`ui/src/lib/managed-sandbox-environment.ts`: managed rows → name alone;
other rows → "Name · driver".
- `AgentConfigForm`: inherited-default label uses the helper; fallback
copy "Managed sandbox" → "Paperclip Computer"; environment options use
the helper.
- `ProjectProperties`, `CompanyEnvironments`: environment selector
options use the helper; the environments-list row hides the driver
suffix on managed rows; the managed detail page's fallback description
no longer says "sandbox".
- `Agents` page: managed environments are described as "Managed by
Paperclip" instead of "<provider> sandbox provider".
- `CommentThread` run details: the driver suffix is omitted for
sandbox-driver environments.
- claude-local and codex-local adapters: auth-missing check message/hint
reworded from "sandbox" to "environment" (ACP and environment-test
paths); claude-local probe/effort/login hints reworded the same way.
- Run status lines: "Syncing workspace to sandbox", "Exporting git
changes from sandbox", "Starting adapter in sandbox", and friends now
say "environment"; "Finalizing sandbox workspace" → "Finalizing
workspace". Templated transfer-progress lines map the `sandbox`
transport key to "environment" for display (`runtime-progress.ts`).
- Agent form sign-in panel: "Sign in to the sandbox" → "Sign in to the
environment"; "Authenticated. The sandbox has credentials now." → "…The
environment has credentials now."
- Feature catalog + instance settings card: "Managed Sandbox Only" →
"Managed Environment Only" (setting key unchanged; the card keeps its
alphabetical slot).
- Server agents routes: execution-target failure and test-identity copy
no longer say "sandbox"; workspace-mode label "Cloud sandbox" → "Cloud
environment".
- Tests: new `environmentDisplayLabel` unit cases; new `AgentConfigForm`
render case asserting the managed default renders without "(sandbox)" or
"· sandbox"; status-line assertions updated across adapter-utils, server
heartbeat/live-run, and UI chat suites.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `pnpm --filter @paperclipai/adapter-claude-local typecheck` and
`--filter @paperclipai/adapter-codex-local typecheck` — clean.
- `vitest run` for `managed-sandbox-environment.test.ts`,
`AgentConfigForm.render.test.tsx`, `CompanyEnvironments.test.tsx`,
`Agents.test.tsx`, `CommentThread.test.tsx`, `NewAgent.test.tsx` — all
green (118 tests across the two runs).
## Risks
Low risk. Cosmetic label changes only; no data or API changes. Rows
without `metadata.managedByPaperclip` render exactly as before, so
self-hosted deployments with their own environments see no change. The
only self-hosted-visible wording changes are the adapter auth-check
message and the driver suffix omission on sandbox-driver rows in run
details.
## Model Used
- Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code CLI,
extended thinking, tool use.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
docs reference these labels)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
5bc6031f79 |
fix(server,ui,claude-local): verify auth on the adapter Test lane and enforce managed-sandbox tenant binding (#11810)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter Test checks whether an agent adapter can run with its configured environment, and every local-driver adapter (Claude, Codex, Gemini, OpenCode, Pi, Cursor, etc.) shares this Test route and its UI resolution logic > - The Claude ACP Test lane could report pass without checking local or remote authentication, and the shared Test route and UI had gaps in environment binding, probe safety, and managed-sandbox resolution that affect every adapter that uses the Test button, not only Claude > - This pull request verifies authentication on every Claude ACP target, and closes the shared Test-route/UI gaps: tenant-binding on the route, a managed-sandbox-only redirect that matches the real run path, and a three-tier environment resolution in the UI > - The benefit is a truthful Test result with safer probe execution and tenant isolation, for Claude specifically and for every other local adapter that shares this Test surface ## Linked Issues or Issue Description **What happened?** The Claude ACP Test lane returned `status: "pass"` without checking authentication for some local and non-sandbox targets. Separately, the shared `/companies/:companyId/adapters/:type/test-environment` route — used by every local-driver adapter, not only Claude — accepted a foreign environment id, and its UI resolution did not mirror the server's managed-sandbox-only redirect. **Expected behavior** The Test lane checks the resolved credential and hello probe for every Claude ACP target. The shared adapter Test route rejects a foreign environment before it reveals environment details or starts a lease, for any adapter type. The Test's environment resolution (UI and server) matches the real run's three-tier resolution, including the managed-sandbox-only redirect. **Steps to reproduce** 1. Run the Claude ACP Test lane against a local target without a valid credential. 2. Run the adapter Test route with an environment id from another company (any adapter type). 3. Observe the pass result on step 1, or the missing tenant-binding rejection on step 2. **Paperclip version or commit** `933749e01f74e82ce5d315c071be534d04e01158` **Deployment mode** Local dev (`pnpm dev`) and server route tests. **Agent adapter(s) involved** Claude Code directly (the ACP auth-verification work). The tenant-binding guard, managed-sandbox-only redirect, and UI three-tier resolution apply to the shared adapter Test route and affect every local-driver adapter (Codex, Gemini, OpenCode, Pi, Cursor, etc.), not only Claude — see "What Changed" below for the split between Claude-only and shared changes. **Database mode** Not database-related. **Access context** Both board and agent paths use the affected Test surface, for every local-driver adapter. **Additional context** Two commits that were previously bundled into this PR — a `plugin-worker-manager` duplex-channel frame-bound fix and a `workspace-runtime` exit-persist crash fix — are unrelated to the adapter Test lane and have been split out into their own PRs: #11860 and #11861. ## What Changed Claude-only (`packages/adapters/claude-local`): - Verify `CLAUDE_CODE_OAUTH_TOKEN` and run the hello probe for every Claude ACP target. - Keep `adapter_auth_missing` sandbox-only and report missing non-sandbox credentials as a warning. - Add a deny-by-default probe environment builder for the ACP and CLI local probes. - Log only fixed probe context and allowlisted classifications. - Seed the host OAuth token into the hello probe environment. Shared, cross-adapter (`server/src/routes/agents.ts`, `ui/src/lib/adapter-test-environment.ts`, `ui/src/components/AgentConfigForm.tsx`, `ui/src/components/OnboardingWizard.tsx`): - Add a company-binding guard and a binding assertion for the generic `/companies/:companyId/adapters/:type/test-environment` route, so a foreign-company environment id is rejected before any secret resolution or sandbox lease, for every adapter type. - Resolve all three server environment tiers (agent default, instance default, local default) in the UI, and add the managed-sandbox-only redirect so the Test probes the same target a real run would use. - Enforce onboarding Test results: block hire on a failed environment test. - Add regression tests for authentication, tenant binding, probe safety, diagnostics, and UI resolution. ## Verification - Adapter suites pass for the Claude local server probe, remote, ACP, auth, probe environment, and config paths. - Server route tests pass, including the five tenant-binding cases. - UI adapter Test environment resolver tests pass for all three resolution tiers. - Adapter package `tsc --noEmit` exits 0. - Full CI must pass on this pull request. ## Risks The probe environment now denies caller variables by default. A required variable that is not on the allowlist could stop a probe from starting. The route now rejects foreign environment ids with a fixed 403 response. The managed-sandbox-only redirect changes where the Test (and the login affordance) probes for every local-driver adapter under that policy, not only Claude — operators running other local adapters under managed-sandbox-only will see their Test target move from local to the managed sandbox, matching what real runs already do. The change limits secret and diagnostic exposure. ## Model Used Original implementation: OpenAI Codex, GPT-5; exact context window not exposed in that run; tool use and code execution. This revision (commit split and title/description correction): Claude, Sonnet 5 (claude-sonnet-5). The original title and description described this PR as Claude-only; review found it also changes the shared adapter Test route and UI resolution used by every local-driver adapter, and carried two unrelated server fixes. Claude split those two commits into #11860 and #11861 via `git rebase --onto` (verified byte-identical to the original tree minus those commits) and rewrote this description to reflect the actual scope. No functional code in this PR was authored by Claude. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |