## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
4.4 KiB
Adding a Harness
Every Paperclip Runner harness must declare its permission behavior as part of the provider contract. A driver is not complete until its permission modes, maximum non-interactive default, durable request translation, recovery identity, isolation behavior, and conformance coverage are defined.
Current permission catalog
| Provider | Agent configuration key | Supported values | Default |
|---|---|---|---|
| Codex | codexPermissionMode |
never, on-request, untrusted |
never |
| OpenCode | opencodePermissionMode |
allow, ask, deny |
allow |
| ACPX (Claude, Codex) | acpxPermissionMode |
approve-all, approve-paperclip, approve-reads, deny-all |
approve-all |
The browser-safe source of truth for labels, defaults, and configuration
validation is PAPERCLIP_RUNNER_PERMISSION_CAPABILITIES in
@paperclipai/adapter-utils. Native process boundaries validate the pinned
provider value again; they must not silently accept an unknown mode.
“Full auto” means the harness does not pause for a duplicate approval inside the environment assigned to the run. It does not grant host-wide filesystem access, unrestricted network access, Paperclip credentials, or write access in read-only planning mode. Workspace containment, protected-path and symlink checks, network policy, credential bindings, and authenticated PRP authorization remain authoritative and run before any harness auto-approval.
Driver requirements
When adding or upgrading a harness:
- Declare the exact native modes supported by the provider. Choose the provider's highest non-interactive mode as the default for a fresh Paperclip Runner execution.
- Add its labels, configuration key, options, and default to the shared provider-discriminated capability catalog. The agent form must show only the selected provider's setting and retain stored choices when the provider selection changes.
- Pin the effective value in native execution input and provider recovery state. Include it in provider-session compatibility so changing a mode replaces an incompatible warm session rather than mutating an active turn.
- Translate every lower-mode prompt into the canonical runtime-request lifecycle: emit one durable request, recover it after reconnect, accept allow-once, allow-for-session, deny, and cancellation where the provider supports them, and send exactly one provider-native resolution.
- Authorize Paperclip semantic and question tools through authenticated PRP. These tools may bypass duplicate harness approval, but never their control-plane authorization.
- Emit the provider and effective permission mode in session-start diagnostics without including credentials or unredacted provider input.
- Keep standalone local adapters unchanged unless their own contract is
deliberately revised. Paperclip Runner defaults apply only to
paperclip_runnerexecutions.
Compatibility rules
Native execution input v4 pins approvalPolicy for Codex and
permissionMode for OpenCode and ACPX. Persisted v1-v3 executions remain
replayable. Their missing Codex and OpenCode settings use the historical
effective behavior, and legacy ACPX permissionPolicy: "interactive" means
the historical approve-reads behavior. Legacy fields are readable only;
new agent configuration must not persist them.
Do not rewrite an active execution when configuration changes. A fresh execution carries the new policy and replaces an incompatible idle or checkpointed session through normal session-compatibility handling.
Conformance checklist
A new driver must cover:
- agent create/edit defaults, selected-provider visibility, provider switching, invalid-value rejection, and no Runner sandbox-bypass control;
- native serialization, legacy replay, recovery identity, and replacement after a permission change;
- start and resume propagation for every supported mode;
- no runtime request in the maximum mode;
- once, session, and deny resolution in prompting modes, including every provider event version the driver accepts;
- immediate rejection in deny mode;
- workspace and protected-path denial before auto-approval;
- authenticated semantic-tool behavior; and
- an end-to-end default-mode write plus validation command that completes without an Input needed permission card.
Run TypeScript, Rust, UI, protocol-generation, durable transport, and documentation validation before declaring the harness qualified.