## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
4.6 KiB
Capability Authorization and Exposure
Two decisions gate every capability: exposure (is the tool even offered to this actor?) and invocation (may this specific call proceed?). The authorization engine answers both, produces typed denials that carry no protected state, and redacts secrets before any observable boundary.
Sources: src/tools/capability-tool-authorization.ts,
src/tools/capability-semantic-tool-runtime.ts, src/tools/capability-tool-bindings.ts.
Three-way classification
- Control-plane-owned — a frozen operation list (
checkout_task,release_task,select_work,route_wake,enforce_budget,append_audit_record,persist_run,replay_run,schedule_blocker_wake,reconcile_run).authorizeInvocationshort-circuits these to outcomeabsentwith reasoncontrol_plane_owned_operation. No agent tool exists for them and none can be invoked. - Always-agent tool — exposed to any actor whose task is active.
- Optional-agent tool — exposed only when a grant unlocks it.
Exposure vs. invocation
computeVisibleTools()walks the whole catalog and records, per tool, an exposure ofexposedorabsent, returning only the allowed tools.authorizeInvocation()re-evaluates at call time, producingallowedordenied.
Invocation is evaluated in order and fails at the first rule that rejects:
policy-denied operation IDs; task-mode membership
(operation_not_available_in_task_mode); role check
(actor_role_not_authorized); policy-denied claims
(required_claim_denied_by_policy); missing grants
(required_claim_missing); read_secret_value requires
policy.allowSecretValueAccess (secret_value_access_disabled);
generic_api_request requires the escape hatch to be enabled and the target to
be allowlisted (generic_escape_hatch_disabled / ..._target_denied); and the
interaction-kind policy for request_human_input. Allowed calls carry the
reason task_scoped_always_tool or required_claims_granted.
Self-approval is prohibited even with the role and an explicit decision claim.
How grants unlock optional tools
The actor's grants are the union of capabilityGrants (the actor's own) and
scenarioGrants (seeded per scenario). An optional tool unlocks only when
every entry in its requiredClaims is present in that union. Each
authorization record stores the consideredClaims, grantedClaims, and
missingClaims, so a denial explains exactly which claim was absent.
Typed denials carry no protected state
A denial is a CapabilityPolicyDenial: ok: false with an error.code of
policy_denied, operation_absent, input_invalid, or operation_unsupported,
the operationId, a generic reason, and the authorization record. It carries
no fixture state and no protected payload. If the underlying mock operation
throws, the runtime emits operation_unsupported with reason
mock_operation_rejected and swallows the underlying error, so protected state
cannot leak through an exception.
Secret redaction (the TASK-16909 rule)
Redaction is two layers:
- Model-only delivery is separated from every observable boundary.
invoke()returns a redactedobservableResult. The only path to an unredacted secret isinvokeForModel(), which returns a frozen capsule whosetoJSON()and default serialization always yield the redacted form; only an explicitreadModelResult()opens the raw value. The capsule uses a distinct schema (paperclip.capability.model-tool-result.v1) so it cannot be assigned to an observable sink by mistake. - Path redaction rules.
read_secret_valueredacts$.valueto[SECRET_VALUE]across output, error, and authorization-record channels;generic_api_requestredacts$.headers.authorizationand$.headers.cookieto[REDACTED]across all channels. The scenario runner applies the same redaction again defensively at the explorer artifact boundary and injects only a placeholder secret value.
This is the security-gate outcome (TASK-16902) and its remediation (TASK-16909): a real secret value reaches the model surface only, never a trace, artifact, or browser view.
Running the tests
pnpm --filter @paperclipai/paperclip-runner exec vitest run \
src/tools/capability-semantic-tools.test.ts
The "exposure and authorization" and "security policy" describes cover typed denials without protected state, model-only secret delivery, the self-approval prohibition, and the test-only, explicitly granted, allowlisted escape hatch.