Files
PaperClipAI/packages/paperclip-runner/docs/capability-scenario-explorer.md
Dotta 560e7e48b5 feat(runner): add SDK and developer tooling (#12608)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package already provides the production protocol and
execution spine.
> - Contributors still need stable SDK surfaces, deterministic test
tools, and local inspection tools.
> - Those surfaces share generated contracts and must change as one
package boundary.
> - This pull request adds the package-local SDK, labs, examples, and
drift checks.
> - The benefit is a reviewable developer platform that does not change
application execution selection.

## Linked Issues or Issue Description

**Subsystem affected**

`packages/paperclip-runner` — runner SDK, conformance tools, and
developer tooling.

**Problem or motivation**

The production runner spine is present, but package consumers cannot
build deterministic integrations, inspect sessions, or verify
provider-neutral behavior through supported surfaces.

**Proposed solution**

Add browser, React, standalone, live-session, scenario, conformance, and
evaluation surfaces. Add generated contract inventories and
package-local verification scripts. Keep production application routing
unchanged.

**Alternatives considered**

We considered splitting each generated catalog, SDK surface, and demo
into separate pull requests. Those changes share exports, fixtures, and
drift gates. Splitting them would create intermediate package states
that do not build.

**Roadmap alignment**

No overlapping item appears in `ROADMAP.md`. This work extends the
runner package that is already on `master`.

## What Changed

- Add browser, React, standalone, live-session, and issue-thread SDK
surfaces.
- Add deterministic mock control-plane, scenario, conformance, replay,
and evaluation tools.
- Add bounded Codex, OpenCode, and ACPX development transports and
fixtures.
- Keep deferred managed-provider execution fail-closed. Persisted
compatibility data remains readable.
- Add generated capability inventories with their source files and drift
checks.
- Add examples, package documentation, browser checks, and
clean-consumer checks.
- Preserve the reviewed protocol bounds, replay compatibility aliases,
process environment isolation, and semantic redaction limits.
- Update the ACPX package patch that the existing workspace patch
registry already tracks.
- Do not change `pnpm-lock.yaml`, repository workflows, server runtime
selection, or the application UI.

## Verification

GitHub Actions is the verification authority for this pull request. The
repository CI, package TypeScript and Rust checks, package tests,
generated-output drift checks, browser checks, security scans, and
Greptile review must pass on the exact head.

Local test suites were not run because this series uses parallel GitHub
Actions for verification.

## Risks

This is a large greenfield package change. The main risks are public
export drift, generated-output drift, and optional React consumer
compatibility. Package boundary checks, clean-consumer checks, and
browser tests cover those risks. Production adapter selection and server
execution are outside this pull request.

## Stack

1. **This PR:** runner SDK and developer tooling.
2. [Codex production server
integration](https://github.com/paperclipai/paperclip/pull/12616).
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617).

## Model Used

OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the feature request
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 21:33:11 -05:00

3.6 KiB

Capability Browser Scenario Explorer

The scenario explorer is a read-only browser surface over all 106 conformance cases. It runs each scenario in the page against the mock control plane through the semantic tool runtime and renders the resulting run artifact. It never re-judges parity, never leaves its own origin, and holds no credential.

Sources: src/scenarios/ (scenario index, runner, parity, state diff) and examples/scenario-explorer/. Interaction contract: Capability scenario explorer UX.

The run artifact

Running a scenario produces one immutable artifact carrying:

  • Tool exposure — which semantic tools were available, and for optional tools, the grant that unlocked each one.
  • Control-plane actions — control-plane-owned steps (for example checkout), each labelled "no agent tool exists for this."
  • Authorization records — allow and typed-deny decisions; a denial carries the missing claim and no protected task data.
  • State diff — an immutable diff over the ten entity domains, with unchanged domains collapsed.
  • Parity verdict — the runtime's own verdict, plus Capability's per-case result carried through in a separately labelled block. The explorer displays these; it does not recompute them.

Read-only and frozen SDK

The explorer imports the frozen 0.1.2 public SDK through the package-local alias @paperclip-runner-local/capability, which is deliberately not the published package name. It adds no SDK export and changes no published surface. Fake mode runs entirely in the page against checked-in fixtures and renders fixture time only, so two loads of the same route produce identical settled DOM.

What the explorer shows

  • Home — 16 group facets whose counts sum to 106, corpus stats, and example deep links.
  • Picker — a listbox with live filter chips, disabled zero-count values, and a clear-filters control.
  • Transcript — both channels (agent-visible and control-plane).
  • Inspector — four tabs: context (exposure and grants), authorization (allow/deny), state diff, and traceability.

Boundary

  • No network request leaves the explorer's own origin.
  • localStorage holds no run artifact, grant, or fixture payload.
  • Codex mode is a disabled option with a stated reason; wiring it to the SDK relay is deferred, and the browser holds no provider credential either way.
  • Secrets are redacted before display; redaction chips name the rule and never the value. See authorization and exposure.

Accessibility

The 7F browser suite asserts: listbox arrow-key navigation with aria-activedescendant, WAI-ARIA tab arrow-key activation with exactly one tablist, every interactive control named, a polite live region announcing the settled verdict, three landmarks, a single h1, and — at 390px — one segment at a time with zero horizontal overflow.

Running it

# Open the explorer on 127.0.0.1:4183.
pnpm --filter @paperclipai/paperclip-runner demo:scenarios

# Scenario runtime, explorer components, and route determinism (49 tests).
pnpm --filter @paperclipai/paperclip-runner test:scenarios

# Browser IA, determinism, evidence routes, boundary, a11y, responsive (25).
pnpm --filter @paperclipai/paperclip-runner test:browser:scenarios

# Deterministic 24-image acceptance set (12 routes x 2 viewports).
# Recorded evidence generation is deferred from this release.