Files
PaperClipAI/packages/paperclip-runner/docs/tutorials/scenario-chat.md
T
Dotta 560e7e48b5 feat(runner): add SDK and developer tooling (#12608)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package already provides the production protocol and
execution spine.
> - Contributors still need stable SDK surfaces, deterministic test
tools, and local inspection tools.
> - Those surfaces share generated contracts and must change as one
package boundary.
> - This pull request adds the package-local SDK, labs, examples, and
drift checks.
> - The benefit is a reviewable developer platform that does not change
application execution selection.

## Linked Issues or Issue Description

**Subsystem affected**

`packages/paperclip-runner` — runner SDK, conformance tools, and
developer tooling.

**Problem or motivation**

The production runner spine is present, but package consumers cannot
build deterministic integrations, inspect sessions, or verify
provider-neutral behavior through supported surfaces.

**Proposed solution**

Add browser, React, standalone, live-session, scenario, conformance, and
evaluation surfaces. Add generated contract inventories and
package-local verification scripts. Keep production application routing
unchanged.

**Alternatives considered**

We considered splitting each generated catalog, SDK surface, and demo
into separate pull requests. Those changes share exports, fixtures, and
drift gates. Splitting them would create intermediate package states
that do not build.

**Roadmap alignment**

No overlapping item appears in `ROADMAP.md`. This work extends the
runner package that is already on `master`.

## What Changed

- Add browser, React, standalone, live-session, and issue-thread SDK
surfaces.
- Add deterministic mock control-plane, scenario, conformance, replay,
and evaluation tools.
- Add bounded Codex, OpenCode, and ACPX development transports and
fixtures.
- Keep deferred managed-provider execution fail-closed. Persisted
compatibility data remains readable.
- Add generated capability inventories with their source files and drift
checks.
- Add examples, package documentation, browser checks, and
clean-consumer checks.
- Preserve the reviewed protocol bounds, replay compatibility aliases,
process environment isolation, and semantic redaction limits.
- Update the ACPX package patch that the existing workspace patch
registry already tracks.
- Do not change `pnpm-lock.yaml`, repository workflows, server runtime
selection, or the application UI.

## Verification

GitHub Actions is the verification authority for this pull request. The
repository CI, package TypeScript and Rust checks, package tests,
generated-output drift checks, browser checks, security scans, and
Greptile review must pass on the exact head.

Local test suites were not run because this series uses parallel GitHub
Actions for verification.

## Risks

This is a large greenfield package change. The main risks are public
export drift, generated-output drift, and optional React consumer
compatibility. Package boundary checks, clean-consumer checks, and
browser tests cover those risks. Production adapter selection and server
execution are outside this pull request.

## Stack

1. **This PR:** runner SDK and developer tooling.
2. [Codex production server
integration](https://github.com/paperclipai/paperclip/pull/12616).
3. [Provider-neutral task-thread
UI](https://github.com/paperclipai/paperclip/pull/12617).

## Model Used

OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the feature request
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-31 21:33:11 -05:00

130 lines
5.3 KiB
Markdown

# Scenario chat Tutorial: Chat With the Mock Control Plane
**Time to first success: about 3 minutes.** One command opens a chat where you
send prompts to a Capability scenario and watch exactly how the mock Paperclip
control plane was used on every turn. The tutorial runs from the repository
root. It starts no Paperclip service, contacts no Paperclip control plane, and
holds no provider credential.
Scenario chat extends the Capability browser surface; it does not integrate the runner
into Paperclip. Real integration is future upload integration and requires separate approval. See
[the future binding boundary reference](../capability-future-binding-boundary.md).
## What you need
- Node.js 20 or newer and pnpm 9 or newer. Verified with Node 22.22.2.
- No Rust toolchain and no network access after `pnpm install`.
Install the package workspace from the repository root:
```sh
pnpm install --filter @paperclipai/paperclip-runner --lockfile=false --offline --ignore-scripts --dev
```
## 1. Open the chat (about 1 minute)
```sh
pnpm --filter @paperclipai/paperclip-runner run demo:scenarios
```
Open <http://127.0.0.1:4183/scenario-explorer/#/chat/ap-mcp-gate-01>.
Before you type anything, the conversation already has content: a **session
seed** strip reading "Control plane checked out the task and opened the run — no
agent tool exists for this," and a **Session seed activity** strip counting what
it did. That is the point of turn 0. Work happened, and no agent tool was
involved.
The session status chip reads **Idle · seeded**. The right rail — **Activity**
on desktop, the **Activity** segment on a phone — shows the seed group with its
exposure decision, state diff, and verdict.
## 2. Run a turn that succeeds
Type into the composer and press Enter:
```
Where does the MCP tool approval for this task stand?
```
The turn appends: your message as a right-aligned bubble, two semantic calls
(`get_task_context`, `list_approvals`) as tool cards with their disposition
badges, the agent's reply, and a **Turn 1 activity** strip:
```
Turn 1 activity ⚙ 2 tools ⌘ 0 control plane Δ 0 domains ✓ Pass
```
Click the strip. The Activity stream opens turn 1's group with the six evidence
sections in order — exposed tools, calls and results, authorization, control
plane, state diff, parity. Nothing changed state, and the diff says so.
## 3. Run a turn that is denied
```
Raise the approval request yourself and get the tool switched on.
```
`request_approval` is denied. The denial renders in place on the danger surface
with the failing operation, the code `policy_denied`, the reason
`required_claim_missing`, and an explicit line that **no protected state was
returned**. The turn strip counts it (`✕ 1 denied`), and on mobile the Activity
segment carries a `· 1 deny` chip so you can see it without opening the segment.
The session status still reads **Scripted · deterministic** and the session
parity still passes. A denial is a correct outcome, not a broken session — this
run simply does not hold `governance:approvals:request`. Open the turn's
**Authorization** section to see the claims considered and the missing one.
## 4. See control-plane-owned work with no agent tool
Open <http://127.0.0.1:4183/scenario-explorer/#/chat/ix-checkbox-01?replay=fake>.
`replay=fake` plays the recorded three-turn conversation.
Turn 2 opens a checkbox interaction. Turn 3 is the board answering it — and
that answer is performed by the control plane, not by a tool:
```
⌘ Control plane recorded the board's answer and scheduled the continuation
wake — no agent tool exists for this.
```
Open turn 3's **State diff** section. The `interactions` domain shows
`status: pending → answered`, the unchanged domains collapse to one line, and a
**Control plane** badged row records `wake.scheduled`. The wake is scheduled by
reconciliation; the agent never asked for it.
## 5. Reset and replay
**Reset session** asks first, because it discards mutable mock state. Accept it
and the transcript returns to the seeded turn 0 — no turns, no denials, a clean
diff. Switching scenario from the rail mid-session asks the same question, since
two scenarios never share mock state.
**Replay scripted session** plays the remaining recorded turns and needs no
confirmation: it reproduces an identical timeline.
## 6. Run the tests and record the evidence
```sh
pnpm --filter @paperclipai/paperclip-runner run test:scenarios
pnpm --filter @paperclipai/paperclip-runner run test:browser:scenarios
# Recorded evidence generation is deferred from this release.
```
The first covers the session data contract and the chat components, the second
drives the surface in a real browser at both viewports, and the third writes the
ten acceptance screenshots to `.paperclip-local/evidence/capability/` — failing if any
route scrolls horizontally.
## What this does not do
- No ACPX implementation and no real control-plane integration.
- No provider credential in the browser. **Codex (bounded)** stays disabled with
a named reason unless the local provider relay is running; **Scripted
(deterministic)** drives the same mock core offline.
- Nothing you type is persisted. A reload drops an unsent draft, which is the
accepted trade against storing conversation content.
Reference: [Scenario chat scenario chat](../scenario-chat.md).