mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner package already provides the production protocol and execution spine. > - Contributors still need stable SDK surfaces, deterministic test tools, and local inspection tools. > - Those surfaces share generated contracts and must change as one package boundary. > - This pull request adds the package-local SDK, labs, examples, and drift checks. > - The benefit is a reviewable developer platform that does not change application execution selection. ## Linked Issues or Issue Description **Subsystem affected** `packages/paperclip-runner` — runner SDK, conformance tools, and developer tooling. **Problem or motivation** The production runner spine is present, but package consumers cannot build deterministic integrations, inspect sessions, or verify provider-neutral behavior through supported surfaces. **Proposed solution** Add browser, React, standalone, live-session, scenario, conformance, and evaluation surfaces. Add generated contract inventories and package-local verification scripts. Keep production application routing unchanged. **Alternatives considered** We considered splitting each generated catalog, SDK surface, and demo into separate pull requests. Those changes share exports, fixtures, and drift gates. Splitting them would create intermediate package states that do not build. **Roadmap alignment** No overlapping item appears in `ROADMAP.md`. This work extends the runner package that is already on `master`. ## What Changed - Add browser, React, standalone, live-session, and issue-thread SDK surfaces. - Add deterministic mock control-plane, scenario, conformance, replay, and evaluation tools. - Add bounded Codex, OpenCode, and ACPX development transports and fixtures. - Keep deferred managed-provider execution fail-closed. Persisted compatibility data remains readable. - Add generated capability inventories with their source files and drift checks. - Add examples, package documentation, browser checks, and clean-consumer checks. - Preserve the reviewed protocol bounds, replay compatibility aliases, process environment isolation, and semantic redaction limits. - Update the ACPX package patch that the existing workspace patch registry already tracks. - Do not change `pnpm-lock.yaml`, repository workflows, server runtime selection, or the application UI. ## Verification GitHub Actions is the verification authority for this pull request. The repository CI, package TypeScript and Rust checks, package tests, generated-output drift checks, browser checks, security scans, and Greptile review must pass on the exact head. Local test suites were not run because this series uses parallel GitHub Actions for verification. ## Risks This is a large greenfield package change. The main risks are public export drift, generated-output drift, and optional React consumer compatibility. Package boundary checks, clean-consumer checks, and browser tests cover those risks. Production adapter selection and server execution are outside this pull request. ## Stack 1. **This PR:** runner SDK and developer tooling. 2. [Codex production server integration](https://github.com/paperclipai/paperclip/pull/12616). 3. [Provider-neutral task-thread UI](https://github.com/paperclipai/paperclip/pull/12617). ## Model Used OpenAI Codex, GPT-5, high-reasoning mode, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the feature request template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
223 lines
10 KiB
Markdown
223 lines
10 KiB
Markdown
# Capability Clean-Start Tutorial: The Paperclip-Style Issue Thread Over a Mock Control Plane
|
||
|
||
**Time to first success: about 5 minutes.** Two commands take you from a clean
|
||
checkout to 106 passing conformance cases and a Paperclip-style issue thread you
|
||
can click through in a browser. The whole tutorial runs from the repository
|
||
root. It starts no Paperclip service, contacts no Paperclip control plane, clones
|
||
no external eval repository, and holds no provider credential. Everything the
|
||
default path needs is checked in under `packages/paperclip-runner/`.
|
||
|
||
This is the canonical Capability tutorial for the final build. It covers both agent
|
||
modes: **scripted (deterministic) fake mode**, which is fully offline, and
|
||
**Codex (bounded) live mode**, which is optional and needs a locally
|
||
authenticated Codex. The two modes share one mock control plane, one semantic
|
||
tool catalog, one authorization engine, and one view contract. See
|
||
[execution modes and identity](../capability-execution-modes.md) for the rules that
|
||
separate them, and [the future binding boundary](../capability-future-binding-boundary.md)
|
||
for why none of this touches real Paperclip — that is future upload integration and needs separate
|
||
approval.
|
||
|
||
For a deeper tour of the read-only scenario explorer and the capability
|
||
contract, see the companion
|
||
[scenario-explorer tutorial](capability-scenario-explorer.md); this page focuses
|
||
on the issue-thread surface and the live loop.
|
||
|
||
## What you need
|
||
|
||
- Node.js 20 or newer and pnpm 9 or newer. This tutorial was verified with Node
|
||
22.22.2 and pnpm 9.15.4.
|
||
- No Rust toolchain for the fake-mode steps (steps 1–4). The optional live steps
|
||
(step 5) build the `paperclip-runnerd` binary and need a Rust toolchain and a
|
||
locally authenticated Codex.
|
||
- No network access after `pnpm install`.
|
||
|
||
Install the package workspace from the repository root:
|
||
|
||
```sh
|
||
NODE_ENV=development pnpm install --filter @paperclipai/paperclip-runner --frozen-lockfile --offline --ignore-scripts
|
||
```
|
||
|
||
This keeps the checked-in lockfile authoritative while using only packages
|
||
already present in the pnpm store, so the offline install does not re-resolve
|
||
dependencies to registry-latest versions. Setting `NODE_ENV=development`
|
||
ensures the complete test toolchain is installed even when the checkout
|
||
inherits `NODE_ENV=production`.
|
||
|
||
## 1. Prove the 106-case conformance suite (about 1 minute)
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner test:capability-evals
|
||
```
|
||
|
||
This drives all 106 eval-derived cases across the 16 groups entirely in-process
|
||
against the mock control plane. Expected final line:
|
||
|
||
```
|
||
Test Files 1 passed (1)
|
||
Tests 1 passed (1)
|
||
```
|
||
|
||
For the per-group counts, assertion classes, the 18-operation fake-agent matrix,
|
||
and the 16-case bounded Codex binding sample, generate the parity report:
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner report:capability-evals
|
||
```
|
||
|
||
Expected final line:
|
||
|
||
```
|
||
Capability eval conformance passed: 106 cases across 16 groups.
|
||
```
|
||
|
||
The report is written to `.paperclip-local/evidence/capability/eval-parity-report.{json,md}`.
|
||
It is generated on demand and is not committed; delete it before running
|
||
`docs:validate`, or run validation first.
|
||
|
||
## 2. Run the issue-thread contract and UI tests (about 1 minute)
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner test:scenarios
|
||
```
|
||
|
||
Expected final lines:
|
||
|
||
```
|
||
Test Files 11 passed (11)
|
||
Tests 137 passed (137)
|
||
```
|
||
|
||
This covers the mock `ControlPlanePort`, the semantic tools with their
|
||
authorization and redaction rules, and the `CapabilityIssueThreadSnapshot`
|
||
view-model contract — including that every one of the twelve deterministic
|
||
screenshot states builds a schema-valid snapshot with no credential anywhere in
|
||
the projected view.
|
||
|
||
## 3. Open the issue thread in fake mode (about 2 minutes)
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner console:issue-thread
|
||
```
|
||
|
||
Open <http://127.0.0.1:4184/#/issue/hb-baseline?shot=thread-baseline>.
|
||
|
||
You are looking at a Paperclip issue thread, not a dashboard. The header carries
|
||
three identity chips that stay visible in every mode — `Fake agent`,
|
||
`In-process runner`, and `Mock Paperclip` — plus the mock issue identifier
|
||
(reserved `MCK-` prefix), status, priority, and the Scenario / Replay / Reset /
|
||
Stop controls. The main column is one readable work thread; the Evidence panel is
|
||
collapsed on the right.
|
||
|
||
Swap the `shot` value to seed any of the twelve contract states — each one is a
|
||
different item type or interaction state from the
|
||
[issue-thread UX contract](../design/capability-issue-thread-ux-contract.md):
|
||
|
||
| `shot=` | What it shows |
|
||
| --- | --- |
|
||
| `thread-baseline` | User message, model prose, a durable `Recorded to mock thread` comment, and a collapsed tool strip. |
|
||
| `turn-streaming` | A streaming turn with the composer in its `streaming` state (input editable, Stop primary). |
|
||
| `interaction-question-pending` | An inline `ask_user_questions` card that refuses an incomplete submit. |
|
||
| `interaction-confirmation-pending` | A revision-bound confirmation that requires a reject reason. |
|
||
| `interaction-resolved-mixed` | Accepted, rejected, expired, and stale interactions as durable history. |
|
||
| `denial-optional-tool` | A verbatim typed denial with its Evidence badge and no protected data. |
|
||
| `document-revision` | A document revision card bound to the mock document. |
|
||
| `deliverable-registered` | A registered deliverable reporting one size in one unit. |
|
||
| `disposition-terminal` | A terminal disposition that disables the composer. |
|
||
| `debug-panel-open` | The eight-section Evidence panel opened on the authorization records. |
|
||
| `reconnect-banner` | The reconnect banner and its `Reconnected` confirmation. |
|
||
| `replay-mode` | The replay progress strip with `Step back`, `Next turn`, and `Play all`. |
|
||
|
||
Open the Evidence panel (or the `Evidence` segment on a phone). Its eight
|
||
sections stay in a fixed order: Tools exposed, Calls & results, Authorization,
|
||
Control plane, Runner & events, State diff, Traceability, Parity. The Tools
|
||
section separates `Agent tool — always`, `Agent tool — granted` (with its
|
||
grant), and a `Control plane (not exposed to the agent)` list — what the model
|
||
*cannot* call is first-class evidence. Every strip, denial, and card deep-links
|
||
into its record, and each record links back to its thread anchor.
|
||
|
||
Resize the browser below 768px: the page collapses to a `Thread` / `Evidence`
|
||
segmented control, keeps Stop outside the `⋯` menu while a turn is active, and
|
||
never scrolls horizontally.
|
||
|
||
## 4. Record the deterministic screenshot matrix (about 1 minute)
|
||
|
||
```sh
|
||
# Recorded evidence generation is deferred from this release.
|
||
pnpm --filter @paperclipai/paperclip-runner check:capability:ui
|
||
```
|
||
|
||
The first writes 24 images — the twelve slugs at 1440×900 and 390×844 — to
|
||
`.paperclip-local/evidence/capability/ui/`. Every mobile capture asserts no horizontal
|
||
scroll before the shot. The second re-records and confirms all 24 reproduce
|
||
byte-for-byte, which is the property that makes the fake-mode matrix a stable
|
||
acceptance artifact. Both modes render fixture time only, so a fresh checkout
|
||
reproduces the committed images exactly.
|
||
|
||
## 5. Optional: drive a real Codex turn (live mode)
|
||
|
||
Live mode needs a Rust toolchain and a locally authenticated Codex. It starts a
|
||
real `paperclip-runnerd` process and a real Codex app-server session; the
|
||
package server owns the process, the tool loop, and the provider credential, and
|
||
the browser receives none of them. **Skip this section if you do not have Codex
|
||
installed** — nothing above depends on it.
|
||
|
||
Headless smoke over the package server:
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner smoke:capability:ui
|
||
```
|
||
|
||
This creates a session, runs live turns, and asserts the identity reads
|
||
`Real Codex` / `Real runnerd` / `Mock Paperclip`, that a real tool call was
|
||
recorded with its authorization record, that the control-plane-owned list is
|
||
withheld from the agent, and that no credential appears in the view.
|
||
|
||
Capture the live browser frames (a real browser driving a real Codex turn):
|
||
|
||
```sh
|
||
# Recorded evidence generation is deferred from this release.
|
||
```
|
||
|
||
The frames land in `.paperclip-local/evidence/capability/ui-live/`. They are intentionally
|
||
**not** byte-stable — a real provider turn varies — so they are excluded from the
|
||
determinism gate and are the mode that satisfies the final live acceptance
|
||
criteria. A scripted (`mode=fake`) frame cannot.
|
||
|
||
Prove the process and network boundary directly:
|
||
|
||
```sh
|
||
pnpm --filter @paperclipai/paperclip-runner trace:live-runner -- --json
|
||
```
|
||
|
||
This checks a real semantic-tool mutation, a typed result, a same-thread second
|
||
turn, process ownership and cleanup (`runnerExited`, `authorityCleared`), and —
|
||
critically — `noRealPaperclipRequest` and `noPaperclipAuthorityInChild`.
|
||
|
||
## Reset, stop, replay, and cleanup
|
||
|
||
These controls work the same in either mode because they act on the mock core
|
||
and the runner session, not on the agent:
|
||
|
||
- **Reset** asks first, then restores the original clean mock seed under new
|
||
run/session authority and rotates or clears the old authority. It does not
|
||
touch another browser session.
|
||
- **Stop** cancels only the active turn and reaps the runnerd process group; the
|
||
session and transcript survive with a `Stopped by user` marker.
|
||
- **Replay** reproduces a recorded canonical timeline. It is always scripted and
|
||
always labelled fake-derived.
|
||
- **Refresh / reconnect** restores the same session, pending interaction,
|
||
transcript, and mock state; reconnect starts a fresh runnerd and Codex and
|
||
resumes the persisted thread without restarting the conversation.
|
||
|
||
## What this does not do
|
||
|
||
- No ACPX implementation and no real control-plane integration. Real Paperclip
|
||
binding is future upload integration and requires separate approval.
|
||
- No provider, runner, or control-plane credential in the browser, ever.
|
||
- Nothing you type is persisted beyond the mock session. A reload drops an
|
||
unsent draft — the accepted trade against storing conversation content.
|
||
|
||
References: [issue-thread UI](../capability-issue-thread-ui.md),
|
||
[live runnerd and Codex loop](../capability-live-runnerd-codex.md),
|
||
[execution modes and identity](../capability-execution-modes.md).
|