## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
Paperclip Native Runner
This package is the standalone development boundary for Paperclip's native
runner protocol, process supervision, durable transport, provider drivers, and
normalized session backends. Rust owns the production runner under runner/;
TypeScript provides the control-plane reference, browser SDK, scenario tools,
and conformance oracle.
The package includes one coherent set of capabilities: PRP v1 validation and replay, a supervised local runner with a scripted fake harness, durable WebSocket delivery and recovery, qualified Codex, OpenCode, ACPX, Claude Managed, and AWS AgentCore drivers, live session and issue-thread surfaces, a public browser/React SDK, a standalone adapter demo, and a deterministic mock control plane. None of these surfaces imports or starts Paperclip's server, UI, CLI, or production database.
Public package surfaces
@paperclipai/paperclip-runner— production contracts, clients/backends, PRP validation/replay, canonical catalog/dispatcher, and compatibility check.@paperclipai/paperclip-runner/testing— deterministic mocks plus PRP and semantic conformance kits. Tests and external conformance consumers import this explicitly.@paperclipai/paperclip-runner/evals— versioned native-attempt metadata, fail-closed package/binary compatibility checks, and explicit runnerd artifact resolution for eval consumers.
The package root has no mock or scenario exports. Generic credential-free
matrix orchestration lives in the workspace-private
@paperclipai/paperclip-eval-kernel; scenario content and provider-backed eval
campaigns remain outside the runtime package.
See ADR 0001.
The two conformance surfaces intentionally prove different contracts. The
existing runControlPlanePortConformance suite checks narrow PRP run/event
persistence. CAPABILITY_HIGH_RISK_SEMANTIC_VECTORS and
runSemanticConformanceKit compare normalized tool authorization, state,
effects, audit, retries, conflicts, redaction, continuation, and terminal
decisions. The production adapter stays App-owned and invokes Paperclip's real
route/service authorities; it does not copy those rules into this package.
Quick start
Native provider debug-trace correlation uses an incremental index owned by its transport. Pending event lookups read only newly appended bytes, with a 1 MiB read budget per lookup; they retry until the observed suffix is indexed. Partial records remain pending, and trace replacement or truncation invalidates the index. Closing a transport clears its index. Other active transports cannot evict its progress. Records over 64 KiB are skipped by the correlation index without buffering or parsing their full contents; the original trace file retains them. Do not restore a full synchronous trace scan for each pending event: it blocks event delivery and can leave the board showing an active run after the provider turn has already ended.
The package also builds paperclip-runner-acpx-sidecar. This bounded v2
stdin/stdout bridge admits the pinned Claude and Codex ACPX profiles. It
validates the exact model, session identity, tool catalog, structured input,
and terminal settlement at the process boundary. Pi remains unavailable.
Native Claude skill assignments travel in the runtime-context snapshot through
runnerd to the ACPX sidecar. After acquiring the provider lifetime lease, the
host materializes the assigned bundles under the isolated Claude home's
skills/ directory before launch. Reopening a provider refreshes that snapshot;
project and ambient host settings remain excluded. This path is separate from
the legacy claude_local adapter's remote skill staging.
The isolated Claude settings pin both model and availableModels to the
user's requested ID. This keeps ACP from replacing an exact ID with a picker
alias during selection and verification. Users can keep selecting models from
the normal Claude catalog or entering custom IDs; unavailable models still fail
at the provider rather than silently falling back.
For ACPX Claude, approve-reads is shown as Allow Paperclip reads. The host
intersects the run's public tools with the implementation catalog's read effects
and writes exact MCP permission rules into the isolated Claude settings. The
paperclip connection is always the runner's authenticated tool bridge; ambient
MCP configuration is excluded. Tool hints and provider permission metadata cannot
grant access. Unassigned tools, writes, external tools, and provider-native
operations do not receive automatic read permission. Protocol completion and
task-delivery controls keep their existing separate allowance.
This runtime has no interactive permission handler. An operation that still
requires approval stops the turn with approval_required. The server marks the
task blocked, exposes the permission action to the operator, and disables
automatic retry. The operator must review the operation and the agent's
permission setting before retrying. Company access checks still run when each
Paperclip tool executes.
Runnerd selects only qualified provider profiles. Claude Managed and AWS AgentCore receive immutable company-profile snapshots with explicit retention, spend, and invocation limits. No provider process receives a Paperclip API credential or unrestricted server environment.
Claude Managed resolves its API key from the company secret bound to the selected profile. AWS AgentCore uses workload identity only; long-lived static AWS access keys are intentionally removed from the runner environment.
The Rust core includes a bounded client for the sidecar protocol. It enforces request identity, event order, frame and queue limits, timeouts, redacted diagnostics, and process-group cleanup. Runnerd selects this package-local transport only through an exact qualified provider descriptor.
Before a later provider adapter consumes a valid sidecar event, the Rust core also requires its optional or mandatory run and turn scope to match the active execution. Process and diagnostic events can remain global. All operational, tool, input, permission, and terminal events require the exact active binding.
A package-local payload boundary decodes events only after that scope check. It validates control identities, terminal status, question sets, and the admitted runtime event types and bounded fields. It redacts diagnostic and retained event values again before they can enter provider state.
Validated ACPX runtime events normalize into the same provider-neutral activity families as the direct Codex transport. Reasoning contents stay private. Tool targets are resolved within the workspace under the provider host's path semantics and receive a versioned sidecar boundary marker before becoming bounded, display-only PRP safe paths. Raw or unmarked provider locations fail closed. URI-scheme and Windows drive-shaped values require a separate sidecar attestation backed by an existing in-workspace entry or, for a not-yet-created edit target, an existing in-workspace parent. This preserves real POSIX colon filenames without treating arbitrary URI text as a path. Windows separators are canonicalized, and consumers must not reinterpret the display value as file-access authority. Operational semantic-result and terminal events remain reserved for the stateful adapter rather than being duplicated.
The ACPX provider reducer preserves that order while it tracks one active turn, bounded assistant text, semantic results, and pending tool or input correlations. Terminal events flush the final assistant message first and clear unresolved turn-scoped requests.
The package-local session bootstrap starts the bounded sidecar transport, verifies the qualified capability handshake and effective model, opens one identity-bound session, and confirms its run attachment. Any failed bootstrap terminates the process; session shutdown preserves persistent provider state. The session can then start one immutable-workspace turn, request interruption, and reduce polled events through the scope-first state boundary. A mismatched command acknowledgement or invalid event terminates the session fail closed. Polled semantic calls pass through the run-scoped authorized tool bridge before they can be returned to a caller. Before a follow-up turn releases settled tool receipts, runner-core suspends and reaps the idle sidecar/provider generation, then resumes the same verified persistent identity in a fresh generation. This prevents a late session-lifetime MCP callback from inheriting the next turn's event authority.
The Rust question-response validator checks the versioned response envelope against the exact persisted question IDs, answer modes, options, required answers, custom-answer policy, and text constraints before provider delivery. Tool results and structured question responses then use two-phase resolution: validate retained identity and schema, require the exact sidecar acknowledgement, and only then clear pending local state. Codex permission requests violate its pinned sidecar policy and terminate the session fail closed. Safe suspension is available only with no active turn or pending request. The sidecar must return the exact persistent session identity before runnerd terminates the local process. Already validated ACPX reducer events project into provider-neutral durable events only with an exact run, session, turn, and item binding. Raw sidecar envelopes and permission requests are not admitted at this boundary. A safely suspended session can be recorded as a bounded private checkpoint. The checkpoint binds the exact provider identity, run, catalog revision, and catalog digest and is replaced atomically before a later recovery attempt. Recovery releases the stored identity only after those bindings match the prospective session configuration exactly.
Run the complete contract gate with:
pnpm install --filter @paperclipai/paperclip-runner --lockfile=false --offline --ignore-scripts --dev
pnpm --filter @paperclipai/paperclip-runner verify
The verification command requires a stable Rust toolchain with cargo on
PATH, in addition to Node.js 24.11+ and pnpm 9+.
Minimal Debian/Ubuntu hosts without root access can extract the required Playwright browser libraries into a user-owned cache and run the same acceptance sequence with:
pnpm --filter @paperclipai/paperclip-runner verify:rootless
The tracer's final line is stable:
{
"schemaVersion": "paperclip.runner.conformance.output.v1",
"runIdentity": {
"runId": "run_conformance_0001",
"sessionId": "session_conformance_0001"
},
"result": {
"status": "succeeded",
"summary": "Standalone Conformance fixture accepted."
}
}
Run only the tracer with:
pnpm --filter @paperclipai/paperclip-runner trace:conformance
Replay the Replay happy path, run a Local session, or open the browser devtool:
pnpm --filter @paperclipai/paperclip-runner replay:fixture
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario happy-path
pnpm --filter @paperclipai/paperclip-runner trace:codex
pnpm --filter @paperclipai/paperclip-runner demo:live-console -- --host 127.0.0.1 --port 4174
# Live console: chat with a live session in the browser.
pnpm --filter @paperclipai/paperclip-runner console:live-console
pnpm --filter @paperclipai/paperclip-runner browser:dev --host 127.0.0.1 --port 4179
# SDK: open the public-SDK reference console and mini consumer.
pnpm --filter @paperclipai/paperclip-runner console:sdk
# Standalone: run the standalone legacy/native/kill-switch tracer and page.
pnpm --filter @paperclipai/paperclip-runner trace:standalone
pnpm --filter @paperclipai/paperclip-runner trace:standalone -- --feature-flag enabled
pnpm --filter @paperclipai/paperclip-runner trace:standalone -- --feature-flag enabled --kill-switch enabled
pnpm --filter @paperclipai/paperclip-runner demo:standalone
Live console provider-backed routes are loopback-only and reject wildcard/LAN
binds. Browser mutations require same-origin Fetch Metadata, matching Origin,
and JSON content; see the protocol-server tutorial for direct curl examples.
Direct live protocol qualification
The canonical direct live protocol suite lives in the separate
paperclip-evals repository under evals/paperclip-runner/. Its
live-mini.json roster is the complete 35-case Codex qualification lane. Build
this package's TypeScript output, release paperclip-runnerd, package tarball,
and dist-issue-thread viewer, then use the roster runner documented in that
repository. The package ships the required orchestration entry point as
paperclip-runner-eval-session (dist/cli/eval-session.js). Evalbook owns the
consistent HTML matrix and read-only attempt drill-down pages.
The hosted full-campaign workflow, parallel matrix, credential boundaries,
canonical report merge, and versioned S3 index are documented in
docs/runner-protocol-live-evals.md.
This direct protocol qualification is separate from the stress-derived Runner workflow schedule below and from the full-stack browser model E2E suite.
Stress-derived workflow, chaos, and AWS AgentCore operations
The deterministic workflow scorer and the chaos schedule do not require provider credentials:
pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals
pnpm --filter @paperclipai/paperclip-runner report:runner-chaos-evals
report:runner-live-evals is a paid, provider-backed command. Native Codex
requires OPENAI_API_KEY; ACPX Claude requires
ANTHROPIC_API_KEY; OpenCode candidates require OPENROUTER_API_KEY. The live
matrix admits no Pi profile and does not persist credential values. Set
PAPERCLIP_EVAL_MAX_CAMPAIGN_COST_USD to a positive finite number to bound
additional scheduling after the observed campaign total reaches that value:
PAPERCLIP_EVAL_MAX_CAMPAIGN_COST_USD=12 \
PAPERCLIP_EVALS_ROOT=/path/to/paperclip-evals \
pnpm --filter @paperclipai/paperclip-runner report:runner-live-evals
# Run two scheduled native Codex executions only.
PAPERCLIP_EVALS_ROOT=/path/to/paperclip-evals \
pnpm --filter @paperclipai/paperclip-runner report:runner-live-evals -- \
--candidate codex-luna --limit 2
GitHub-hosted live campaigns additionally require the default branch, an
allowlisted numeric actor ID, the protected runner-e2e-paid environment, and
an explicit repository variable before scheduled runs are enabled. Manual
dispatches accept the same candidate, case, and execution-limit selectors. The
paid job uses the reviewed RunsOn Fleet label when RUNNER_E2E_AWS_ENABLED=true
and otherwise stays on ubuntu-latest. Uploaded reports contain redacted
observations and trace digests, not raw provider
frames, prompts, credentials, tool arguments, or hidden reasoning.
The AgentCore proof-of-concept uses an AWS CLI v2 profile to provision a
dedicated invocation role and scoped resources. Its local mode-0600 metadata
file contains no access keys; probes assume short-lived STS credentials and
clear them after use. Validate locally, provision or inspect the stack, run the
bounded lab/smoke, and tear it down explicitly with:
pnpm --filter @paperclipai/paperclip-runner test:aws-agentcore-provisioning
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision -- --dry-run
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:probe
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:lab
pnpm --filter @paperclipai/paperclip-runner smoke:capability:aws-agentcore
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:destroy -- --yes
To admit the hosted direct-eval workflow, provision with the account-local GitHub Actions OIDC provider and keep the default exact repository and protected environment binding:
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision -- \
--aws-profile paperclip-dev \
--github-oidc-provider-arn arn:aws:iam::<account-id>:oidc-provider/token.actions.githubusercontent.com
This adds only repo:paperclipai/paperclip:environment:runner-e2e-paid as a
web-identity subject on the scoped invocation role. The generated nonsecret
profile records that role as both the local invocation role and the hosted
execution role.
Provisioning can incur Bedrock, AgentCore Runtime/Memory, storage, and private
networking charges. Provisioning refuses to modify a colliding stack unless its
Paperclip ownership tags and template description match. A verified
ROLLBACK_COMPLETE stack still requires --replace-failed-stack plus an
interactive confirmation (or --yes) before it can be deleted and recreated.
Destruction requires --yes and refuses to remove a stack with an active
recorded lab unless --force is also supplied.
Package-owned commands
| Command | Purpose |
|---|---|
build |
Compile the TypeScript public surface, Rust workspace, and browser devtool. |
typecheck |
Check TypeScript, Rust, generated schema sources, and browser types. |
test |
Run Rust/TypeScript fixture, supervisor, fake-driver, live/replay, and boundary tests. |
check:forbidden-imports |
Reject TypeScript imports and Cargo path dependencies that cross into Paperclip core. |
check:tracked-imports |
Reject tracked imports and package.json entry points that only resolve against untracked files, so a clean checkout of any commit builds. |
check:numbered-milestones |
Reject numbered construction-milestone names in tracked package paths and source. |
check:package-boundaries |
Enforce the acyclic runtime/testing/eval dependency and manifest boundary. |
check:clean-consumers |
Pack the runner and install its root, evals, and testing exports in a clean consumer. |
test:eval-slice |
Run the credential-free eval bundle, scoring, and behavior/fault slice. |
test:runner-workflow-evals |
Run the deterministic provider-neutral workflow matrix. |
report:runner-workflow-evals |
Validate deterministic fail-closed results and write JSON, Markdown, JUnit, and GitHub-safe reports. |
report:runner-live-evals |
Execute the paid provider schedule and render its immutable attempts with the canonical paperclip-evals HTML grid. |
report:runner-chaos-evals |
Write the credential-free eight-scenario chaos schedule. |
test:aws-agentcore-provisioning |
Validate the AgentCore template and wrapper safety contracts without provisioning. |
aws-agentcore:provision / probe / lab / destroy |
Manage the scoped AgentCore proof-of-concept lifecycle. |
smoke:capability:aws-agentcore |
Exercise the qualified AgentCore profile through the capability harness. |
check:conformance-parity |
Require byte-for-byte equivalent Rust and TypeScript tracer output. |
check:replay-goldens |
Require all reducer snapshots and cross-language summaries to match checked goldens. |
check:replay-parity |
Run TypeScript and Rust against the same Replay fixture summaries. |
check:browser-tokens |
Reject component-local visual literals and require the standalone token layer. |
docs:validate |
Validate local documentation links. |
trace:conformance |
Run the Rust mock-core tracer, print the stable result, and exit. |
trace:conformance:typescript |
Run the TypeScript reference tracer directly. |
replay:fixture |
Validate and reduce a fixture to a final snapshot. |
trace:local-runner |
Run one native local session through the Rust runner and fake harness. |
trace:codex |
Run the mock core with a real, local skillless Codex app-server session. |
demo:live-console |
Start the package-local HTTP/SSE server with server-only Codex authentication. |
console:live-console |
Start the standalone browser devtool with the Live console on 127.0.0.1:4180. |
console:sdk |
Start the public-SDK reference console and mini consumer on 127.0.0.1:4181. |
test:sdk |
Run targeted browser-client, reducer-projection, and React component contract tests. |
test:browser:sdk |
Exercise both consumers with the fake driver, keyboard/a11y checks, reconnect/replay, measurements, and screenshots. |
record:sdk:codex |
Run both public consumers against a safe real Codex session and capture live screenshots. |
check:capability-contract |
Verify the generated capability, legacy MCP, and eval traceability contract. |
check:semantic-contracts |
Verify the provider-neutral semantic tool contract is current. |
trace:live-runner |
Run the real runnerd/Codex semantic loop against the mock control plane. |
demo:scenarios |
Start the Capability scenario explorer over the mock control plane on 127.0.0.1:4183. |
console:issue-thread |
Start the Paperclip-style issue thread on 127.0.0.1:4184. |
test:scenarios |
Run the scenario index, run-artifact, parity, explorer component, and route tests. |
test:browser:scenarios |
Exercise both the scenario explorer and issue-thread browser contracts. |
browser:dev |
Start the standalone live/replay browser devtool. |
test:browser |
Exercise static replay and live scenarios, then capture temporary screenshots under ignored test output. |
verify |
Run the complete deterministic Conformance through SDK acceptance sequence. |
verify:rootless |
Extract Debian/Ubuntu browser libraries without root, then run verify. |
Navigate
- Architecture and dependency boundary
- ADR 0001: runner and testing package boundaries
- Tutorial index
- Conformance hand-run tutorial
- Replay hand-run tutorial
- Local runner hand-run tutorial
- Local protocol reference
- Durable transport reference
- Codex skillless Codex tutorial
- Codex skillless Codex driver reference
- Live console protocol/server tutorial
- Live console protocol/server reference
- Live console tutorial
- Live console reference
- SDK console tutorial
- SDK browser SDK reference
- SDK component decision record
- Capability semantic catalog and authorization
- Capability live runnerd/Codex loop
- Capability issue-thread UI
- PRP compatibility/versioning policy
- Adding a harness and permission-mode requirements
- PRP v1 expressiveness audit
- Cumulative end-to-end tutorial
Codex adds the package-local real-model reference driver, Live console adds the package-local browser console, and SDK extracts a reusable public SDK plus two standalone consumers. Runtime production Paperclip integration remains deferred; the App-owned production conformance adapter is test-only.
The SDK reference console opens in direct chat mode. Enter a normal prompt,
then open the protocol inspector to review events and reducer state. Expand a
Terminal row and its nested Debug details disclosure to inspect every
canonical event retained for that command. The header marker 🖇️ v0.1.2
identifies the current console iteration.