Files
DottaandPaperclip d6e6d1c78a Stream OpenCode immediately, scope reasoning to turns, and deduplicate instructions (#15708)
## Thinking Path

> - Paperclip helps people manage AI agents and observe their work.
> - The native runner carries provider output into the task transcript.
> - OpenCode sends incremental text in `message.part.delta` events.
> - The driver ignored those events and waited for a full part snapshot.
Accepted chats also waited after lifecycle verification was complete.
> - This pull request forwards each valid chunk as it arrives and
removes unnecessary isolated startup work.
> - Valid chunks reach Paperclip sooner without changing completion or
permission rules. The matched browser experiment below does not
establish a reliable end-to-end latency improvement.

## Linked Issues or Issue Description

**What happened?**

The native OpenCode driver withheld text until OpenCode sent a full
snapshot. A local delayed-provider probe with the pinned OpenCode
1.18.34 runtime showed a 1,496 ms gap between the first native text
delta and the first Paperclip text delta (1,510 ms from provider text).

**Expected behavior**

Forward valid assistant text and reasoning deltas immediately. Preserve
repeated identical tokens, message roles, turn ownership, and final
response selection.

**Steps to reproduce**

Send an empty text part, then text deltas, and withhold the final full
snapshot. The new fixture requires the consumer to acknowledge its first
text delta before releasing the snapshot. The previous driver cannot
pass this handshake.

**Paperclip version or commit**

Baseline: `011e5669bad61c435a4745fb10812bef70376841`. The patch is
rebased onto current master.

**Deployment mode**

Local native OpenCode 1.18.34 on macOS arm64. Live diagnostics use
`openrouter/deepseek/deepseek-v4-flash-0731`. These are narrow
diagnostics, not full Product E2E qualification.

## What Changed

- Keep the macOS verified OpenCode launch pathname until child exit,
with guarded reuse and cleanup on failed launches. This fixes a
reproduced asynchronous executable-lifetime failure.
- Dispatch accepted queued chats immediately after lifecycle readiness
commits, through normal scheduler admission. Preserve ready state if
dispatch fails.
- Validate native OpenCode model access with the real hello probe
without a redundant catalog scan. Keep legacy adapter discovery
unchanged and isolate unspecified native probe directories.
- Map incremental text and reasoning to the existing canonical event
path, using observed part and message identities.
- Keep additive chunks without IDs distinct. Preserve ID-based event
deduplication.
- Clear current streaming part metadata when a new turn starts. Reject
mismatched session/message identities and unknown parts or fields.
- Disable automatic updates, remote model refreshes, and default plugins
in each isolated OpenCode process. Keep assigned MCP tools and default
model reasoning intact.
- Retain the latest cumulative text/reasoning part while its message
role is unknown, including completion metadata after more than 100
chunks.
- Expose `startTurn({ message, reasoningMode: "disabled" })` for one
OpenCode/OpenRouter turn. Carry the selection in the durable command
through Rust and the proxy. An omitted/default option restores provider
defaults on the next turn, including warm and recovered sessions. There
is no environment-level setting or automatic selection policy.
- Group optional Rust turn inputs in `CodexTurnOptions` and dispatch
through `start_turn_with_options`; retain `start_turn` for defaults.
Future turn settings can be added without another feature-specific
method.
- Expose a native-session capability for per-turn reasoning and reject
explicit unsupported selections before dispatch, including direct ACPX
and Codex backends. Runnerd checks the durable selected model as well as
the provider.
- Remove duplicate agent instructions from the shared OpenCode proxy
task envelope for all users. Send one system copy on every turn,
including prepared and resumed turns, because OpenCode reads the latest
user message’s system field. Preserve task constraints and completion
contracts.
- Document startup and streaming behavior. Add regression coverage for
role ordering, repeated chunks, final snapshot deduplication, and stale
turns.

## Verification

### Lower-load browser retry — 2026-10-10

**All 8/8 trials succeeded. In this batch, median Send → first visible
reply fell from 30.586 s to 12.476 s (59.2% lower).** Four fresh-agent
trials per variant, phase order before/after/after/before with two
trials per phase. Every adjacent-phase pair favored the candidate,
although one difference was only 0.301 s. This is evidence for the
readiness/probe changes on this host, not a production percentile or a
guarantee across providers.

| Phase/trial | Variant | Send → reply (s) | Observed ready → run (s) |
Session open → turn start (s) | Model requests |
|---|---|---:|---:|---:|---:|
| 0/0 | before | 45.042 | 25.337 | 3.629 | 2 |
| 0/1 | before | 13.689 | 0.000 | 2.443 | 2 |
| 1/0 | after | 14.696 | 0.000 | 2.526 | 1 |
| 1/1 | after | 13.389 | 0.000 | 2.470 | 2 |
| 2/0 | after | 10.871 | 0.000 | 2.706 | 2 |
| 2/1 | after | 11.564 | 0.004 | 2.443 | 2 |
| 3/0 | before | 41.582 | 21.518 | 2.376 | 2 |
| 3/1 | before | 19.589 | 9.753 | 2.707 | 1 |

Observed ready-to-run median: **15.636 s → 0 s**; zero means execution
began by the first ready observation, not literally zero latency.
Polling interval is nominally 250 ms plus HTTP time.
Session-open-to-turn-start includes controller handling, not just child
startup. All eight ran within the existing startup timeouts; no timeout
or reasoning policy was changed. Host load at fixture creation stayed
**5.47–8.73 on 18 CPUs**, versus 68.9–100.8 in the prior campaign. No
builds/tests ran concurrently with measurement.

Same merged source **b0a0158729a7913882a0e756fe7e27228d72bf53**, release
runner, rebuilt static UI, OpenCode **1.18.34**, model
**openrouter/deepseek/deepseek-v4-flash-0731**, provider-default
reasoning, and empty isolated workspaces in both variants. Only the
readiness worker and explicit-model verification source were toggled;
the launch-lifetime fix and all proxy streaming changes were held
constant. Thus this does not separately measure the original streaming
changes. One browser per phase, two fresh companies/agents/chats per
phase. Measurement begins at capture-phase Send click and ends on
visible Hi text plus two animation frames. Status/thinking is excluded;
emoji can arrive just after the first text. Successful screenshots
inspected. Each variant had three two-request turns using
`paperclip_finish` and one one-request turn; no status lookup was
observed. Reported usage includes settlement after the first visible
reply.

Runner SHA-256
`57f658c220eecd2370ffd3333cd7e5d06d16bb2798af0cf43c84e01c8207c1c3`;
proxy
`bcc02fda2bf35b38a116cd880841aa6c534f6e0b497a96d97369a1e06dcc72ac`;
rebuilt UI tree
`eee238f2eda628c4e8c3660707986da4f8dd9100c814e26938fce3756df2eae4`
(sorted relative path, NUL, binary per-file SHA-256). Worker/probe
variant source hashes are unchanged from the preserved comparison below.
Measured turn cost **$0.00612574**, plus verification probes not
included in that figure. All eight attempts retained; all source/proxy
files restored and test supervisors exited.

Current commit validation: **53 successful CI checks, 4 skipped, zero
failing; Greptile 5/5 on b0a0158729 with no actionable findings.** The
previous 20 targeted post-merge tests passed. No source edit was made
for this rerun. Earlier failed and inconclusive cohorts below remain
historical evidence and are not pooled with this batch.


### Master update and failure diagnosis — 2026-10-09

Merged `origin/master` at `417db14e7d` into this branch without
conflicts; pushed merge commit
`b0a0158729a7913882a0e756fe7e27228d72bf53`. The incoming changes do not
alter OpenCode process startup or its timeout budgets. Historical
measurements below remain attributed to their original commits.

The four last-campaign failures never issued `turn.start`, so these were
**cold-session admission failures, not model replies failing**. Two
retained `session.open` commands eventually completed after controller
timeout/cleanup had begun. The other two ended with the inner
`thread/start` response timeout, and their diagnostics recorded OpenCode
startup retries. Startup event timestamps are buffered publication times
and must not be interpreted as exact process-spawn times.

On the merged source, two credential-free startup-only checks used the
real pinned OpenCode 1.18.34, empty temporary workspaces, a placeholder
key, and **no model turns**:

- Existing pinned executable: spawn at 1.083 s; failed health
attempt/retry at 11.311 s and 21.914 s; third process healthy at 32.210
s; session opened at **35.168 s**.
- Verified private executable path: snapshot preparation took **0.497
s**; process spawned at 1.588 s; healthy at 9.967 s; session opened at
**11.120 s**.

These are sequential diagnostic samples, not evidence that one
executable path is faster. Host load was **56.5–64.5 on 18 logical
CPUs**. The local health wait permits 10 s per attempt and the driver
retries up to three times; the outer controller's cold admission budget
is 30 s, and the Rust facade also waits 30 s for a response. Thus
retries plus setup/session creation can outlast the caller's budget. The
isolated 35.168 s success demonstrates that a session can eventually
open yet fail the app's admission deadline. Heavy load aggravates the
timing; its exact contribution is not isolated. No timeout was increased
and no new model-latency claim is made.

Merge validation: **20/20 targeted lifecycle-policy and heartbeat
run-completion tests passed**, including temporary PostgreSQL cases. The
first sandbox invocation skipped database cases; the authorized local-DB
rerun executed all 20. The package-manager launcher could not verify its
version via the unreachable registry, so validation used the
already-installed Vitest binary. Full CI/review subsequently passed on
this merge commit; see the October 10 retry above. Older results below
apply to their named commits.


### Follow-up: reproduced macOS launch failure — 2026-10-09

The retry exposed a startup bug rather than only host-load noise. The
macOS proxy deleted the private verified executable path immediately
after asynchronous spawn. A credential-free test of the actual pinned
OpenCode `--version` reproduced a fast SIGKILL and a 15-second timeout
in three immediate-cleanup attempts; three deferred-cleanup controls
succeeded. The timeout was killed by the diagnostic's own deadline. This
experiment does not establish the OS-internal reason for the fast
SIGKILL.

Commit `c30451e0ebf3f74e9ef70d37bbe4d9fe7fb56359` retains the private
launch path until child exit, cleans failed spawns, and waits for forced
shutdown before reuse. Source descriptor authentication, ownership,
mode, inode checks, and the prohibition on ambient executable fallback
remain. Overlapping reuse fails closed. The existing synchronous shell
test was changed to exercise the actual asynchronous boundary.

**Verification:** 59 focused proxy/driver tests passed; runner
TypeScript check passed. All six subsequent real OpenCode version
launches succeeded without timeout (4.11, 4.33, 5.73, 6.91, 4.95, 6.70
seconds). The native Rust binary is unchanged from the retry; no model
calls were made by these diagnostics.

Retained interrupted retry: before 42.56/18.69 seconds; after one
180-second timeout (`provider_process_exited` / health SIGKILL) and one
37.33-second success. The latter overlapped the credential-free
diagnostic and is not a clean matched timing. The next phase was stopped
before Send to diagnose the reproduced bug. Reported measured-turn cost:
$0.0016874, verification probes additional. These attempts are not
pooled with other campaigns.

<!-- launch-fixed-proof:start -->
**Completed retry, stopped by the declared failure rule after 6 of 8
planned trials. No reliable end-to-end speedup demonstrated.** At commit
`c30451e0ebf3f74e9ef70d37bbe4d9fe7fb56359`, the launch fix above was
held constant in both variants; only the readiness handoff and
explicit-model verification path changed. Planned phase order:
before/after/after/before, two fresh agents/chats per phase, one new
browser per phase. The second after phase had two consecutive failures,
so the final before phase was not run.

| Phase/trial | Variant | Send → first visible reply | Observed ready →
run | Outcome |
|---|---|---|---|---|
| 0/0 | before | No reply within 180 s | 2.777 s | session.open timeout
|
| 0/1 | before | 47.654 s | 9.383 s | Succeeded |
| 1/0 | after | 50.051 s | 0.087 s | Succeeded |
| 1/1 | after | No reply within 180 s | 0.151 s | session.open timeout |
| 2/0 | after | No reply within 180 s | 3.337 s | session.open timeout |
| 2/1 | after | No reply within 180 s | 1.434 s | session.open timeout |

Both successful screenshots show “Hi 👋”. Timing starts at the browser's
capture-phase Send click and ends at the first visible answer text plus
two animation frames; the first fragment can be “Hi” before the emoji
arrives. Thinking/status text is not counted. All four failures were
`runner_local_connect_failed: PRP command session.open timed out`
(`native_session_interrupted`), before a reply; none reported the
earlier health-stage SIGKILL. They are retained, not discarded as
outliers. One success per variant does not support a speedup estimate or
median comparison. Observed ready-to-run median was **6.080 s before
(n=2) vs 0.793 s after (n=4)**; this is sampled every nominal 250 ms
plus HTTP latency and is a lower bound, not precise tracing. Host load
at fixture creation ranged **68.9–100.8 on 18 logical CPUs**, a
substantial confound that worsened over the campaign. No local
build/test was run concurrently with this campaign.

Provenance: pinned OpenCode **1.18.34**, model
**openrouter/deepseek/deepseek-v4-flash-0731**, default reasoning in
both variants, empty synthetic workspaces and isolated probe homes.
Release runner SHA-256
`57f658c220eecd2370ffd3333cd7e5d06d16bb2798af0cf43c84e01c8207c1c3`;
current proxy SHA-256
`bcc02fda2bf35b38a116cd880841aa6c534f6e0b497a96d97369a1e06dcc72ac`;
built UI tree SHA-256
`7f0c2b7f586cb753408351574b2239e22311944add99af4f6f3562cd6d47bbcb`. This
holds the rebuilt runner and static UI constant. Provider-reported
measured turn cost **$0.00120649**; verification probes are additional.
Including the earlier interrupted retry, this follow-up reports
**$0.00289389** in measured turn usage; missing usage is not assumed
free. Prior cohorts remain separate below.

Validation on `c30451e0e`: **59 focused tests passed**, runner
TypeScript check passed, **6/6 real credential-free pinned OpenCode
launches passed**, fresh Greptile **5/5 with zero unresolved threads**.
CI: **51 success, 4 skipped, 1 failed**. The failed native compilation
job never reached compilation: Docker's auth endpoint timed out during
Buildx setup, then failed the same way on one retry
([job](https://github.com/paperclipai/paperclip/actions/runs/37994181674/job/114039737503)).
Full local checks were not rerun while measuring. The PR remains draft:
the launch defect is reproduced and fixed, but cold session startup
still exceeds the controller's admission budget on this host. Raising
that budget would not establish a faster first response.
<!-- launch-fixed-proof:end -->

### Readiness investigation and live rerun — 2026-10-09

**Measured queue fix; no reliable end-to-end speedup claim. Keep this PR
in draft.** Accepted chats could wait for the periodic queue sweep after
verification had already committed `ready`. The lifecycle worker now
dispatches through the existing scheduler immediately after that commit.
The normal company, pause, budget, concurrency, and task-ownership
checks still apply. Dispatch failure preserves readiness and leaves
recovery to the queue sweep.

Native OpenCode verification also avoids enumerating the full model
catalog before running its real hello request. Identifier validation
remains. Authentication failures, unavailable models, and timeouts still
prevent readiness. An unspecified local probe directory uses an empty
temporary directory.

**Queue-only experiment:** four before and four after attempts, in
before/after/after/before server phases, two fresh agents per phase.
Median observed ready-to-start delay fell from **19.429 s to 0.058 s**.
State polling was nominally 250 ms plus HTTP time, so the latter means
within observation resolution; it is not a precise 58 ms latency claim.
After timings were 100.20, 65.29, 63.78, and 76.97 seconds. Before had
one 59.65-second success and three 120-second response timeouts (one
provider health SIGKILL, one late success, one still running after
settlement and canceled). These censored results do not support a
success-only median speedup comparison. Focused tests overlapped the
last baseline phase, which further limits total-time interpretation.
Reported measured-turn cost: $0.0055383; verification probes and
incomplete/canceled usage are additional.

**Final matched experiment:** built static UI in both variants to avoid
pre-Send Vite compilation; current proxy, release runner, selected
model, default reasoning, and explicitly isolated empty probe
cwd/HOME/XDG held constant. Both worker handoff and model-probe changes
vary. Fresh browser/company/agent/chat per attempt. ABBA order, two
attempts per variant, 180-second response limit declared before the
batch, four-turn bound. Same Send-to-rendered-Hi metric as the earlier
experiment.

| Order | Variant | Send → visible Hi | Created → execution | Observed
ready → execution | Outcome |
| --- | --- | ---: | ---: | ---: | --- |
| 1 | Before | 62.66 s | 23.467 s | 2.830 s | Succeeded |
| 2 | After | >180 s | 14.200 s | Within polling resolution | Provider
health SIGKILL |
| 3 | After | >180 s | 13.877 s | Within polling resolution | Provider
initialization timeout |
| 4 | Before | >180 s | 144.256 s | 14.067 s | Still running after
settlement; canceled |

Only one of four attempts produced a timed reply. **There is no valid
overall speedup estimate from this batch.** The dispatch delay is
removed, but provider startup failures and heavy shared-host load
remain. A recorded host sample was load 93.55 on 18 logical CPUs. This
is a confound, not proof that every failure was caused by host load.
Provider-reported measured-turn cost was $0.00032752, with no reported
usage for the other three attempts; verification probes are additional.

Retained earlier setup attempts: one shared-memory startup failure
before model calls, and a separate development-UI batch with a
60.06-second baseline followed by a navigation timeout before Send.
These attempts are retained separately, not counted as successful
response timings. No live attempts were silently dropped.

Provenance: source base `53580f8fbd40b963f3fb7ba5172ff3ede2e9ef1f`,
measured patch committed as `40b54eafb`, then rebased as
`c590770dcc2bb1debce14037935a7031e5c3c1ac`. **These live timings predate
the rebase.** Source and binary hashes are retained below. The inner
harness label `after` selects the identical current proxy in every
attempt; the outer campaign handoff label is the actual before/after
variant. Owned benchmark processes stopped and original proxy restored.

<!-- readiness-check-status:start -->
Post-rebase validation at `c590770dcc2bb1debce14037935a7031e5c3c1ac`:

- Full local `pnpm -r typecheck`: **passed**, including the native
release build (1,098.6 seconds on the loaded host).
- Focused lifecycle and adapter tests: **26 passed**. The 14 database
integration tests did not run locally because their hard-coded 20-second
database setup hook timed out. A retry with a CLI timeout override
encountered the same explicit hook limit. The earlier pre-rebase focused
batch passed 77 tests; these results are not substituted for post-rebase
execution.
- Local `pnpm test:run` was started, then stopped to avoid duplicating
CI on the saturated host. Local `pnpm build` was not run separately.
**Current-head CI passed Typecheck + Release Registry, Build, all
general and serialized test groups, Rust checks, both runner Vitest
groups, and all eight browser E2E shards.** The last general shard
passed 1,179 tests (72 files).
- Fresh [Greptile
review](https://github.com/paperclipai/paperclip/pull/15708#issuecomment-6085742217):
**5/5**, reviewed this exact head, no actionable defect.
- Final CI rollup: **49 success, 4 skipped, 4 failure**. The failures
are [Compile isolated native
Runner](https://github.com/paperclipai/paperclip/actions/runs/37990274515/job/114024329242),
[Docker context
integrity](https://github.com/paperclipai/paperclip/actions/runs/37990274623/job/114024196511),
[Canary Dry
Run](https://github.com/paperclipai/paperclip/actions/runs/37990274623/job/114024196424),
and the dependent `ci / verify` aggregate. All three root failures are
Docker Hub HTTP 429 / unauthenticated pull-rate limits while fetching
Node base images. The native-image job was retried once and hit the same
limit. No source compiler/test failure was reported by those gates. CI
maintainers need authenticated pulls or to rerun after the registry
limit clears.
- Owned benchmark servers stopped. Local validation commands completed
or were stopped. The PR remains a draft: user-visible speed/reliability
is not qualified, and the Docker-dependent CI gates are not green.
<!-- readiness-check-status:end -->

<details><summary>Sanitized readiness evidence with all attempts and
source hashes</summary>

```json
{"campaign":{"order":["before","after","after","before"],"trialsPerPhase":1,"maxMeasuredTurns":4,"maxCampaignUsd":1,"sourceHead":"53580f8fbd40b963f3fb7ba5172ff3ede2e9ef1f","reasoning":"provider default, no override","proxySha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","purpose":"Final matched comparison of both optimizations; distinct from earlier queue-only campaign.","responseTimeoutMs":180000,"timeoutRationale":"Declared before this final campaign because earlier unchanged baselines sometimes exceeded 120 seconds; earlier timeouts remain retained.","ui":{"mode":"built static UI","treeSha256":"7f0c2b7f586cb753408351574b2239e22311944add99af4f6f3562cd6d47bbcb"},"priorCampaign":"combined-proof stopped after a development-UI navigation timeout before Send; no failed attempts erased."},"complete":true,"statistics":{"before":{"attempted":2,"successfulMeasurements":1,"medianTtfrMs":62657.2000002861,"rangeTtfrMs":[62657.2000002861,62657.2000002861],"medianObservedReadyToRunMs":8448.5,"medianRunStartToVisibleMs":37290.199951171875},"after":{"attempted":2,"successfulMeasurements":0,"medianTtfrMs":null,"rangeTtfrMs":null,"medianObservedReadyToRunMs":0.0,"medianRunStartToVisibleMs":null}},"measuredTurnCostUsd":0.00032752,"costCoverage":"Provider-reported measured chat turns only; harness verification probes are additional and not included.","readinessMetricLimit":"First observed ready state sampled every 250 ms plus HTTP duration; clamped to zero if execution started before the observation. This is a lower bound on ready-to-start delay, not a precise sub-millisecond metric.","results":[{"phase":0,"trial":0,"handoff":"before","ttfrMs":62657.2000002861,"answer":"Hi","statuses":["succeeded"],"error":null,"runErrorCode":null,"providerHealthKilled":false,"queueMs":23467.0,"queuedBeforeObservedReady":true,"observedVerificationMsAfterFixture":26390.0,"observedReadyToRunMs":2830.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":37290.199951171875,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00032752,"costUsdExact":"0.000327520","inputTokens":26864,"outputTokens":46,"cachedInputTokens":0,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"},{"phase":1,"trial":0,"handoff":"after","ttfrMs":null,"answer":null,"statuses":["failed"],"error":"TimeoutError: page.waitForFunction: Timeout 180000ms exceeded.","runErrorCode":"provider_process_exited","providerHealthKilled":true,"queueMs":14200.0,"queuedBeforeObservedReady":true,"observedVerificationMsAfterFixture":24028.0,"observedReadyToRunMs":0,"readyAndStartedInSameSample":true,"runStartToVisibleMs":null,"usage":[{}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":2,"trial":0,"handoff":"after","ttfrMs":null,"answer":null,"statuses":["failed"],"error":"TimeoutError: page.waitForFunction: Timeout 180000ms exceeded.","runErrorCode":"provider_initialize_timeout","providerHealthKilled":false,"queueMs":13877.0,"queuedBeforeObservedReady":true,"observedVerificationMsAfterFixture":18107.0,"observedReadyToRunMs":0,"readyAndStartedInSameSample":true,"runStartToVisibleMs":null,"usage":[{}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":3,"trial":0,"handoff":"before","ttfrMs":null,"answer":null,"statuses":["running"],"error":"TimeoutError: page.waitForFunction: Timeout 180000ms exceeded.","runErrorCode":null,"providerHealthKilled":false,"queueMs":144256.0,"queuedBeforeObservedReady":true,"observedVerificationMsAfterFixture":136901.0,"observedReadyToRunMs":14067.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":null,"usage":[{}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"}],"hostLoadSample":{"at":"2026-10-09T20:48:40.031Z","parallelism":18,"load":[93.546875,82.68017578125,75.9091796875]},"measuredChangesCommit":"40b54eafb","postRebaseCommit":"c590770dcc2bb1debce14037935a7031e5c3c1ac","postRebaseLimit":"Live benchmark predates the rebase. No post-rebase live result is claimed.","variantSourceHashes":{"before":{"server/src/modules/agent-lifecycle/application/worker.ts":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326","packages/adapters/opencode-local/src/server/test.ts":"4c10bacb55f2c7589e7b5c79d74781206ae075db8955f5b998065d6db3c0335a","server/src/adapters/registry.ts":"4a858763f69947d5745c7b8cb251a62fa404bc9683ade0ffb76282ed12761394"},"after":{"server/src/modules/agent-lifecycle/application/worker.ts":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7","packages/adapters/opencode-local/src/server/test.ts":"cd2ea3d51c3f1b4fb42261be701fb7f6b85704ad8e641e6db73d7c58e47364d3","server/src/adapters/registry.ts":"4a858763f69947d5745c7b8cb251a62fa404bc9683ade0ffb76282ed12761394"}},"model":"openrouter/deepseek/deepseek-v4-flash-0731","prompt":"Reply with Hi and an emoji and nothing else.","runnerBinary":{"profile":"release","sha256":"6134c17debda14a76da69ac9631a2d10dad482ebd58dbc9a4e4480ed652211c4"},"probeIsolation":"Explicit empty cwd and HOME/XDG folders under the synthetic fixture workspace in BOTH variants. No server checkout or ambient provider home is used for the hello request.","comparisonLabelNote":"Inner run.mts variant after selects the identical current proxy in every trial; outer campaign readinessHandoff controls before/after source changes."}
```

Queue-only campaign:

```json
{"campaign":{"order":["before","after","after","before"],"trialsPerPhase":2,"maxMeasuredTurns":8,"maxCampaignUsd":1,"sourceHead":"53580f8fbd40b963f3fb7ba5172ff3ede2e9ef1f","reasoning":"provider default, no override","proxySha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04"},"complete":true,"statistics":{"before":{"attempted":4,"successfulMeasurements":1,"medianTtfrMs":59654.5,"rangeTtfrMs":[59654.5,59654.5],"medianObservedReadyToRunMs":19429.0,"medianRunStartToVisibleMs":8843.5},"after":{"attempted":4,"successfulMeasurements":4,"medianTtfrMs":71128.7000002861,"rangeTtfrMs":[63778.90000009537,100195.59999990463],"medianObservedReadyToRunMs":58.0,"medianRunStartToVisibleMs":26542.849975585938}},"measuredTurnCostUsd":0.0055383,"costCoverage":"Provider-reported measured chat turns only; harness verification probes are additional and not included.","readinessMetricLimit":"First observed ready state sampled every 250 ms plus HTTP duration; clamped to zero if execution started before the observation. This is a lower bound on ready-to-start delay, not a precise sub-millisecond metric.","results":[{"phase":0,"trial":0,"handoff":"before","ttfrMs":59654.5,"answer":"Hi 👋","statuses":["succeeded"],"error":null,"queueMs":50590.0,"observedReadyToRunMs":25709.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":8843.5,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00087852,"costUsdExact":"0.000878520","inputTokens":25644,"outputTokens":264,"cachedInputTokens":28416,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"},{"phase":0,"trial":1,"handoff":"before","ttfrMs":null,"answer":null,"statuses":["failed"],"error":"TimeoutError: page.waitForFunction: Timeout 120000ms exceeded.","queueMs":41651.0,"observedReadyToRunMs":14656.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":null,"usage":[{}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"},{"phase":1,"trial":0,"handoff":"after","ttfrMs":100195.59999990463,"answer":"Hi 🙂","statuses":["succeeded"],"error":null,"queueMs":65893.0,"observedReadyToRunMs":116.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":33243.60009765625,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00075782,"costUsdExact":"0.000757820","inputTokens":27014,"outputTokens":175,"cachedInputTokens":26368,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":1,"trial":1,"handoff":"after","ttfrMs":65288.60000038147,"answer":"Hi 👋","statuses":["succeeded"],"error":null,"queueMs":43511.0,"observedReadyToRunMs":0,"readyAndStartedInSameSample":true,"runStartToVisibleMs":21227.60009765625,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00085326,"costUsdExact":"0.000853260","inputTokens":25166,"outputTokens":250,"cachedInputTokens":28160,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":2,"trial":0,"handoff":"after","ttfrMs":63778.90000009537,"answer":"Hi 👋","statuses":["succeeded"],"error":null,"queueMs":39218.0,"observedReadyToRunMs":0,"readyAndStartedInSameSample":true,"runStartToVisibleMs":23479.89990234375,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.000839,"costUsdExact":"0.000839000","inputTokens":27068,"outputTokens":238,"cachedInputTokens":26368,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":2,"trial":1,"handoff":"after","ttfrMs":76968.80000019073,"answer":"Hi 👋","statuses":["succeeded"],"error":null,"queueMs":46257.0,"observedReadyToRunMs":185.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":29605.800048828125,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00030582,"costUsdExact":"0.000305820","inputTokens":24694,"outputTokens":32,"cachedInputTokens":1792,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"ceffb90a6bc97ffaaf8ec7d0267c7bfa7738d321aaf2a618b5efb9b114bc88e7"},{"phase":3,"trial":0,"handoff":"before","ttfrMs":null,"answer":null,"statuses":["succeeded"],"error":"TimeoutError: page.waitForFunction: Timeout 120000ms exceeded.","queueMs":100979.0,"observedReadyToRunMs":3872.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":null,"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00190388,"costUsdExact":"0.001903880","inputTokens":25524,"outputTokens":852,"cachedInputTokens":55808,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"},{"phase":3,"trial":1,"handoff":"before","ttfrMs":null,"answer":null,"statuses":["running"],"error":"TimeoutError: page.waitForFunction: Timeout 120000ms exceeded.","queueMs":129391.0,"observedReadyToRunMs":24202.0,"readyAndStartedInSameSample":false,"runStartToVisibleMs":null,"usage":[{}],"workerSha256":"02adc7e1162fb60439869cec62b75d8fb22fec71155f80535c8e00211cde8326"}]}
```
</details>

### Matched live browser latency evidence — 2026-10-09

**Conclusion: the overall first-response speedup is not established.**
Twelve successful real browser runs produced a lower after median, but
the updated path won only 3 of 6 adjacent pairs. Median paired
improvement was -0.17 seconds (effectively tied). Do not treat the
deterministic streaming-forwarding improvement as proof that this PR
fixes end-to-end first-response latency.

- **Actual path:** local Paperclip browser → real server/database →
current optimized Rust runner → pinned OpenCode 1.18.34 → OpenRouter
`deepseek/deepseek-v4-flash-0731`. Default reasoning in both variants;
no disabled-reasoning override.
- **Prompt:** `Reply with Hi and an emoji and nothing else.` Fresh
company, agent, chat, and per-turn native provider process for each
trial; same QA fixture instructions and selected model.
- **Comparison:** current product/runner at
`53580f8fbd40b963f3fb7ba5172ff3ede2e9ef1f` held constant. The baseline
substitutes the exact `opencode-app-server-proxy.ts`,
`opencode-proxy-task-envelope.ts`, and `opencode-server-driver.ts`
contents from pre-PR `41fec8c8921f9fb19c13fe11e2c9b71bf1017317`. This
isolates those changed modules; it is not an entire old-checkout
comparison.
- **Metric:** browser capture-phase Send click to a rendered `Hi`
markdown answer plus two animation frames, with a screenshot of the
answer. Thinking, tool labels, and status indicators do not count. This
is a DOM/render measurement, not instrumented compositor pixel timing.
- **Order/bound:** predeclared `before, after, after, before` repeated
three times; six trials per variant; 120-second response timeout; $1
campaign guard. All 12 measured trials returned `Hi` and an emoji and
finished successfully; none omitted.
- **Result:** before median **48.08 s** (37.42–66.40); after median
**42.50 s** (27.68–63.32). The 11.6% lower median alone is not
persuasive evidence of a reliable improvement given the mixed paired
results and broad overlap.
- **Remaining delay:** retained queue timestamps for 10/12 trials show
**9.76–41.65 seconds from run creation to execution start**. This is
before the OpenCode path changed here. Its cause is not established by
this benchmark. The final two queue reads were unavailable after
automatic server shutdown; all twelve primary browser timings were
retained.
- **Cost:** $0.010517369 provider-reported across the 12 measured turns;
$0.011371499 including the successful setup pilot. This is separate from
the older diagnostics listed below.
- **Retained setup attempts:** two failures before Send (composer
locator and transpiler helper); one baseline startup failure before any
observed model request (`SIGKILL` while probing OpenCode health); one
successful current-code pilot whose selector missed the visible reply.
These are retained separately, not silently included/excluded as valid
timings. The measured batch used the corrected selector and the same
release binary for both variants.
- **Limits:** local development server and headless Chromium on macOS
arm64, small sample, variable queue timing/provider routing/cache. This
does not qualify production/cloud performance or isolate the
contributions of each changed module. Server, browser, and owned runs
were stopped; original proxy build artifact restored; no product source
files changed for this experiment.

| Pair | Before: Send → visible Hi | After: Send → visible Hi | Faster
variant |
| --- | ---: | ---: | --- |
| 1 | 66.40 s | 27.68 s | After |
| 2 | 37.42 s | 28.00 s | After |
| 3 | 39.75 s | 55.22 s | Before |
| 4 | 47.32 s | 63.32 s | Before |
| 5 | 48.85 s | 29.79 s | After |
| 6 | 51.31 s | 61.07 s | Before |

<details>
<summary>Sanitized measurement records, source/binary hashes, usage, and
retained attempt classification</summary>

```json
{"manifest":{"candidate":"53580f8fbd40b963f3fb7ba5172ff3ede2e9ef1f","baseline":"41fec8c8921f9fb19c13fe11e2c9b71bf1017317","comparison":"Same current product and runner; replace the three OpenCode proxy/driver source modules with exact pre-PR contents for baseline. Reasoning uses defaults in both.","model":"openrouter/deepseek/deepseek-v4-flash-0731","prompt":"Reply with Hi and an emoji and nothing else.","samplesPerVariant":6,"order":["before","after","after","before","before","after","after","before","before","after","after","before"],"maxTurnMs":120000,"maxCostUsd":1,"variants":{"before":{"sources":{"src/cli/opencode-app-server-proxy.ts":"e9b46c937d7ce3b51dc44181a2a3ad82befda9ebe184adf80c0a6b30da140fd6","src/cli/opencode-proxy-task-envelope.ts":"fc23f8a515c2c4d1551f6b95e5b38e6254849dec23f7396e76b646e170a6e006","src/drivers/opencode/opencode-server-driver.ts":"99de415e8fc3bb5fb6f8d7512d751c4c37d2aa77a18b0546b9fb232f1a6fb6c5"},"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b"},"after":{"sources":{"src/cli/opencode-app-server-proxy.ts":"fa800e9bccce02f05f1af83e225d3b5a1300d0a7d4c35ac548f400d6c0a83f79","src/cli/opencode-proxy-task-envelope.ts":"5c387614fc9832322686326b8482eb00d88699d491b05cf59cadcd94a41194ea","src/drivers/opencode/opencode-server-driver.ts":"a944f3bf7f1bf897daec3a5dc2aa34ee9d1a2133efcb3cf150b002095a0ec586"},"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04"}},"runnerBinary":{"profile":"release","sha256":"6134c17debda14a76da69ac9631a2d10dad482ebd58dbc9a4e4480ed652211c4"},"metric":"Browser Send capture-phase click to visible Hi answer, sampled on animation frames, plus two paint opportunities. Excludes reasoning/tool/status text.","setupAttempts":["setup-failure-1","setup-failure-2","startup-failure-3","pilot-selector-failure"]},"summary":{"before":{"attempts":6,"measured":6,"medianMs":48083.050000190735,"minMs":37420.5,"maxMs":66397.69999980927,"totalCostUsd":0.004945229,"toolCalls":[4,1,0,0,0,1]},"after":{"attempts":6,"measured":6,"medianMs":42504.049999952316,"minMs":27680.099999904633,"maxMs":63320.90000009537,"totalCostUsd":0.00557214,"toolCalls":[0,1,0,4,0,1]},"medianReductionPercent":11.602841334350234,"medianPairedImprovementMs":-168.59999990463257,"reportedCampaignCostUsdIncludingPilot":0.011371498999999998},"pairs":[{"indices":[0,1],"beforeMs":66397.69999980927,"afterMs":27680.099999904633,"improvementMs":38717.59999990463},{"indices":[2,3],"beforeMs":37420.5,"afterMs":27997,"improvementMs":9423.5},{"indices":[4,5],"beforeMs":39751.60000038147,"afterMs":55217.199999809265,"improvementMs":-15465.599999427795},{"indices":[6,7],"beforeMs":47318.90000009537,"afterMs":63320.90000009537,"improvementMs":-16002.0},{"indices":[8,9],"beforeMs":48847.2000002861,"afterMs":29790.900000095367,"improvementMs":19056.300000190735},{"indices":[10,11],"beforeMs":51311.7000002861,"afterMs":61072.40000009537,"improvementMs":-9760.699999809265}],"trials":[{"index":0,"variant":"before","startedAt":"2026-10-09T19:15:23.172Z","measurementError":null,"ttfrMs":66397.69999980927,"domMs":66390.59999990463,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.002612595,"costUsdExact":"0.002612595","inputTokens":55962,"outputTokens":1344,"cachedInputTokens":83456,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":["paperclip_get_task_context","paperclip_comment_on_task","paperclip_finish","paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":4,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:15:58.194Z","turn.submitted":"2026-10-09T19:15:58.311Z","turn.started":"2026-10-09T19:15:59.063Z","firstTextDelta":"2026-10-09T19:16:36.039Z","turn.completed":"2026-10-09T19:16:36.118Z"},"completedRunDurationsMs":[42915.0]},{"index":1,"variant":"after","startedAt":"2026-10-09T19:16:36.598Z","measurementError":null,"ttfrMs":27680.099999904633,"domMs":27636.599999904633,"answer":"Hi","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00030429,"costUsdExact":"0.000304290","inputTokens":26845,"outputTokens":28,"cachedInputTokens":0,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":3,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:17:00.717Z","turn.submitted":"2026-10-09T19:17:00.841Z","turn.started":"2026-10-09T19:17:01.508Z","firstTextDelta":"2026-10-09T19:17:05.934Z","turn.completed":"2026-10-09T19:17:06.205Z"},"completedRunDurationsMs":[14435.0]},{"index":2,"variant":"after","startedAt":"2026-10-09T19:17:08.762Z","measurementError":null,"ttfrMs":27997,"domMs":27984.599999904633,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00101016,"costUsdExact":"0.001010160","inputTokens":27544,"outputTokens":366,"cachedInputTokens":26624,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":["paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:17:26.192Z","turn.submitted":"2026-10-09T19:17:26.305Z","turn.started":"2026-10-09T19:17:26.936Z","firstTextDelta":"2026-10-09T19:17:40.688Z","turn.completed":"2026-10-09T19:17:40.826Z"},"completedRunDurationsMs":[17763.0]},{"index":3,"variant":"before","startedAt":"2026-10-09T19:17:41.376Z","measurementError":null,"ttfrMs":37420.5,"domMs":37364.40000009537,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.000693018,"costUsdExact":"0.000693018","inputTokens":24236,"outputTokens":282,"cachedInputTokens":27648,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":["paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:18:12.381Z","turn.submitted":"2026-10-09T19:18:12.616Z","turn.started":"2026-10-09T19:18:13.441Z","firstTextDelta":"2026-10-09T19:18:21.746Z","turn.completed":"2026-10-09T19:18:22.262Z"},"completedRunDurationsMs":[29216.0]},{"index":4,"variant":"before","startedAt":"2026-10-09T19:18:23.000Z","measurementError":null,"ttfrMs":39751.60000038147,"domMs":39705.7000002861,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.000265747,"costUsdExact":"0.000265747","inputTokens":26123,"outputTokens":77,"cachedInputTokens":0,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:19:01.000Z","turn.submitted":"2026-10-09T19:19:01.163Z","turn.started":"2026-10-09T19:19:01.893Z","firstTextDelta":"2026-10-09T19:19:07.169Z","turn.completed":"2026-10-09T19:19:07.384Z"},"completedRunDurationsMs":[15843.0]},{"index":5,"variant":"after","startedAt":"2026-10-09T19:19:10.050Z","measurementError":null,"ttfrMs":55217.199999809265,"domMs":55204.89999961853,"answer":"Hi 🙂","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00030421,"costUsdExact":"0.000304210","inputTokens":4294,"outputTokens":28,"cachedInputTokens":22543,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:19:59.877Z","turn.submitted":"2026-10-09T19:19:59.990Z","turn.started":"2026-10-09T19:20:00.713Z","firstTextDelta":"2026-10-09T19:20:07.290Z","turn.completed":"2026-10-09T19:20:07.658Z"},"completedRunDurationsMs":[15104.0]},{"index":6,"variant":"after","startedAt":"2026-10-09T19:20:09.550Z","measurementError":null,"ttfrMs":63320.90000009537,"domMs":63313.09999990463,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00260723,"costUsdExact":"0.002607230","inputTokens":137971,"outputTokens":951,"cachedInputTokens":1024,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":["paperclip_comment_on_task","paperclip_get_task_context","paperclip_comment_on_task","paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":4,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:20:59.730Z","turn.submitted":"2026-10-09T19:20:59.859Z","turn.started":"2026-10-09T19:21:00.501Z","firstTextDelta":"2026-10-09T19:21:19.416Z","turn.completed":"2026-10-09T19:21:19.608Z"},"completedRunDurationsMs":[26519.0]},{"index":7,"variant":"before","startedAt":"2026-10-09T19:21:20.519Z","measurementError":null,"ttfrMs":47318.90000009537,"domMs":47246.800000190735,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.000257965,"costUsdExact":"0.000257965","inputTokens":4700,"outputTokens":71,"cachedInputTokens":21407,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:22:03.924Z","turn.submitted":"2026-10-09T19:22:04.439Z","turn.started":"2026-10-09T19:22:04.772Z","firstTextDelta":"2026-10-09T19:22:09.915Z","turn.completed":"2026-10-09T19:22:10.343Z"},"completedRunDurationsMs":[16702.0]},{"index":8,"variant":"before","startedAt":"2026-10-09T19:22:11.124Z","measurementError":null,"ttfrMs":48847.2000002861,"domMs":48802.59999990463,"answer":"Hi 🎉","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.000327654,"costUsdExact":"0.000327654","inputTokens":44796,"outputTokens":32,"cachedInputTokens":0,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:22:57.182Z","turn.submitted":"2026-10-09T19:22:57.281Z","turn.started":"2026-10-09T19:22:57.988Z","firstTextDelta":"2026-10-09T19:23:02.915Z","turn.completed":"2026-10-09T19:23:03.035Z"},"completedRunDurationsMs":[9541.0]},{"index":9,"variant":"after","startedAt":"2026-10-09T19:23:03.579Z","measurementError":null,"ttfrMs":29790.900000095367,"domMs":29748.800000190735,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00034054,"costUsdExact":"0.000340540","inputTokens":24710,"outputTokens":59,"cachedInputTokens":1792,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":[],"eventCounts":{"turn.started":1,"item.delta":2,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:23:28.840Z","turn.submitted":"2026-10-09T19:23:28.954Z","turn.started":"2026-10-09T19:23:29.666Z","firstTextDelta":"2026-10-09T19:23:35.306Z","turn.completed":"2026-10-09T19:23:35.531Z"},"completedRunDurationsMs":[12732.0]},{"index":10,"variant":"after","startedAt":"2026-10-09T19:23:37.247Z","measurementError":null,"ttfrMs":61072.40000009537,"domMs":61061.300000190735,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00100571,"costUsdExact":"0.001005710","inputTokens":25307,"outputTokens":368,"cachedInputTokens":28160,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6691475ed55f10179ec6ee0023a1f6462982f763848da7cd5b1d0baf38633b04","toolCalls":["paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:24:33.721Z","turn.submitted":"2026-10-09T19:24:33.832Z","turn.started":"2026-10-09T19:24:34.601Z","firstTextDelta":"2026-10-09T19:24:43.135Z","turn.completed":"2026-10-09T19:24:43.278Z"},"completedRunDurationsMs":[17387.0]},{"index":11,"variant":"before","startedAt":"2026-10-09T19:24:44.710Z","measurementError":null,"ttfrMs":51311.7000002861,"domMs":51300.5,"answer":"Hi 👋","runCount":1,"statuses":["succeeded"],"usage":[{"model":"openrouter/deepseek/deepseek-v4-flash-0731","costUsd":0.00078825,"costUsdExact":"0.000788250","inputTokens":24316,"outputTokens":356,"cachedInputTokens":27648,"cacheWriteTokens":0,"usageSource":"per_run","costStatus":"reported"}],"bundleSha256":"6dd8726e72545341f6772e54ec5aeb6264dec74db67290df44930d87135efb6b","toolCalls":["paperclip_finish"],"eventCounts":{"turn.started":1,"tool.execution.started":1,"item.delta":1,"turn.completed":1},"eventTimes":{"session.started":"2026-10-09T19:25:31.347Z","turn.submitted":"2026-10-09T19:25:31.463Z","turn.started":"2026-10-09T19:25:32.208Z","firstTextDelta":"2026-10-09T19:25:39.405Z","turn.completed":"2026-10-09T19:25:39.563Z"},"completedRunDurationsMs":[15446.0]}],"additionalAttempts":[{"attempt":"setup-1","stage":"before_send","result":"non-editable composer locator","modelRequestObserved":false},{"attempt":"setup-2","stage":"before_send","result":"transpiler helper absent in browser","modelRequestObserved":false},{"attempt":"startup-pilot-before","runnerProfile":"debug","result":"no visible reply within 120s; OpenCode health startup SIGKILL; retained infrastructure failure","modelRequestObserved":false},{"attempt":"selector-pilot-after","runnerProfile":"debug","result":"Hi reply succeeded; browser selector missed it; no valid TTFR","providerReportedCostUsd":0.00085413}],"queueTimingCoverage":{"measuredTrials":10,"totalTrials":12,"records":[{"index":0,"runs":[{"createdAt":"2026-10-09T19:15:28.998Z","startedAt":"2026-10-09T19:15:53.558Z","queuedMs":24560.0}]},{"index":1,"runs":[{"createdAt":"2026-10-09T19:16:38.520Z","startedAt":"2026-10-09T19:16:54.157Z","queuedMs":15637.0}]},{"index":2,"runs":[{"createdAt":"2026-10-09T19:17:10.587Z","startedAt":"2026-10-09T19:17:23.561Z","queuedMs":12974.0}]},{"index":3,"runs":[{"createdAt":"2026-10-09T19:17:43.807Z","startedAt":"2026-10-09T19:17:53.564Z","queuedMs":9757.0}]},{"index":4,"runs":[{"createdAt":"2026-10-09T19:18:28.113Z","startedAt":"2026-10-09T19:18:54.011Z","queuedMs":25898.0}]},{"index":5,"runs":[{"createdAt":"2026-10-09T19:19:12.343Z","startedAt":"2026-10-09T19:19:53.993Z","queuedMs":41650.0}]},{"index":6,"runs":[{"createdAt":"2026-10-09T19:20:14.258Z","startedAt":"2026-10-09T19:20:53.866Z","queuedMs":39608.0}]},{"index":7,"runs":[{"createdAt":"2026-10-09T19:21:23.019Z","startedAt":"2026-10-09T19:21:54.298Z","queuedMs":31279.0}]},{"index":8,"runs":[{"createdAt":"2026-10-09T19:22:14.671Z","startedAt":"2026-10-09T19:22:53.933Z","queuedMs":39262.0}]},{"index":9,"runs":[{"createdAt":"2026-10-09T19:23:06.155Z","startedAt":"2026-10-09T19:23:23.941Z","queuedMs":17786.0}]}],"limitation":"The server shut down before a final API read; queue timing is retained for the first ten trials only."},"conclusion":"No reliable end-to-end speedup established. The after median is lower, but only three of six adjacent matched pairs were faster; median paired improvement is approximately zero. The deterministic streaming-forwarding fix remains separately supported.","limitations":["Local dev server and headless Chromium on macOS arm64; six fresh chats per variant; not a production/cloud latency guarantee.","Current product and current release runner held constant; only three OpenCode modules replaced by their exact pre-PR contents for baseline.","Default model reasoning unchanged; no no-reasoning override or decision policy.","Upstream routing/cache and queue time vary; trials run serially in predeclared counterbalanced ABBA order.","Render proxy: capture-phase Send click to visible nonempty Hi markdown, then two animation frames; screenshot verifies the rendered answer. This is not instrumented compositor pixel timing.","Two setup-only failures, one startup failure, and one invalid-measurement successful pilot are retained separately, not silently treated as successful timings."],"reproducibility":{"harnessSha256":"f93996b306c7d31c19bfcb9cc197653f867944ffb900d81a23a34f831e17c0c5","variantBuilderSha256":"4e88a7bbdc0de5d72601194ebe3b96d9c6f6eda5a33d8e0f00cf2d5d14db34c2","analyzerSha256":"2d129177a1b8033f768aad3c82dc8fff70ba0ba76e7270b2d8612ee760d4aaf4"}}
```

</details>

- Latest options-struct refactor: 33 provider tests and 86 backend tests
passed. Four port-dependent fixtures initially failed inside the sandbox
and passed with loopback access. Latest check snapshot: 51 successes, 4
skips, and 2 failures (`ci / e2e` and `ci / e2e shard (5/8)`). This
latency experiment does not make the PR merge-ready.

- Focused TypeScript suites passed: OpenCode driver, per-turn reasoning,
prepared-context recovery, proxy envelope, live sessions, Codex context
control, and native instruction measurements. The real runnerd/proxy
integration test confirms disabled followed by default in the same
session. A stale staged binary initially hid the new parameter;
rebuilding the release runner resolved it.
- Capability-guard regression suites passed all 53 tests; the real
runner/proxy boundary still accepts the per-turn selection. Both
OpenRouter and non-OpenRouter real runner/proxy cases passed, along with
62 existing OpenCode fixture tests. Runner TypeScript build passed after
the final correction.
- Affected Rust modules passed: 33 provider tests and 86 durable backend
tests. Malformed or unsupported reasoning selections do not dispatch a
provider turn or poison admission state.
- Pinned OpenCode 1.18.34 with a local OpenRouter-compatible fixture
sent three successive provider requests with reasoning omitted,
`enabled: false`, then omitted. Each request contained one copy of the
agent instructions. No paid provider call was needed.
- The reconstructed first-turn task envelope shrank from 4,356 to 460
characters. This is a payload-size measurement, not a claimed token or
latency reduction. Existing stored conversation history is not
rewritten.
- Prior head `8c07fe2e70267a7c8a3118f1071eca73ecdc6c8a`: all 53 checks
passed, with four intentional skips. Greptile is 5/5 with no open
findings. One bounded retry cleared first-test timeouts, a legacy
signoff lease failure, and CI runner shutdowns. The timed-out suites
also passed locally (14 auth-signal tests, 40 device-login tests, and
the runner resume case).
- Known reported diagnostic usage totals about $0.0029. Two prematurely
closed baseline warm-turn attempts are excluded from latency comparisons
and their usage is not fully captured; background provider usage is also
not fully accounted. No further paid runs are active.
- `pnpm -r typecheck`: passed. Runner TypeScript typechecks and build
also passed after the review correction.
- `pnpm test:run`: attempted and stopped after a reproducible failure in
the untouched Codex recovery test `archives a damaged prior epoch before
completing guarded same-task replacement with real runnerd`
(`server/src/services/native-runtime/native-session-resume.test.ts:1072`).
It returns `needs_review` / failed instead of done / succeeded,
including when rerun alone. The broad local suite is not green. The two
initial collection failures disappeared after building; their isolated
rerun passed all 1,565 tests.
- `pnpm build`: passed.
- Pinned real OpenCode plus a delayed local provider: first native delta
to Paperclip delta fell from 1,496 ms to less than 1 ms at millisecond
measurement resolution. This isolates forwarding delay, not
network/model inference.
- Two live driver samples after the patch reached first text in
2.28–2.62 seconds including session startup. Baseline first responses
took 12.37–13.43 seconds. Small unmatched samples had different
reasoning/cache outcomes, so these are observations, not an attributed
speedup ratio.
- Full compiled runnerd protocol diagnostic: completed with `Hi 👋`;
first assistant delta at 9.19 seconds and cleanup complete at 9.44
seconds. The seeded mock-control-plane harness used 27,512 input tokens
across two provider requests, split into 13,521 and 13,991. The extra
`answer_status_question` call posted `Hi 👋` to the mock task in 12 ms,
then caused another model request. A reconstruction contains 25 tools
(35.3k JSON characters), 12.8k system characters, and a 4.4k task
envelope; 3.9k instruction characters appear in both system and task
context. Provider-reported cost was $0.00073976; the pricing-table
estimate was $0.00395332. This is one integration diagnostic, not a
production latency claim.
- A separate full-runner sample with existing lazy tool discovery also
completed, but took 14.64 seconds through cleanup and called
`paperclip_finish`. Its usage ledger recorded 12,888 input and 11,776
cached-input tokens, with $0.00096344 provider-reported cost. This did
not establish a latency improvement, so the default tool exposure was
not changed.
- Four direct OpenRouter requests separately measured default reasoning
at 2.45/3.90 seconds to first text and disabled reasoning at 0.66/0.53
seconds. Default reasoning remains unchanged; disabling it requires the
explicit per-turn option.

## Risks

- Immediate readiness dispatch changes scheduling timing. It uses
existing admission checks and preserves the periodic sweep as recovery.
- Native model verification relies on the real request rather than
catalog enumeration. The final live batch had provider startup failures;
overall speed and reliability are not qualified.
- Incremental chunks without IDs cannot safely be deduplicated by
content. Repeated identical chunks are valid. Known-ID duplicates remain
suppressed.
- OpenCode uses its bundled model catalog plus the explicit
selected-model entry. New remote catalog metadata will not refresh
during isolated startup.
- Default provider plugins are disabled. This runner supplies
credentials and assigned MCP tools through its own controlled
configuration.
- Disabled reasoning is supported only on the OpenCode/OpenRouter path
and requires a compatible model. It applies to one turn and is not an
automatic latency policy.
- The matched live browser sample is small and does not establish a
reliable end-to-end improvement. It includes local queueing/startup and
browser rendering, but does not qualify production or remote sandbox
latency.

## Model Used

OpenAI GPT-6 (Codex). Tool use, code editing, and command execution were
used. The session does not expose a more specific model ID or
context-window size. Live diagnostic target: DeepSeek
`deepseek/deepseek-v4-flash-0731` through OpenRouter and OpenCode
1.18.34.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-10 07:15:41 -05:00
..
…

Paperclip Native Runner

This package is the standalone development boundary for Paperclip's native runner protocol, process supervision, durable transport, provider drivers, and normalized session backends. Rust owns the production runner under runner/; TypeScript provides the control-plane reference, browser SDK, scenario tools, and conformance oracle.

The package includes one coherent set of capabilities: PRP v1 validation and replay, a supervised local runner with a scripted fake harness, durable WebSocket delivery and recovery, qualified Codex, OpenCode, ACPX, Claude Managed, and AWS AgentCore drivers, live session and issue-thread surfaces, a public browser/React SDK, a standalone adapter demo, and a deterministic mock control plane. None of these surfaces imports or starts Paperclip's server, UI, CLI, or production database.

Connection continuations inspect the harness descriptor's optional toolRefreshOnResume capability. Native Codex, Claude Managed Agents, AgentCore, OpenCode, and qualified Claude/Codex/Grok ACPX profiles expose it; unqualified harnesses leave it false or absent. Changing tools can replace a provider process while retaining the provider conversation. Company, agent, task, workspace, model, instruction, and skill compatibility still gate recovery. An MCP-only assignment change can resume only when the selected harness explicitly supports refreshing tools.

Compact continuation messages include the active completion revision and criterion IDs plus a reminder to obtain an accepted paperclip_finish or paperclip_block result for this turn. Reports from earlier turns do not finish the new turn. Provider final text alone remains insufficient; governed waits and strict native completion validation keep their existing behavior. This reminder changes the resumed model input, not the tool catalog or automatic retry policy.

When recovery needs a fresh conversation, the server supplies a deterministic handoff through a lazy history loader at the fresh attempt boundary. It includes the original request, recent messages, resolved interaction summaries, agent replies, and document excerpts, with source identities and retrieval instructions. Reads and excerpts are bounded; the handoff has a 24,000-byte ceiling and explicit truncation/omission markers. Conversation reset boundaries, deleted messages, source quarantine, and secret redaction apply before model submission. Successful recovery does not fetch or replay the handoff. Legacy adapters advertise supportsToolRefreshOnResume for their selected harness: Claude and Codex CLI/ACP, Grok CLI, Gemini/Kimi CLI/ACP, and Cursor/OpenCode/Pi CLI. CLI adapters using environment tool delivery start each invocation with current endpoints and credentials, including resumed turns. ACP reloads current MCP bindings, including run-scoped credentials, while preserving the conversation; an unqualified/custom harness retains its restart fence. Legacy fresh attempts receive the same bounded handoff, including resume-failure fallbacks. Provider authentication repairs retain their existing fresh-session recovery behavior.

ACPX release declarations

acpx-profiles.json owns the package versions, command/profile digests, and required execution policies used by TypeScript and Rust. It contains no models. After changing it, run pnpm --filter @paperclipai/paperclip-runner generate:acpx-profiles. Normal build and typecheck reject stale generated declarations. Generation also checks installed dependency pins and agreement with Cursor's distribution manifest and immutable release attestation (cursor-contract.json). Cursor's per-platform closure pins are generated from cursor-distributions.json.

Codex startup and browser login prefer the CLI from the installed dependency graph. If that dependency is absent, they use an executable codex from the selected execution host's PATH. An explicit execution command takes precedence, and resumed sessions retain their recorded command. Older or newer CLI versions are allowed. Relative PATH entries use the selected process working directory. Compatibility is established by the actual protocol or login attempt. Missing executables and real protocol failures still return actionable errors. Package identity and executable containment remain checked. Browser login keeps its existing provider-specific credential home. Linux ARM64 retains its existing legacy login path because native execution is not qualified there.

Exact dependency pins remain release and ACPX artifact-qualification checks. They make published builds reproducible and verify sandbox artifacts; they do not add a version-number gate to ordinary native Codex startup or browser login.

Release packages retain the pinned Codex JavaScript dependency graph, but do not bundle Codex native binaries. The published manifest declares the official, exact-version platform packages as optional dependencies. npm installs the package for the consumer's operating system and architecture. Those declarations belong to the published server manifest; the bundled JavaScript wrapper delegates platform installation there so npm can keep native and legacy versions separate. The wrapper code and patched ACP bridge stay unchanged. Codex sandbox read roots include the separately installed native vendor resources without granting access to enclosing npm directories or credential homes. This packaging does not change agent defaults or the selected runner of an existing agent.

Every ACPX harness accepts an explicit caller-selected model without a Paperclip model allowlist. The adapter sends that ID unchanged and verifies the provider's effective model before prompting. An incomplete discovery catalog does not block selection; a provider rejection or mismatch fails without choosing a fallback. Qualification model selections live in test catalogs, separately from optional product defaults. The legacy resolved snapshot field qualificationModel contains the caller's selected model; its serialized name preserves recovery identities. Historical profile fixtures remain immutable evidence, not release declarations.

Pi credentials are selected from the explicit run environment and remain session-bound. The native Rust launcher forwards the controller-bound credential names, including custom provider references, without restricting Pi to an OpenRouter credential. Built-in providers can use their usual API-key environment variables. Custom Pi providers are supported through PAPERCLIP_PI_PROVIDERS: a JSON object containing the native models.json provider entries (without the outer providers key). The runner writes that configuration to the private Pi agent directory and binds its digest to recovery. API-key and header environment references are forwarded only from the explicit run environment; command-based credential resolution is unsupported. Select the provider/model ID explicitly in the agent configuration. The bundled ACPX package tests exercise unlisted model selection, rejection, exact acknowledgement, and replay on a loaded connection. Catalog membership and Cursor model-alias expansion do not determine the selected model. Provider activity adapters own native tool identities, evidence, and diagnostic usage projection; the shared driver and sidecar consume those hooks.

Public package surfaces

  • @paperclipai/paperclip-runner — production contracts, clients/backends, PRP validation/replay, canonical catalog/dispatcher, and compatibility check.
  • @paperclipai/paperclip-runner/testing — deterministic mocks plus PRP and semantic conformance kits. Tests and external conformance consumers import this explicitly.
  • @paperclipai/paperclip-runner/evals — versioned native-attempt metadata, fail-closed package/binary compatibility checks, and explicit runnerd artifact resolution for eval consumers.

The package root has no mock or scenario exports. Generic credential-free matrix orchestration lives in the workspace-private @paperclipai/paperclip-eval-kernel; scenario content and provider-backed eval campaigns remain outside the runtime package. See ADR 0001.

The two conformance surfaces intentionally prove different contracts. The existing runControlPlanePortConformance suite checks narrow PRP run/event persistence. CAPABILITY_HIGH_RISK_SEMANTIC_VECTORS and runSemanticConformanceKit compare normalized tool authorization, state, effects, audit, retries, conflicts, redaction, continuation, and terminal decisions. The production adapter stays App-owned and invokes Paperclip's real route/service authorities; it does not copy those rules into this package.

Quick start

Native provider debug-trace correlation uses an incremental index owned by its transport. Pending event lookups read only newly appended bytes, with a 1 MiB read budget per lookup; they retry until the observed suffix is indexed. Partial records remain pending, and trace replacement or truncation invalidates the index. Closing a transport clears its index. Other active transports cannot evict its progress. Records over 64 KiB are skipped by the correlation index without buffering or parsing their full contents; the original trace file retains them. Do not restore a full synchronous trace scan for each pending event: it blocks event delivery and can leave the board showing an active run after the provider turn has already ended.

The package also builds paperclip-runner-acpx-sidecar. This bounded v2 stdin/stdout bridge admits the pinned Claude, Codex, Grok and Cursor ACPX profiles. It validates the exact model, session identity, tool catalog, structured input, and terminal settlement at the process boundary. Copilot and Pi remain gated. Verified distributions are build-owned; no provider accepts an arbitrary executable. See the rich ACP capability report.

Install Cursor explicitly with paperclipai runtime setup cursor; npm installation does not download it. Run setup as the OS user that runs Paperclip (the service account for a managed service). Setup writes to that account's ~/.paperclip/runtimes/cursor/<platform>-<arch>/<closure-sha256>, so a system-wide npm installation can remain read-only. Isolated provider HOME/XDG settings do not redirect this cache. Container images continue to use their packaged assets. Configure a company secret binding for CURSOR_API_KEY or CURSOR_AUTH_TOKEN, select Cursor in the Runner configuration, and select an exact model ID. Agent is the default; Plan and Ask are explicit modes. Paperclip semantic questions are supported. Native AskQuestion and authoritative per-run dollar usage are unavailable. Accepting a native plan ends the planning run successfully while the task waits for the next instruction. See the Cursor release report.

A release includes all three platform daemons and the Linux provider-pack identity from its matching Daytona image. Run stage:release-binaries with a manifest that binds each daemon path and SHA-256, plus remoteProviderPack: {path, sha256} for the actual image's provider-pack.json. Assemble these assets after the normal build and include them in the server's vendored Runner output before npm packing. Assembly requires the provider pack's source revision to match sourceRevision and its ACPX profiles and Cursor distribution to match the current source pins. An independently rehashed older pack is rejected. Provider-pack builds replace the installed Copilot platform wrappers with relative launchers. These wrappers remain usable after the pack moves into an image. The build still rejects any wrapper that retains its temporary deployment path. Ordinary remote Cursor startup uses the packaged Linux daemon and verifies every image asset against that manifest. A mismatched image fails before the provider starts; install the matching package and image together.

Pi cold provider admission has an absolute 60-second budget. Warm run attachment checkpoints the old ACPX sidecar and starts a new one, so it uses that same budget while retaining the existing Runner authority. Live adoption and ordinary commands retain their 30-second bounds. Closing the transport cancels admission; an acknowledgement received after the admission deadline cannot revive it.

The board's transcript parser coalesces consecutive identical Pi runtime-failure display rows within one run, turn and session, including the exit handler and its late prompt rejection. Both original PRP facts remain in the run log. Distinct failure details, intervening retry activity and later turns remain visible. The profile-14 notice projection, wrapper bytes and terminal settlement remain unchanged.

Cursor candidate configuration accepts acpxSessionMode: "agent" | "plan" | "ask" (default agent). This selects the native Cursor mode independently of acpxPermissionMode and Paperclip task planning or company approvals. The mode is validated at the API boundary and is bound to provider admission and recovery; changing it cannot reuse an incompatible warm session. Other providers reject this setting. Cursor configuration requires an explicit model. Pending providers retain operator-controlled qualification admission.

Remote Codex sessions relay assigned app tools through the server's configured gateway. Small catalogs are sent directly. When a catalog would exceed the runner's 256-operation or 768 KiB contract limit, the server exposes paperclip_search_assigned_tools and paperclip_call_assigned_tool instead. Search returns bounded pages of names, descriptions, and input schemas. Each page intersects the session's pinned assignments with current gateway grants. An individual schema that exceeds a page returns an inputSchemaRef. The same search tool retrieves that schema in chunks via schemaTool and schemaOffset; discovery can continue past the large tool. Calls retain task ownership, work-mode restrictions, gateway authorization, approvals, and audit. Core task tools and the runner's completion tools keep their reserved space; no assigned tools are silently removed to fit the limit.

Native Claude skill assignments travel in the runtime-context snapshot through runnerd to the ACPX sidecar. After acquiring the provider lifetime lease, the host materializes the assigned bundles under the isolated Claude home's skills/ directory before launch. Reopening a provider refreshes that snapshot; project and ambient host settings remain excluded. This path is separate from the legacy claude_local adapter's remote skill staging.

The isolated Claude settings pin both model and availableModels to the user's requested ID. This keeps ACP from replacing an exact ID with a picker alias during selection and verification. Users can keep selecting models from the normal Claude catalog or entering custom IDs; unavailable models still fail at the provider rather than silently falling back.

ACPX Claude defaults to approve-all, shown as Full auto (approve all). OpenCode defaults to allow; native Codex defaults to never (no approval pauses). These defaults cover all assigned tools and connections, including provider-native operations. Full auto is resolved consistently for agent creation, adapter conversion, direct driver launches, and fresh/resumed turns. Explicitly stored restrictive modes still apply.

Native OpenCode forwards incremental text and reasoning parts as they arrive; the final full snapshot selects the final response without repeating streamed text. Delta parts must match an observed message and part identity. Identical chunks without event IDs remain distinct tokens. Each isolated launch disables automatic updates, remote model-catalog refreshes, and default plugins; the runner supplies the pinned executable, selected model, and assigned MCP tools. Provider inference and reasoning still contribute to time to first text. These startup settings do not disable model reasoning.

On macOS, the verified OpenCode executable keeps its private launch pathname until the provider process exits. Removing that pathname immediately after spawn can kill or stall the signed binary before its health check succeeds. Ownership, permissions, and file-identity checks remain enforced; cleanup also runs on failed launches, and retries create a fresh launch snapshot.

Call session.startTurn({ message, reasoningMode: "disabled" }) to disable reasoning for one OpenCode/OpenRouter turn. CapabilityLiveSession.sendMessage accepts the same option in its second argument. The choice travels in the durable turn.start command and selects an OpenCode model variant for that prompt only. Omitting it (or passing "default") on the next turn restores provider defaults, including on a warm or recovered session. It is not an environment variable or agent-wide setting. No automatic decision policy selects it. The model must support disabling reasoning; other runner providers reject an explicit selection before starting work. Native sessions expose support through the perTurnReasoning capability.

OpenCode's proxy sends agent instructions once per turn through the system prompt. The task envelope retains task-specific constraints and the completion contract, without repeating the system instructions. This applies to all users of the proxy; already-recorded conversation history is not rewritten.

The runner's authenticated bridge and controller still enforce company access, action claims, task modes, and governed approvals. Provider permission defaults do not change workspace isolation or grant credentials or connection access. approve-paperclip remains an optional narrower mode for assigned planning and task tools; approve-reads allows assigned reads; deny-all rejects requests. None of these restrictive modes is the default.

Restrictive profiles route supported permission decisions through durable runtime requests and the existing task interaction controls. Requests are persisted before presentation; answers are checked against the offered decisions and acknowledged by the sidecar before settlement. Unknown, stale and duplicate responses fail. Missing provider decision support remains a blocked disposition, not implicit approval. Company access checks still run for each Paperclip tool. Provider death expires pending promises; approvals are never replayed into a replacement process.

Automatic Paperclip/read allowances currently require the Claude SDK dispatch boundary. Grok preserves these restricted settings, but its ACP requests lack independently bound tool authority. Those operations require a supported operator permission decision; a missing interactive responder stops with approval_required. An explicitly selected approve-all policy permits unattended Grok work in an assigned sandbox. Paperclip authorization and governed approvals still apply.

Runnerd selects only qualified provider profiles. Claude Managed and AWS AgentCore receive immutable company-profile snapshots with explicit retention, spend, and invocation limits. No provider process receives a Paperclip API credential or unrestricted server environment.

Claude Managed resolves its API key from the company secret bound to the selected profile. AWS AgentCore uses workload identity only; long-lived static AWS access keys are intentionally removed from the runner environment.

The Rust core includes a bounded client for the sidecar protocol. It enforces request identity, event order, frame and queue limits, timeouts, redacted diagnostics, and process-group cleanup. Runnerd selects this package-local transport only through an exact qualified provider descriptor.

Before a later provider adapter consumes a valid sidecar event, the Rust core also requires its optional or mandatory run and turn scope to match the active execution. Process and diagnostic events can remain global. All operational, tool, input, permission, and terminal events require the exact active binding.

A package-local payload boundary decodes events only after that scope check. It validates control identities, terminal status, question sets, and the admitted runtime event types and bounded fields. It redacts diagnostic and retained event values again before they can enter provider state. Authoritative semantic tool arguments are validated for transport bounds and forwarded unchanged, including credential-bearing document and instruction content. The provider harness owns credential policy; redaction of logs and audit previews must not reject or rewrite execution arguments. Diagnostic detection requires explicit credential fields/assignments or recognizable key, Bearer, JWT, or PEM formats, not ordinary prose such as "credential handling" or dotted filenames.

Human question tools accept one complete payload.questionSet for text and choice questions. The control plane generates legacy questions entries with stable free-text option IDs. Legacy callers remain supported. Calls that supply both forms must describe the same complete form; partial forms remain invalid. The native recovery bridge uses the same projection for answer delivery.

The server validates canonical answer constraints before persistence. Regex matching runs in isolated workers with a one-second deadline and at most four active workers. A timeout or capacity error leaves the question pending. The ordinary and native answer paths both await this validation before persistence. Saved native answer delivery does not repeat regex matching.

Validated ACPX runtime events normalize into the same provider-neutral activity families as the direct Codex transport. Reasoning contents stay private. Tool targets are resolved within the workspace under the provider host's path semantics and receive a versioned sidecar boundary marker before becoming bounded, display-only PRP safe paths. Raw or unmarked provider locations fail closed. URI-scheme and Windows drive-shaped values require a separate sidecar attestation backed by an existing in-workspace entry or, for a not-yet-created edit target, an existing in-workspace parent. This preserves real POSIX colon filenames without treating arbitrary URI text as a path. Windows separators are canonicalized, and consumers must not reinterpret the display value as file-access authority. Operational semantic-result and terminal events remain reserved for the stateful adapter rather than being duplicated.

The ACPX provider reducer preserves that order while it tracks one active turn, bounded assistant text, semantic results, and pending tool or input correlations. Terminal events flush the final assistant message first and clear unresolved turn-scoped requests.

The package-local session bootstrap starts the bounded sidecar transport, verifies the qualified capability handshake and effective model, opens one identity-bound session, and confirms its run attachment. Any failed bootstrap terminates the process; session shutdown preserves persistent provider state. The session can then start one immutable-workspace turn, request interruption, and reduce polled events through the scope-first state boundary. A mismatched command acknowledgement or invalid event terminates the session fail closed. Polled semantic calls pass through the run-scoped authorized tool bridge before they can be returned to a caller. Before a follow-up turn releases settled tool receipts, runner-core suspends and reaps the idle sidecar/provider generation, then resumes the same verified persistent identity in a fresh generation. This prevents a late session-lifetime MCP callback from inheriting the next turn's event authority.

The Rust question-response validator checks the versioned response envelope against the exact persisted question IDs, answer modes, options, required answers, custom-answer policy, and text constraints before provider delivery. Tool results and structured question responses then use two-phase resolution: validate retained identity and schema, require the exact sidecar acknowledgement, and only then clear pending local state. Codex permission requests violate its pinned sidecar policy and terminate the session fail closed. Safe suspension is available only with no active turn or pending request. The sidecar must return the exact persistent session identity before runnerd terminates the local process. Already validated ACPX reducer events project into provider-neutral durable events only with an exact run, session, turn, and item binding. Raw sidecar envelopes and permission requests are not admitted at this boundary. A safely suspended session can be recorded as a bounded private checkpoint. The checkpoint binds the exact provider identity, run, catalog revision, and catalog digest and is replaced atomically before a later recovery attempt. Recovery releases the stored identity only after those bindings match the prospective session configuration exactly.

Run the complete contract gate with:

pnpm install --filter @paperclipai/paperclip-runner --lockfile=false --offline --ignore-scripts --dev
pnpm --filter @paperclipai/paperclip-runner verify

The verification command requires a stable Rust toolchain with cargo on PATH, in addition to Node.js 24.11+ and pnpm 9+.

Minimal Debian/Ubuntu hosts without root access can extract the required Playwright browser libraries into a user-owned cache and run the same acceptance sequence with:

pnpm --filter @paperclipai/paperclip-runner verify:rootless

The tracer's final line is stable:

{
  "schemaVersion": "paperclip.runner.conformance.output.v1",
  "runIdentity": {
    "runId": "run_conformance_0001",
    "sessionId": "session_conformance_0001"
  },
  "result": {
    "status": "succeeded",
    "summary": "Standalone Conformance fixture accepted."
  }
}

Run only the tracer with:

pnpm --filter @paperclipai/paperclip-runner trace:conformance

Replay the Replay happy path, run a Local session, or open the browser devtool:

pnpm --filter @paperclipai/paperclip-runner replay:fixture
pnpm --filter @paperclipai/paperclip-runner trace:local-runner -- --scenario happy-path
pnpm --filter @paperclipai/paperclip-runner trace:codex
pnpm --filter @paperclipai/paperclip-runner demo:live-console -- --host 127.0.0.1 --port 4174

# Live console: chat with a live session in the browser.
pnpm --filter @paperclipai/paperclip-runner console:live-console
pnpm --filter @paperclipai/paperclip-runner browser:dev --host 127.0.0.1 --port 4179

# SDK: open the public-SDK reference console and mini consumer.
pnpm --filter @paperclipai/paperclip-runner console:sdk

# Standalone: run the standalone legacy/native/kill-switch tracer and page.
pnpm --filter @paperclipai/paperclip-runner trace:standalone
pnpm --filter @paperclipai/paperclip-runner trace:standalone -- --feature-flag enabled
pnpm --filter @paperclipai/paperclip-runner trace:standalone -- --feature-flag enabled --kill-switch enabled
pnpm --filter @paperclipai/paperclip-runner demo:standalone

Live console provider-backed routes are loopback-only and reject wildcard/LAN binds. Browser mutations require same-origin Fetch Metadata, matching Origin, and JSON content; see the protocol-server tutorial for direct curl examples.

Direct live protocol qualification

The canonical direct live protocol suite lives in the separate paperclip-evals repository under evals/paperclip-runner/. Its live-mini.json roster is the complete 35-case Codex qualification lane. Build this package's TypeScript output, release paperclip-runnerd, package tarball, and dist-issue-thread viewer, then use the roster runner documented in that repository. The package ships the required orchestration entry point as paperclip-runner-eval-session (dist/cli/eval-session.js). Evalbook owns the consistent HTML matrix and read-only attempt drill-down pages.

The hosted full-campaign workflow, parallel matrix, credential boundaries, canonical report merge, and versioned S3 index are documented in docs/runner-protocol-live-evals.md.

This direct protocol qualification is separate from the stress-derived Runner workflow schedule below and from the full-stack browser model E2E suite.

Stress-derived workflow, chaos, and AWS AgentCore operations

The deterministic workflow scorer and the chaos schedule do not require provider credentials:

pnpm --filter @paperclipai/paperclip-runner test:runner-workflow-evals
pnpm --filter @paperclipai/paperclip-runner report:runner-chaos-evals

report:runner-live-evals is a paid, provider-backed command. Native Codex requires OPENAI_API_KEY; ACPX Claude requires ANTHROPIC_API_KEY; OpenCode candidates require OPENROUTER_API_KEY. The live matrix remains qualified-only and does not persist credential values. Candidate qualification uses eval-session --candidate-profile <pi|cursor|copilot> with an explicit model and a separately materialized pinned candidate pack. This option is a constructor-bound diagnostic opt-in; session JSON cannot enable a candidate. Missing credentials or unverifiable spend block paid qualification. Set PAPERCLIP_EVAL_MAX_CAMPAIGN_COST_USD to a positive finite number to bound additional scheduling after the observed campaign total reaches that value:

PAPERCLIP_EVAL_MAX_CAMPAIGN_COST_USD=12 \
  PAPERCLIP_EVALS_ROOT=/path/to/paperclip-evals \
  pnpm --filter @paperclipai/paperclip-runner report:runner-live-evals

# Run two scheduled native Codex executions only.
PAPERCLIP_EVALS_ROOT=/path/to/paperclip-evals \
  pnpm --filter @paperclipai/paperclip-runner report:runner-live-evals -- \
  --candidate codex-luna --limit 2

GitHub-hosted live campaigns additionally require the default branch, an allowlisted numeric actor ID, the protected runner-e2e-paid environment, and an explicit repository variable before scheduled runs are enabled. Manual dispatches accept the same candidate, case, and execution-limit selectors. The paid job uses the reviewed RunsOn Fleet label when RUNNER_E2E_AWS_ENABLED=true and otherwise stays on ubuntu-latest. Uploaded reports contain redacted observations and trace digests, not raw provider frames, prompts, credentials, tool arguments, or hidden reasoning.

The AgentCore proof-of-concept uses an AWS CLI v2 profile to provision a dedicated invocation role and scoped resources. Its local mode-0600 metadata file contains no access keys; probes assume short-lived STS credentials and clear them after use. Validate locally, provision or inspect the stack, run the bounded lab/smoke, and tear it down explicitly with:

pnpm --filter @paperclipai/paperclip-runner test:aws-agentcore-provisioning
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision -- --dry-run
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:probe
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:lab
pnpm --filter @paperclipai/paperclip-runner smoke:capability:aws-agentcore
pnpm --filter @paperclipai/paperclip-runner aws-agentcore:destroy -- --yes

To admit the hosted direct-eval workflow, provision with the account-local GitHub Actions OIDC provider and keep the default exact repository and protected environment binding:

pnpm --filter @paperclipai/paperclip-runner aws-agentcore:provision -- \
  --aws-profile paperclip-dev \
  --github-oidc-provider-arn arn:aws:iam::<account-id>:oidc-provider/token.actions.githubusercontent.com

This adds only repo:paperclipai/paperclip:environment:runner-e2e-paid as a web-identity subject on the scoped invocation role. The generated nonsecret profile records that role as both the local invocation role and the hosted execution role.

Provisioning can incur Bedrock, AgentCore Runtime/Memory, storage, and private networking charges. Provisioning refuses to modify a colliding stack unless its Paperclip ownership tags and template description match. A verified ROLLBACK_COMPLETE stack still requires --replace-failed-stack plus an interactive confirmation (or --yes) before it can be deleted and recreated. Destruction requires --yes and refuses to remove a stack with an active recorded lab unless --force is also supplied.

Package-owned commands

Command Purpose
build Compile the TypeScript public surface, Rust workspace, and browser devtool.
typecheck Check TypeScript, Rust, generated schema sources, and browser types.
test Run Rust/TypeScript fixture, supervisor, fake-driver, live/replay, and boundary tests.
check:forbidden-imports Reject TypeScript imports and Cargo path dependencies that cross into Paperclip core.
check:tracked-imports Reject tracked imports and package.json entry points that only resolve against untracked files, so a clean checkout of any commit builds.
check:numbered-milestones Reject numbered construction-milestone names in tracked package paths and source.
check:package-boundaries Enforce the acyclic runtime/testing/eval dependency and manifest boundary.
check:clean-consumers Pack the runner and install its root, evals, and testing exports in a clean consumer.
test:eval-slice Run the credential-free eval bundle, scoring, and behavior/fault slice.
test:runner-workflow-evals Run the deterministic provider-neutral workflow matrix.
report:runner-workflow-evals Validate deterministic fail-closed results and write JSON, Markdown, JUnit, and GitHub-safe reports.
report:runner-live-evals Execute the paid provider schedule and render its immutable attempts with the canonical paperclip-evals HTML grid.
report:runner-chaos-evals Write the credential-free eight-scenario chaos schedule.
test:aws-agentcore-provisioning Validate the AgentCore template and wrapper safety contracts without provisioning.
aws-agentcore:provision / probe / lab / destroy Manage the scoped AgentCore proof-of-concept lifecycle.
smoke:capability:aws-agentcore Exercise the qualified AgentCore profile through the capability harness.
check:conformance-parity Require byte-for-byte equivalent Rust and TypeScript tracer output.
check:replay-goldens Require all reducer snapshots and cross-language summaries to match checked goldens.
check:replay-parity Run TypeScript and Rust against the same Replay fixture summaries.
check:browser-tokens Reject component-local visual literals and require the standalone token layer.
docs:validate Validate local documentation links.
trace:conformance Run the Rust mock-core tracer, print the stable result, and exit.
trace:conformance:typescript Run the TypeScript reference tracer directly.
replay:fixture Validate and reduce a fixture to a final snapshot.
trace:local-runner Run one native local session through the Rust runner and fake harness.
trace:codex Run the mock core with a real, local skillless Codex app-server session.
demo:live-console Start the package-local HTTP/SSE server with server-only Codex authentication.
console:live-console Start the standalone browser devtool with the Live console on 127.0.0.1:4180.
console:sdk Start the public-SDK reference console and mini consumer on 127.0.0.1:4181.
test:sdk Run targeted browser-client, reducer-projection, and React component contract tests.
test:browser:sdk Exercise both consumers with the fake driver, keyboard/a11y checks, reconnect/replay, measurements, and screenshots.
record:sdk:codex Run both public consumers against a safe real Codex session and capture live screenshots.
check:capability-contract Verify the generated capability, legacy MCP, and eval traceability contract.
check:semantic-contracts Verify the provider-neutral semantic tool contract is current.
trace:live-runner Run the real runnerd/Codex semantic loop against the mock control plane.
demo:scenarios Start the Capability scenario explorer over the mock control plane on 127.0.0.1:4183.
console:issue-thread Start the Paperclip-style issue thread on 127.0.0.1:4184.
test:scenarios Run the scenario index, run-artifact, parity, explorer component, and route tests.
test:browser:scenarios Exercise both the scenario explorer and issue-thread browser contracts.
browser:dev Start the standalone live/replay browser devtool.
test:browser Exercise static replay and live scenarios, then capture temporary screenshots under ignored test output.
verify Run the complete deterministic Conformance through SDK acceptance sequence.
verify:rootless Extract Debian/Ubuntu browser libraries without root, then run verify.

Navigate

Codex adds the package-local real-model reference driver, Live console adds the package-local browser console, and SDK extracts a reusable public SDK plus two standalone consumers. Runtime production Paperclip integration remains deferred; the App-owned production conformance adapter is test-only.

The SDK reference console opens in direct chat mode. Enter a normal prompt, then open the protocol inspector to review events and reducer state. Expand a Terminal row and its nested Debug details disclosure to inspect every canonical event retained for that command. The header marker 🖇️ v0.1.2 identifies the current console iteration.

create_task accepts an optional initial status of backlog or todo. Use backlog when the user wants a saved task or plan without execution: assignment and the initial plan are committed without scheduling a wake, even when dependencies are already complete. Omitting status preserves immediate delegation (todo, or blocked for unresolved dependencies).