Files
PaperClipAI/tests/runner-e2e/PROVIDER-CONNECTIONS.md
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00

29 KiB

Live provider connections

provider-connections exercises real connection creation through the Apps catalog or new-agent wizard, then verifies a real agent task and follow-up. It is an explicit-only Product E2E suite. It is excluded from --all and the paid GitHub matrix: subscription login initially runs attended on a developer's machine.

Tasks are submitted through the production composer. When it accepts only a prompt, the harness assigns its fixture title through the public API. This changes naming metadata only; execution, status, attachments and downloaded bytes are produced and verified independently.

Run one journey

Create a private JSON configuration outside the repository. Its values reference credentials; do not paste keys into it. For example:

{
  "target": { "mode": "managed-local" },
  "browser": {
    "headed": true,
    "channel": "chrome",
    "profileDir": "/absolute/private/path/provider-qa-browser",
    "accountAlias": "developer-qa",
    "freshness": "signed-in",
    "loginTimeoutMs": 600000
  },
  "secretFile": "/absolute/private/path/provider-qa.env",
  "models": { "codex": "YOUR_CODEX_MODEL", "claude": "YOUR_CLAUDE_MODEL" },
  "routes": {
    "openrouter": { "model": "YOUR_OPENROUTER_MODEL" },
    "bedrock": { "model": "YOUR_BEDROCK_MODEL", "region": "YOUR_AWS_REGION" },
    "responses": {
      "baseURL": "https://your-gateway.example/v1",
      "credentialEnv": "QA_RESPONSES_KEY",
      "model": "YOUR_GATEWAY_MODEL",
      "auth": "bearer"
    }
  },
  "budgetCents": 200,
  "maxRuns": 3,
  "turnTimeoutMs": 300000
}

The secret file must be an owner-only regular file (chmod 600), with literal NAME=value entries. Only the selected variable is read. Shell substitutions are rejected. Environment variables take precedence. Default names are OPENAI_API_KEY, ANTHROPIC_API_KEY, XAI_API_KEY, GEMINI_API_KEY, OPENROUTER_API_KEY, and AWS_BEARER_TOKEN_BEDROCK. Override direct-key names with credentials: {"codex": "DEDICATED_QA_OPENAI_KEY"}. Subscription cases read no provider API key and never import CLI credential homes.

Routes may be shared by method or overridden per harness, such as "claude.openrouter". Protocol gateways require their own explicit credential reference, or auth: "none" for a deliberately unauthenticated endpoint. Messages gateways also support auth: "api_key" (x-api-key). Bedrock requires a current bearer token and region. The harness does not mint or refresh AWS tokens. Missing credentials/configuration produce a blocked result.

pnpm test:e2e:runner -- --list --suite provider-connections
pnpm test:e2e:runner -- \
  --id provider-connections.connection-codex-native.local.agent-api-key \
  --connection-config /absolute/private/path/connections.json
pnpm test:e2e:runner -- \
  --id provider-connections.connection-claude-legacy.local.apps-subscription \
  --connection-config /absolute/private/path/connections.json

Start with one cell. Running --suite provider-connections selects all 116 eligible cells and can spend money; it does not make missing routes available. Each command generates a separate campaign, preserving previous attempts.

Authentication handoff

Playwright drives Paperclip and opens the real provider sign-in page. When login needs a person, the terminal prints the provider and account alias and writes progress.json. Complete Google/provider login, MFA, consent, and any code entry in the visible QA browser. The test resumes only after Paperclip reports a saved, connected credential. Do not paste passwords or OAuth codes into chat.

The QA browser uses its own private profile. It refuses an existing unmarked profile, keeps a lock while running, and never uses your normal Chrome profile. signed-in permits previously saved website login; it still creates a fresh Paperclip connection. signed-out starts a temporary empty browser profile. Paperclip SSO and provider SSO may share an identity provider, so this records the browser's initial state, not a guarantee of another password challenge.

Human help is recorded as assisted. A login deadline records awaiting_user, not a pass, and stops the selected campaign before opening another login. Rerun the remaining cell IDs to start a fresh authorization after expiry; resumption is within the running process, not across a stopped process. Ctrl-C, termination, and hangup close the test browser pages, cancel recorded login attempts and active runs, performs fixture teardown, records interruption, and stops the campaign. Each cell has a thirty-minute outer deadline. A forced process kill cannot run teardown; inspect the private QA profile lock and target fixtures before rerunning. If Google rejects the automated browser, use the production device-code or code-return flow in your normal browser. Grok's current production flow may require the displayed terminal sign-in command. These are attended results.

Local and staging targets

managed-local owns a fresh server/database and removes its data on shutdown. To use an existing local server or staging, replace target with:

{
  "mode": "attach",
  "baseURL": "https://YOUR-INSTANCE.staging.paperclip.app",
  "expectedCommit": "EXACT_40_CHARACTER_DEPLOYED_COMMIT_SHA",
  "deploymentMode": "authenticated"
}

For local attach use http://127.0.0.1:PORT and local_trusted. Health must match the expected revision and deployment mode before credentials are read. The target is never reset, restarted, or given instance-setting changes. Sign in to Paperclip in the QA browser if needed; that session supplies the same public API permissions as the browser. The operator needs company/agent creation, connection management, and company archival permissions.

Deployment and execution environment are different dimensions. .local. means the agent runs in the target's local environment, including a staging server's local environment. .daytona. requires an already configured active Daytona environment on the selected target; this suite does not provision one. Use target.companyId and target.environmentId to select an empty dedicated company named Connection QA… with that environment. For a new QA company the harness uses the matching visible instance environment. Missing environments or a disabled native runner are explicit target blockers.

Each cell creates a separate QA company unless one was supplied. Teardown pauses its agents, cancels active runs, revokes its test connections, and archives the created company. This avoids requiring hard-delete support on staging and retains inspectable task history. Supplied QA companies must start without active agents or usable connections; revoked connection records and terminated agents from earlier attempts are allowed. Their new test agents are deleted and test connections revoked; historical records may remain. retainCompany: true keeps the fixture and credentials for diagnosis, with agents paused, and is recorded explicitly in evidence. Clean it up before reusing that company.

What qualifies as a pass

Every checkpoint must be verified:

  1. Expected target revision and deployment mode.
  2. Fresh connection saved with the selected method/provider/route.
  3. Connection visible after navigating away and reloading Apps.
  4. Agent created through the wizard with the selected harness/model/binding and requested execution environment.
  5. Real Configure-page environment probe passes.
  6. Real task run succeeds with that exact managed connection's attribution and requested environment in its durable run context.
  7. Agent decodes random input bytes supplied in the task, computes their sum/count/hash, and delivers an attachment whose JSON is checked independently by the harness.
  8. Tool evidence and attachment authorship/run attribution are present.
  9. A browser-submitted follow-up reuses the connection and delivers independently checked minimum/maximum/sum output.

The oracle input is base64 in the task description and must be decoded with a tool, preserving its exact bytes. This makes connection qualification independent of new-task attachment ingestion. The earlier uploaded-input OpenRouter attempt remains a failed attachment-path test: native staging currently admits chat attachments, and the new-task dialog uploads its files after scheduling the agent. Successful provider cells do not qualify that task-level upload path. 10. Downloadable artifacts remain visible after a page reload; cleanup succeeds.

An answer saying “provider routing works” cannot pass. Evidence records the configured model and agent configuration; connection attribution proves the selected credential, not an upstream gateway's internal model mapping.

The matrix covers Claude, Codex, Grok, Gemini, OpenCode, and Hermes local where Paperclip's managed-connection capability table admits them, with applicable legacy/native variants and both UI entry points. Cursor local, Kimi, and Pi remain visible product gaps for managed connection onboarding. Remote OpenClaw and Hermes gateways, Cursor Cloud, managed Claude/AWS AgentCore, process/HTTP, and generic legacy ACPX are deliberately excluded.

Reports, privacy, and CI

Results are under tests/runner-e2e/results/connections-…/: dashboard.html, campaign.json, normalized-results.json, connection-coverage.json, progress.json, and packaged per-cell evidence. The report uses the existing Product E2E grader, billing, history, and dashboard. Coverage reports required/selected/attempted/passed counts separately. A partial local campaign does not qualify staging or the whole matrix. Blocked credentials/login are incomplete; cleanup failures fail.

No Playwright traces, HARs, videos, automatic screenshots, raw assertion call logs, OAuth URLs/codes, or provider transcripts are retained. Only the final synthetic task can be screenshotted; it is private and has no public publication marker. Auth browser state stays outside the repository. Keep reports private. The result includes source revision, target revision, suite definition hash, model, method, assistance, run IDs, checkpoints, timing, and numeric usage. Before removing a managed instance, the harness retains terminal run status, log availability, and closed diagnostic codes for recognized quota, overload, file-conversion, session-configuration, or shell-restriction failures. It reads error records and failed tool receipts only; messages, stderr, model output, and reasoning are not copied into the result. An empty signal list leaves the cause unknown. Missing logs remain explicit and do not imply a healthy run. Attached-company cleanup captures these diagnostics after stopping runs and before deleting the fixture agent, since agent deletion also removes its logs. Cleanup revokes only this attempt's account IDs from successful creation or owned sign-in receipts. Concurrent campaigns' accounts are never adopted or revoked, even if both campaigns initially observed an empty QA company. UI probes may incur unreported spend, so billing is explicitly partial. Regenerate the report from retained results with pnpm test:e2e:runner:dashboard -- /absolute/path/to/connections-campaign.

API-key cells can run headless (browser.headed: false, channel: "chromium") on trusted CI once the target's board authentication is provided through a dedicated QA browser session. Subscription cells require headed mode and a person when challenged. Dedicated subscription accounts and a trusted attended runner are the next operational step; normal untrusted PR jobs never get these credentials. Credential-free support tests and catalog checks run separately:

pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit

The opt-in cancellation smoke uses a synthetic loopback page and cookie, with no provider calls or credentials. It verifies that Ctrl-C, termination, and hangup leave the browser's authenticated API client usable until fixture cleanup finishes: pnpm exec tsx tests/runner-e2e/connection-cancel.smoke.ts.

Initial verification (2026-10-03)

The local native Codex/direct OpenAI key/new-agent cell completed both real artifact turns with exact connection attribution and successful cleanup. The Apps/OpenRouter/native Codex cell created and reused its connection and passed its setup probe, then failed the task proof: its agent could not retrieve the uploaded input and marked the task blocked. This is not a gateway qualification pass. The Claude legacy/Apps subscription journey opened real sign-in and exercised the bounded awaiting_user outcome; login was not completed.

Credential-free verification: runner E2E typecheck, 855 support tests, both updated simulated local-browser-login regressions, catalog discovery, and retained-report regeneration passed. A fresh managed-local server was started and removed with an intentionally missing credential; no provider was called. Staging and the remaining cells have not been qualified. Use the generated campaigns for the exact revision, definition hash, timestamps, and results.

Local follow-up (2026-10-05)

Every eligible non-subscription local cell was attempted against the existing onboarding instance, preserving its original company data. The latest retained attempts now contain 43 passing journeys out of 46 non-subscription cells, including the Gemini investigation on October 6:

Harness Passing key/gateway journeys Eligible local key/gateway journeys Remaining result
Claude 16 16 Direct key, OpenRouter, Bedrock, and Messages passed through both entry points and runner generations.
Codex 12 12 The legacy/new-agent OpenRouter task and follow-up passed after the skill-root and artifact-helper fixes.
OpenCode 8 8 All three remaining OpenRouter journeys passed after the artifact-helper and canonical-workspace permission fixes.
Hermes local 4 4 A follow-up interrupted by lease release passed on one bounded retry. Workspace isolation was then corrected and retested separately.
Grok 2 4 Both legacy API-key journeys passed after preserving private session history and avoiding remote subscription restoration for missing local history. Native journeys require the missing pinned runtime.
Gemini 1 2 New-agent/API-key passed with gemini-3.5-flash-lite; Apps/API-key delivered its first artifact but timed out on follow-up. Installed CLI file creation still fails the isolated compatibility check.

There are another 12 local subscription cells. Attended Claude/OpenAI attempts opened real authorization and reached their human handoff, but have not qualified a completed subscription journey. Staging and Daytona remain untested. The Responses and Messages gateway proofs used the official provider endpoints; they do not qualify an Emissary deployment or every compatible gateway.

The run exposed a real new-agent loading race: an early API-key selection could be discarded when environment settings finished loading and remounted Connect. Connect now waits for those settings before exposing interactive choices. Saved connection selectors and native OpenCode model normalization were repaired in the harness; original failures remain in their original campaigns. Native OpenCode no longer sends the deprecated per-prompt tools override, which replaced its session permission policy; its remaining live OpenRouter failures were subsequently fixed and retested below.

Hermes also ignored the assigned task workspace when an agent had no explicit cwd, writing proof files into the server checkout. It now uses the resolved task workspace as its default. Explicit directory overrides retain their existing behavior. The real retest checks that the local proof file is in the assigned directory and matches the independently downloaded artifact. Earlier checkout spill files were preserved privately outside the repository.

An interrupted subscription attempt exposed another cleanup bug: Playwright's default signal handler closed the browser before the harness could cancel the login using its authenticated API client. The harness now owns browser shutdown, and a real-browser smoke verified cleanup under SIGINT, SIGTERM, and SIGHUP. One cancellation state now covers the entire campaign, including target startup, report generation, and teardown. Credential-free regressions verify that all three signals stop further cells during startup and reporting and still stop the owned target. Signal handlers remain installed until teardown finishes. An expired human handoff also stops the campaign instead of opening the next subscription window. The affected attempt remains failed in the evidence; its login was cancelled and its disposable company archived through the public API. The cleanup audit covers all 113 disposable companies and confirms that the original onboarding company remains available.

Hermes was tested using an isolated official CLI installation. Its temporary PATH launcher was removed after qualification; replay requires Hermes to be installed on the tested server's PATH. The report records the tested CLI version and source commit in qualification-environment.json.

Credential-free verification passed: runner E2E typecheck and all 856 support tests, 46 new-agent UI tests, 39 native OpenCode driver tests, 10 Hermes invocation tests, runner/Hermes package typechecks, and UI token gates. After the final browser signal change, runner E2E typecheck, 13 connection support tests, and the three-signal browser smoke passed again. This initial qualification stage did not repeat repo-wide typecheck, tests, and build; later verification is recorded below.

Credential follow-up, evening of 2026-10-05

The xAI dashboard identified expiry as the original key rejection: its QA key expired on September 27. The existing key was renewed for 90 days, through January 3, 2027, preserving its scope and limits. A Gemini replacement was saved in the owner-only secret file; the temporary exposed key was revoked and its absence verified after a dashboard reload. No provider billing settings changed. Both provider model-list requests returned HTTP 200. A separate bounded Gemini generation request succeeded with gemini-3.8-flash; the older gemini-2.5-flash request returned HTTP 404. Model-list success alone does not qualify inference or a product journey.

Six previously credential-blocked cells were rerun, followed by three bounded Gemini attempts after new diagnostics and a selector repair. At that stage none added a fully qualified journey, leaving 36 of 46 local key/gateway cells passing. Those 91 attempts and the subsequent repair attempts remain inspectable. These evening campaigns have October 6 UTC timestamps and belong to the developer's October 5 local qualification date.

  • Grok legacy, both entry points: the first task delivered an independently verified artifact. Each follow-up reported its session absent locally, attempted remote conversation restoration, and waited for subscription device authorization during an API-key journey. Investigate preserving isolated non-credential session state across disposable managed credential homes.
  • Grok native, both entry points: connection creation and reload passed, but Configure was not reached. An independent non-inference installation probe confirmed the required Grok Build 1.0.13 executable was missing. Provision it using the checked-in helper in doc/grok-native-runner.md, then rerun those exact cells; retain the runtime integrity checks.
  • Gemini: Auto completed ACP startup but timed out before task output. With gemini-3.8-flash, both entry points created and reloaded a connection, passed setup, and started the real task, then failed with acpx_session_config_failed. Installed Gemini CLI 0.58.0 rejected session/set_config_option for the model with ACP -32601 (Method not found). Resolve that runtime compatibility before claiming a pass. The Apps selector now reads the shared app definition's name (Google Gemini); its original selector failures remain retained.

After the catalog selector repair, runner E2E typecheck and all 13 connection support tests passed. The final cleanup audit found no usable QA connections, active QA agents, or active QA runs; the original onboarding company and test server remain available. The private report includes credential-refresh.json, key-retest-diagnostics.json, gemini-generation-probe.json, and cleanup-audit.json. None contain credential values or authorization codes.

The private connections-qualification-2026-10-05 dashboard selects the newest attempt per cell. retained-attempts.json and attempts.md retain all original campaigns. Source and suite definition hashes are preserved: these runs span harness repairs and an uncommitted working tree, so their passing count is historical evidence, not a common-build release qualification. Commit the reviewed fixes and rerun the required cells against that exact deployed revision before claiming the whole feature qualified.

Repair verification, late evening of 2026-10-05

The previously failing Codex OpenRouter journey, both legacy Grok API-key journeys, and all three OpenCode OpenRouter journeys now have passing browser retests with independently downloaded artifacts and successful follow-ups. Original failures remain in the report; no assistant success message alone qualifies a journey.

The fixes refresh Codex's current skill root on resumed turns, keep artifact helper temporary files inside the selected workspace, allow native OpenCode's assigned workspace under both its supplied and canonical filesystem paths, and keep credentials disposable. Missing managed local Grok history starts a fresh task handoff instead of triggering remote subscription recovery during an API-key turn.

On 2026-10-06, host-side Grok transcript retention/restoration was removed after security review: other agents running as the same OS user could read restored transcripts regardless of private file modes. Managed Grok follow-ups now start fresh with the Paperclip task handoff. The historical Grok passes above precede this change and do not qualify current transcript continuation; an isolated history solution and new live qualification remain required.

Gemini now receives its selected model through GEMINI_MODEL before ACP startup instead of an unsupported session/set_config_option call. Real tasks confirmed startup and tool execution. The original gemini-3.8-flash attempts subsequently hit daily inference quota; a small request for gemini-3.1-flash-lite succeeded. These are distinct from model-list and connection-probe success.

Local Gemini turns now also receive the current selected skill root and explicit instructions to read the Paperclip artifact workflow before making control-plane calls. A real follow-up exposed Gemini CLI's rejection of inline command substitution when building a JSON completion request. Its local turn guidance now uses workspace JSON body files with curl --data-binary @file or the selected skill's helpers; the shell restriction remains in place. Regression checks cover current roots across disposable homes for both Codex and Gemini. The shared ACP regression file now passes all 214 tests after isolating Gemini test homes from the developer's CLI state and adding canonical-workspace and remote-workspace regressions.

A credential-free real ACP reproduction isolated two filesystem compatibility problems behind Gemini's file-write errors. The client now translates a genuine missing-file read into the correct ACP resource-not-found error. This verifies the client protocol response; the installed Gemini CLI still mishandles that response, as isolated below. Local Gemini sessions also bind the filesystem client to the canonical workspace directory. Remote workspace paths are preserved. The existing ACPX dependency patch and its lockfile hash contain this narrow fix; four real protocol regressions cover missing-file semantics, external path denial, deny-all permissions, and other filesystem errors. The full affected ACP group passes 221 tests, including the existing three spawn safeguards. An earlier combined verification failure and the failed browser attempts remain retained.

The final filesystem-fix campaign (connections-2026-10-06T04-50-27-196Z-1ec2a9) still failed both Gemini entry points at the unchanged 300-second first-task deadline. Connection creation, reload, Configure probe, and binding passed. The new-agent turn created its local proof files, but neither turn completed the successful run, uploaded-artifact, and follow-up qualification. The earlier filesystem and model-configuration errors were absent from these tool receipts; that does not qualify the full workflow. Further work must investigate Gemini task execution latency and skill use with the configured CLI/model, then rerun the same independent oracle. No timeout increase or cancelled-run pass was used.

A further Gemini attempt exposed a harness settlement race: the first successful turn did not attach the file, while an automatic disposition-repair run did. The harness now waits for active repair runs within the existing configured limits and selects the successful run that actually uploaded the required file. It still rejects cancelled runs, wrong connection/agent attribution, wrong artifact bytes, and exceeded run/cost limits. The repair-attribution regression suite passed (58 tests across the targeted support/catalog files).

The task-composer helper was updated for the current production controls and its remembered project selection. A credential-free real browser smoke proved Plan and Ask mode, assignee/project selection, fixture task naming, and file upload. Public-API title updates only name fixtures; execution, credentials, model selection and artifact delivery remain real product behavior.

Full workspace typecheck and build passed. UI token gates and shell syntax passed. The broad Vitest attempt did not produce a green full-suite result: stale catalog/pool assertions were corrected and checked narrowly, while additional worker-start, hook and process timeouts remain recorded. Serial adapter/server rechecks were interrupted after more deadline failures; they do not constitute a passing full-suite run. UI/CLI and shared/database/skill-catalog groups passed. Further isolated checks do not substitute for a successful complete PR verification run.

The pinned native Grok executable is still absent at /opt/paperclip/providers/grok/1.0.13/grok. Installing it needs administrator access; retain the checked-in provisioning helper and its checksum verification. Subscriptions, staging/Daytona, and common-build release qualification remain outstanding.

Gemini investigation (2026-10-06)

The installed Gemini CLI is 0.58.0. A credential-free check now exercises its actual ACP file tools against the patched ACPX client, with a loopback model response and disposable home/workspace:

pnpm exec tsx tests/runner-e2e/gemini-acp-filesystem.smoke.ts

On the stock CLI, reading an existing file and rejecting an external write pass; creating a new file fails. Its ACP SDK rejects a plain JSON error object, while AcpFileSystemService.normalizeFileSystemError reads the message only from an Error instance. Consequently it misses the genuine resource-not-found message and fails the initial read-before-write. The same implementation is present in the upstream source. A temporary copy that also reads a plain object's string message passes all three checks. The installed CLI has not been changed, and the real browser qualification uses the stock CLI. A workflow that recovers using shell tools does not qualify the broken native file tool.

Retained child-stderr metadata also confirms HTTP 503 overload retries in the earlier deadline failures. A bounded availability probe returned HTTP 503 for gemini-3.8-flash and valid HTTP 200 output for gemini-3.5-flash-lite. With Flash-Lite, the full new-agent/API-key journey passed, including two independently downloaded and verified artifacts and an attributed successful follow-up. The initial Apps/API-key journey created, reloaded, tested, and bound the connection, then failed its first task. It did not retain a cause before managed-instance teardown; that gap prompted the closed diagnostic projection above. One diagnostic retest (connections-2026-10-06T13-25-36-718Z-20ee5d) delivered and independently verified the first artifact, then exceeded the unchanged 300-second follow-up deadline. Both run statuses and log availability were retained before cleanup, but neither log contained a recognized error signal. The follow-up cause remains unknown; overload in other runs is not proof of overload in this run. No deadline or oracle was relaxed.

The CLI requests HIGH thinking by default. Short bounded comparisons did not establish that default thinking causes the long workflow delay. Explicit effort configuration still uses an ACP method this CLI rejects; supported startup modelConfigs settings need a separate adapter fix and local/remote validation. Do not silently drop the user's setting or claim that a model-list/setup probe qualifies execution.

The original 3109 instance is currently stopped, and its temporary config and encryption key are missing. Its database was left untouched; substituting a new key would not recover the existing credentials. Fresh investigation campaigns use managed instances with separate databases and keys, archive their owned companies, and remove their owned servers. The private consolidated report retains all earlier failures, exact model/source provenance, and gemini-investigation.json. These mixed working-tree attempts do not qualify staging or a common committed release build.