Files
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00

77 lines
4.8 KiB
TypeScript

import type { RunnerE2EResult } from "./types.js";
import { isBlockedConnectionResult } from "./connection-evidence.js";
const html = (value: unknown) =>
String(value ?? "").replace(
/[&<>"']/g,
(c) =>
({ "&": "&amp;", "<": "&lt;", ">": "&gt;", '"': "&quot;", "'": "&#39;" })[
c
]!,
);
/** An interrupted first-task journey is not evidence of failed downstream behavior. */
export function isIncompleteFirstTaskResult(
result: RunnerE2EResult,
errors: readonly string[] = [],
) {
if (isBlockedConnectionResult(result)) return result.cleanup === "passed" && result.failureClass !== "secret_leak" && errors.length === 0;
const checks = result.firstTask?.checks ?? [];
return result.status === "failed" &&
result.failureClass === "candidate_failure" &&
result.cleanup === "passed" && errors.length === 0 &&
checks.some((c) => c.notReached) &&
checks.every((c) => c.passed || c.notReached);
}
export function renderCaseOutcome(
result: RunnerE2EResult,
valid: boolean,
errors: readonly string[],
) {
const checks =
(result.firstTask?.checks.length ? result.firstTask.checks : undefined) ??
(result.matcherResults ?? []).map((m) => ({
id: m.matcher.kind === "json_path" ? m.matcher.path : m.matcher.kind,
passed: m.passed,
notReached: isBlockedConnectionResult(result) && !m.passed ? "Blocked before this checkpoint" : undefined as string | undefined,
detail: m.detail,
}));
const failures = checks.filter((c) => !c.passed && !c.notReached);
const notReached = checks.filter((c) => c.notReached);
const evaluated = checks.length - notReached.length;
const passed = checks.filter((c) => c.passed).length;
const incomplete = isIncompleteFirstTaskResult(result, errors);
const failed = !valid || result.status === "failed";
const reason =
result.failureClass === "secret_leak"
? /persisted Paperclip home/.test(result.error ?? "")
? "Credential-persistence check failed"
: "Secret / evidence safety check failed"
: (result.failureClass?.replaceAll("_", " ") ??
(result.cleanup === "failed"
? "Cleanup failed"
: "Run or evidence validation failed"));
const messages = [
...new Set(
[result.error, ...errors].filter((s): s is string => Boolean(s)),
),
];
return `<section class="case-outcome ${failed ? "outcome-failed" : "outcome-passed"}" aria-label="Case result explanation">
<strong>${isBlockedConnectionResult(result) ? `Blocked: ${html(result.providerConnection!.outcome.replaceAll("_", " "))}` : incomplete ? "Incomplete journey" : failed ? "Overall failed" : "Overall passed"} · ${checks.length ? `${passed}/${evaluated} behavioral checks passed${notReached.length ? ` · ${notReached.length} not reached` : ""}` : "No behavioral checks recorded"}</strong>
${result.providerConnection ? `<p>Target: ${html(result.providerConnection.target.origin)} · ${html(result.providerConnection.target.commit ?? "revision unavailable")} · ${result.providerConnection.assisted ? "Assisted login" : "Unassisted"} · ${html(result.providerConnection.authFreshness)} browser · ${html(result.providerConnection.entry)} / ${html(result.providerConnection.method)} · stopped at ${html(result.providerConnection.phase)}</p>` : ""}
${failures.length ? `<p>Failed checks:</p><ul>${failures.map((c) => `<li><strong>${html(c.id)}</strong> — ${html(c.detail)}</li>`).join("")}</ul>` : checks.length ? `<p>No behavioral matcher failed.${notReached.length ? " The journey is incomplete; unexercised checks are not passes." : failed ? " The overall failure came from a separate run, cleanup, or evidence check." : ""}</p>` : ""}
${notReached.length ? `<p>Not reached:</p><ul>${notReached.map((c) => `<li><strong>${html(c.id)}</strong> — ${html(c.notReached)}</li>`).join("")}</ul>` : ""}
${checks.length ? `<details><summary>See all behavioral checks (${checks.length})</summary><table class="matchers"><thead><tr><th>Result</th><th>Check</th><th>Detail</th></tr></thead><tbody>${checks.map((c) => `<tr class="matcher-${c.notReached ? "not-reached" : c.passed ? "passed" : "failed"}"><td>${c.notReached ? "Not reached" : c.passed ? "Pass" : "Fail"}</td><td><code>${html(c.id)}</code></td><td>${html(c.detail)}${c.notReached ? ` — ${html(c.notReached)}` : ""}</td></tr>`).join("")}</tbody></table></details>` : ""}
${
failed
? `<p><strong>${html(incomplete ? "Recording stopped before the journey finished" : reason)}</strong></p>${messages
.map((m) =>
m.length > 1200
? `<details><summary>Full failure details (${m.length.toLocaleString("en-US")} characters)</summary><div class="failure-reason">${html(m)}</div></details>`
: `<div class="failure-reason">${html(m)}</div>`,
)
.join("")}`
: ""
}
</section>`;
}