Files
PaperClipAI/tests/runner-e2e/live-fixtures.ts
DottaandPaperclip caf120105c test: prepare neutral native connection guidance evals (#15407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need to discover connections, obtain consent, and continue
from saved decisions.
> - We want to reduce repeated instructions only when measured behavior
supports the change.
> - The existing decline tasks tell the model not to retry. One
provider-decline check can pass without an explanation or an observed
service counter.
> - This PR adds neutral tasks and stricter saved-evidence checks before
any connection instruction reduction.
> - Production instructions remain unchanged. The new cells are
configured, not live-qualified.

## Linked Issues or Issue Description

Refs #15218. Refs #15389.

**What existing behavior does this improve?**

The Product E2E connection workflow evaluation and its instruction
measurement provenance.

**Current behavior**

Some decline prompts supply the policy they intend to test. The
provider-decline workflow does not require a saved post-decision
explanation. Its old no-call check can use a missing fixture counter as
zero. OpenCode has no connection cases in the original Everyday matrix.

**Proposed behavior**

Add an explicit-only suite with five connection stories on native Codex,
ACPX Claude, and OpenCode. Require an explanation attributed by exact
run ID after a saved decline. Observe the provider fixture counter.
Preserve the original cases and grades.

## What Changed

- Add fifteen configured cells with one attempt, twelve-minute
deadlines, and verified 1,000-cent company and agent budget stops.
- Remove procedure hints from the three new decline prompts. Keep a
user-permitted explanation fallback and the existing positive controls.
- Require saved decline state, one decision, unchanged connections,
observed zero service calls where applicable, and a post-decision
explanation from a successful run on the same task.
- Add negative grader calibration and test the actual fixture budget
payloads. Exclude the suite from default and generic selection.
- Extend the existing full-catalog measurement source manifest with
connection descriptions and schemas. Add an audit of fixed text, tool
descriptions, returned instructions, and unqualified behavior.
- Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and
preserve its new Cursor suites. No production, credential, workflow, or
lockfile change.

## Verification

- Before rebase: Product E2E support passed 1,424 TypeScript tests and
128 Node checks. Six catalog measurement tests, repository
typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery
passed.
- The full pre-rebase repository test run was stopped when master
advanced. Its partial result is not a pass.
- After rebase and the review correction: repository build/typecheck,
Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128
Node checks, six measurement tests, and exact fifteen-cell discovery
pass. The duplicate local full-suite run was stopped incomplete after
about 20 minutes once complete CI passed; no local full-suite pass is
claimed.
- Review found that the initial grader read `runId` instead of public
`createdByRunId`. A regression calibration reproduced both rejection of
valid public comments and acceptance of the wrong alias. The fix uses
the actual field and binds the evidence type to the shared
`IssueComment` contract. A subsequent type-only import path correction
passes Product E2E typecheck.
- Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes
[complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545):
51 successful checks and two intentional Storybook skips, plus separate
Snyk success. Fresh Greptile review is 5/5 with the single review thread
resolved and no new findings. The PR is clean and mergeable.
- Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm
test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm
test:e2e:runner -- --list --suite native-connection-guidance`. The
measurement uses
`PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json
pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/native-procedure-measurement.test.ts`. Validation
used pinned pnpm 9.15.4.
- No paid provider campaign was started. There is no baseline/candidate
behavior result for these new cells.
- The audit records 654 UTF-8 bytes of fixed connection guidance. A
clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41
supplied tools, 53,341 normalized bytes at start/resume, 50,949 at
compact continuation, and a 48,195-byte authenticated OpenCode MCP
catalog. These are byte counts, not tokens, bills, vendor-private prompt
sizes, or savings from this PR.

## Risks

- This is eval preparation. Passing support tests do not establish live
model behavior or qualify an instruction reduction.
- The explanation oracle checks attributed saved output. It does not
prove cognition or arbitrary prose truthfulness. One saved interaction
also does not prove the absence of repeated idempotent tool calls.
- Successful new authentication and tool refresh, existing-connection
agent grants, independent work while waiting, explicit retry after
decline, and blocking when mandatory work remains still need separate
coverage.
- Notion setup decline does not execute a real Notion service. Positive
service approval uses an already installed deterministic service; it
does not qualify new connection creation.
- The original historical failures remain unchanged. Future comparisons
must freeze source, fixture, model, input, and grading controls and
retain every actual attempt.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code
execution. The exact serving model ID and context-window size were not
exposed in this session; they are not inferred. No model provider was
invoked by the eval suite in this PR preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:28:16 -05:00

379 lines
13 KiB
TypeScript

import { NATIVE_COMPLETION_BUDGET_CENTS } from "./native-completion-defaults.js";
import path from "node:path";
import { installedReleaseDaytonaPlugin } from "./installed-release.js";
import { isManagedHiringCase } from "./chat-cases.js";
import { FixtureRegistry } from "./fixture-registry.js";
import { TASK_TITLE_BUDGET_CENTS } from "./task-titles.js";
import { stageGrokSubscriptionFixture } from "./grok-subscription-fixture.js";
import type { RunnerApi } from "./api.js";
import type {
CredentialName,
MatrixExecution,
SecretReference,
SecretReferenceMap,
} from "./types.js";
interface CompanyRecord {
id: string;
issuePrefix?: string | null;
name: string;
}
interface SecretRecord {
id: string;
}
interface PluginRecord {
id: string;
pluginKey: string;
status: string;
}
interface EnvironmentRecord {
id: string;
driver: string;
config?: Record<string, unknown>;
}
interface AgentRecord {
id: string;
name: string;
companyId: string;
}
interface ManagedAccountFixture {
connectionId: string;
binding: {
provider: "openai" | "anthropic" | "openrouter";
method: "api_key";
mode: "responsible_user";
};
}
interface ProjectRecord {
id: string;
name: string;
primaryWorkspace?: {
id: string;
cwd?: string | null;
} | null;
}
export interface LiveFixtureValues {
company: CompanyRecord;
secretRefs: SecretReferenceMap;
environment: EnvironmentRecord;
agent: AgentRecord;
project?: ProjectRecord;
aiConnection?: ManagedAccountFixture;
onboardingRuntime?: {
mode: "production-wizard" | "post-onboarding-runtime-switch";
originalAdapterType: string;
originalModel: string | null;
testedAdapterType: string;
};
teardown(): Promise<void>;
}
function value<T>(resolved: ReadonlyMap<string, unknown>, id: string): T {
const result = resolved.get(id);
if (!result) throw new Error(`Missing resolved fixture ${id}`);
return result as T;
}
async function deleteDaytonaEnvironment(api: RunnerApi, environmentId: string) {
const deadlineAt = Date.now() + 120_000;
let lastError: unknown;
while (Date.now() < deadlineAt) {
try {
await api.delete(
`/api/environments/${environmentId}?destroyReusableSandboxLeases=true`,
{
allowNotFound: true,
},
);
return;
} catch (error) {
lastError = error;
await new Promise((resolve) => setTimeout(resolve, 3_000));
}
}
throw new Error(
`Daytona lease cleanup failed: ${lastError instanceof Error ? lastError.message : String(lastError)}`,
);
}
export async function setupLiveFixtures(input: {
api: RunnerApi;
execution: MatrixExecution;
executionNonce: string;
workspacePath: string;
credentials: Partial<Record<CredentialName, string>>;
daytonaImage?: string;
}): Promise<LiveFixtureValues> {
const { api, execution } = input;
const registry = new FixtureRegistry();
if (execution.environment.id === "daytona") {
registry.register<PluginRecord>({
id: "sandbox-provider",
async setup() {
return api.post<PluginRecord>("/api/plugins/install", {
packageName: process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_CLI
? installedReleaseDaytonaPlugin(process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_CLI, process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_DAYTONA_PLUGIN, process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_DAYTONA_PLUGIN_VERSION)
: path.resolve(
import.meta.dirname,
"../../packages/plugins/sandbox-providers/daytona",
),
isLocalPath: true,
});
},
async teardown() {
// The plugin is installed only in the isolated instance/database. The
// launcher removes that complete instance after the environment lease
// has been destroyed, so no global uninstall mutation is necessary.
},
});
}
registry.register<CompanyRecord>({
id: "company",
async setup() {
return api.post<CompanyRecord>("/api/companies", {
name: `Runner E2E ${execution.id} ${input.executionNonce}`,
description: "Ephemeral paid full-stack runner acceptance fixture",
budgetMonthlyCents: ["native-completion", "native-instruction-consolidation", "native-connection-guidance"].includes(execution.suite.id)
|| (execution.suite.id === "everyday-workflows" && ["hire-reuse", "delegate-feedback"].includes(execution.task.id)) ? NATIVE_COMPLETION_BUDGET_CENTS
: execution.suite.id === "task-titles" ? TASK_TITLE_BUDGET_CENTS
: execution.suite.id === "stock-harness" ? 1_000 : 0,
});
},
async teardown() {
// The isolated instance/database is removed by the launcher. Do not call
// company deletion here: metered runs intentionally retain cost-event
// references until that instance-wide teardown.
},
});
registry.register<SecretReferenceMap>({
id: "secrets",
dependencies: ["company"],
async setup(resolved) {
const company = value<CompanyRecord>(resolved, "company");
const refs: SecretReferenceMap = {};
for (const credentialName of execution.requiredCredentials) {
if (credentialName === "GROK_AUTH_JSON") continue;
const rawValue = input.credentials[credentialName];
if (!rawValue) throw new Error(`Missing credential ${credentialName}`);
const secret = await api.postSensitive<SecretRecord>(
`/api/companies/${company.id}/secrets`,
{
name: `Runner E2E ${credentialName} ${input.executionNonce}`,
key: credentialName,
value: rawValue,
description: `Ephemeral credential for ${execution.id}`,
},
);
refs[credentialName] = {
type: "secret_ref",
secretId: secret.id,
version: "latest",
} satisfies SecretReference;
}
return refs;
},
});
const grokSubscription = execution.profile.credential === "GROK_AUTH_JSON";
if (grokSubscription) {
registry.register<() => Promise<void>>({
id: "subscription-login",
dependencies: ["company"],
async setup(resolved) {
const raw = input.credentials.GROK_AUTH_JSON;
if (!raw) throw new Error("Missing credential GROK_AUTH_JSON");
return stageGrokSubscriptionFixture({
raw, companyId: value<CompanyRecord>(resolved, "company").id,
environment: process.env,
});
},
async teardown(remove) { await remove(); },
});
}
registry.register<EnvironmentRecord>({
id: "environment",
dependencies: [
"company",
"secrets",
...(grokSubscription ? ["subscription-login"] : []),
...(execution.environment.id === "daytona" ? ["sandbox-provider"] : []),
],
async setup(resolved) {
const company = value<CompanyRecord>(resolved, "company");
const secretRefs = value<SecretReferenceMap>(resolved, "secrets");
if (execution.environment.id === "local") {
// Paperclip has one instance-managed local environment. The company
// creation API ensures it exists; creating a second local environment
// is intentionally rejected by the public API.
const environments = await api.get<EnvironmentRecord[]>(
`/api/companies/${company.id}/environments?driver=local`,
);
const local = environments.find(
(candidate) => candidate.driver === "local",
);
if (!local)
throw new Error(
"Isolated Paperclip instance did not create its managed local environment",
);
return local;
}
return api.post<EnvironmentRecord>(
`/api/companies/${company.id}/environments`,
execution.environment.buildEnvironment({
secretRefs,
daytonaImage: input.daytonaImage,
executionId: input.executionNonce,
}),
);
},
async teardown(environment) {
if (execution.environment.id === "daytona") {
await deleteDaytonaEnvironment(api, environment.id);
}
},
});
const managedHiring = isManagedHiringCase(execution.suite.id, execution.task.id);
if (managedHiring) {
registry.register<ManagedAccountFixture>({
id: "ai-connection",
dependencies: ["company"],
async setup(resolved) {
const company = value<CompanyRecord>(resolved, "company");
const key = execution.profile.credential;
const provider = key === "ANTHROPIC_API_KEY" ? "anthropic"
: key === "OPENROUTER_API_KEY" ? "openrouter"
: key === "OPENAI_API_KEY" ? "openai" : null;
if (!provider) throw new Error(`Unsupported managed hiring credential ${key}`);
const apiKey = input.credentials[key];
if (!apiKey) throw new Error(`Missing credential ${key}`);
const account = await api.postSensitive<{ connectionId: string }>(
`/api/companies/${company.id}/ai-connections`,
{
provider,
method: "api_key",
name: `Runner E2E account ${input.executionNonce}`,
ownership: "personal",
apiKey,
agentIds: [],
allAgents: false,
},
);
return {
connectionId: account.connectionId,
binding: { provider, method: "api_key", mode: "responsible_user" },
};
},
});
}
registry.register<AgentRecord>({
id: "agent",
dependencies: [
"company",
"secrets",
"environment",
...(grokSubscription ? ["subscription-login"] : []),
...(managedHiring ? ["ai-connection"] : []),
],
async setup(resolved) {
const company = value<CompanyRecord>(resolved, "company");
const environment = value<EnvironmentRecord>(resolved, "environment");
const secretRefs = value<SecretReferenceMap>(resolved, "secrets");
const agent = execution.profile.buildAgent({
environmentId: environment.id,
environmentFixtureId: execution.environment.id,
workspacePath: input.workspacePath,
secretRefs,
executionId: input.executionNonce,
});
if (["stock-harness", "native-connection-guidance"].includes(execution.suite.id)
|| (execution.suite.id === "everyday-workflows" && ["hire-reuse", "delegate-feedback"].includes(execution.task.id))) {
agent.budgetMonthlyCents = 1_000;
}
if (managedHiring) {
const account = value<ManagedAccountFixture>(resolved, "ai-connection");
const config = agent.adapterConfig as Record<string, unknown>;
delete config.env;
agent.runtimeConfig = {
...(agent.runtimeConfig as Record<string, unknown>),
aiConnection: account.binding,
};
}
return api.post<AgentRecord>(
`/api/companies/${company.id}/agents`,
agent,
);
},
async teardown() {
// Agent state is instance-local. Daytona environment teardown below is
// the only fixture cleanup that must reach an external provider.
},
});
if (execution.task.flow === "warm_three_turn"
|| (execution.suite.id === "extended-harnesses" && execution.task.id === "file-edit-validate")) {
registry.register<ProjectRecord>({
id: "project",
dependencies: ["company", "environment"],
async setup(resolved) {
const company = value<CompanyRecord>(resolved, "company");
const environment = value<EnvironmentRecord>(resolved, "environment");
return api.post<ProjectRecord>(
`/api/companies/${company.id}/projects`,
{
name: `Runner E2E workspace project ${input.executionNonce}`,
description:
"Ephemeral project anchoring the fixture execution workspace and file copy-back",
executionWorkspacePolicy: {
enabled: true,
defaultMode: "shared_workspace",
sharedWorkspaceConcurrency: "serialize",
allowIssueOverride: false,
environmentId: environment.id,
workspaceStrategy: { type: "project_primary" },
},
workspace: {
name: "Primary",
sourceType: "local_path",
cwd: input.workspacePath,
isPrimary: true,
},
},
);
},
async teardown() {
// The isolated instance is deleted after provider resources are gone.
},
});
}
const setup = await registry.setupAll();
return {
company: value<CompanyRecord>(setup.values, "company"),
secretRefs: value<SecretReferenceMap>(setup.values, "secrets"),
environment: value<EnvironmentRecord>(setup.values, "environment"),
agent: value<AgentRecord>(setup.values, "agent"),
...(setup.values.has("project")
? { project: value<ProjectRecord>(setup.values, "project") }
: {}),
...(managedHiring
? {
aiConnection: value<ManagedAccountFixture>(
setup.values,
"ai-connection",
),
}
: {}),
teardown: setup.teardown,
};
}