Files
PaperClipAI/tests/runner-e2e/live-fixtures.test.ts
DottaandPaperclip caf120105c test: prepare neutral native connection guidance evals (#15407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need to discover connections, obtain consent, and continue
from saved decisions.
> - We want to reduce repeated instructions only when measured behavior
supports the change.
> - The existing decline tasks tell the model not to retry. One
provider-decline check can pass without an explanation or an observed
service counter.
> - This PR adds neutral tasks and stricter saved-evidence checks before
any connection instruction reduction.
> - Production instructions remain unchanged. The new cells are
configured, not live-qualified.

## Linked Issues or Issue Description

Refs #15218. Refs #15389.

**What existing behavior does this improve?**

The Product E2E connection workflow evaluation and its instruction
measurement provenance.

**Current behavior**

Some decline prompts supply the policy they intend to test. The
provider-decline workflow does not require a saved post-decision
explanation. Its old no-call check can use a missing fixture counter as
zero. OpenCode has no connection cases in the original Everyday matrix.

**Proposed behavior**

Add an explicit-only suite with five connection stories on native Codex,
ACPX Claude, and OpenCode. Require an explanation attributed by exact
run ID after a saved decline. Observe the provider fixture counter.
Preserve the original cases and grades.

## What Changed

- Add fifteen configured cells with one attempt, twelve-minute
deadlines, and verified 1,000-cent company and agent budget stops.
- Remove procedure hints from the three new decline prompts. Keep a
user-permitted explanation fallback and the existing positive controls.
- Require saved decline state, one decision, unchanged connections,
observed zero service calls where applicable, and a post-decision
explanation from a successful run on the same task.
- Add negative grader calibration and test the actual fixture budget
payloads. Exclude the suite from default and generic selection.
- Extend the existing full-catalog measurement source manifest with
connection descriptions and schemas. Add an audit of fixed text, tool
descriptions, returned instructions, and unqualified behavior.
- Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and
preserve its new Cursor suites. No production, credential, workflow, or
lockfile change.

## Verification

- Before rebase: Product E2E support passed 1,424 TypeScript tests and
128 Node checks. Six catalog measurement tests, repository
typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery
passed.
- The full pre-rebase repository test run was stopped when master
advanced. Its partial result is not a pass.
- After rebase and the review correction: repository build/typecheck,
Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128
Node checks, six measurement tests, and exact fifteen-cell discovery
pass. The duplicate local full-suite run was stopped incomplete after
about 20 minutes once complete CI passed; no local full-suite pass is
claimed.
- Review found that the initial grader read `runId` instead of public
`createdByRunId`. A regression calibration reproduced both rejection of
valid public comments and acceptance of the wrong alias. The fix uses
the actual field and binds the evidence type to the shared
`IssueComment` contract. A subsequent type-only import path correction
passes Product E2E typecheck.
- Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes
[complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545):
51 successful checks and two intentional Storybook skips, plus separate
Snyk success. Fresh Greptile review is 5/5 with the single review thread
resolved and no new findings. The PR is clean and mergeable.
- Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm
test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm
test:e2e:runner -- --list --suite native-connection-guidance`. The
measurement uses
`PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json
pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/native-procedure-measurement.test.ts`. Validation
used pinned pnpm 9.15.4.
- No paid provider campaign was started. There is no baseline/candidate
behavior result for these new cells.
- The audit records 654 UTF-8 bytes of fixed connection guidance. A
clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41
supplied tools, 53,341 normalized bytes at start/resume, 50,949 at
compact continuation, and a 48,195-byte authenticated OpenCode MCP
catalog. These are byte counts, not tokens, bills, vendor-private prompt
sizes, or savings from this PR.

## Risks

- This is eval preparation. Passing support tests do not establish live
model behavior or qualify an instruction reduction.
- The explanation oracle checks attributed saved output. It does not
prove cognition or arbitrary prose truthfulness. One saved interaction
also does not prove the absence of repeated idempotent tool calls.
- Successful new authentication and tool refresh, existing-connection
agent grants, independent work while waiting, explicit retry after
decline, and blocking when mandatory work remains still need separate
coverage.
- Notion setup decline does not execute a real Notion service. Positive
service approval uses an already installed deterministic service; it
does not qualify new connection creation.
- The original historical failures remain unchanged. Future comparisons
must freeze source, fixture, model, input, and grading controls and
retain every actual attempt.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code
execution. The exact serving model ID and context-window size were not
exposed in this session; they are not inferred. No model provider was
invoked by the eval suite in this PR preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:28:16 -05:00

309 lines
13 KiB
TypeScript

import { describe, expect, it } from "vitest";
import type { RunnerApi } from "./api.js";
import { runnerMatrix } from "./catalog.js";
import { setupLiveFixtures } from "./live-fixtures.js";
describe("live runner fixtures", () => {
it.each(["runner-codex", "runner-acpx-claude", "runner-opencode"].flatMap(profile =>
["hire-reuse", "delegate-feedback"].map(task => [profile, task]))) (
"applies %s %s budget stops in actual setup with a nonce and only its selected credential", async (profile, task) => {
const execution = runnerMatrix.find(row => row.id === `everyday-workflows.${profile}.local.${task}`)!;
let agentBody: any;
let companyBody: any;
const api = {
async get() { return [{ id: "local", driver: "local" }]; },
async postSensitive(url: string) { return url.endsWith("/ai-connections") ? { connectionId: "account" } : { id: "secret" }; },
async post(url: string, data: any) {
if (url === "/api/companies") { companyBody = data; return { id: "company", name: "Test" }; }
if (url.endsWith("/agents")) { agentBody = data; return { id: "lead", ...data }; }
throw new Error(`Unexpected POST ${url}`);
},
} as unknown as RunnerApi;
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "random-nonce", workspacePath: "/tmp/test",
credentials: { [execution.profile.credential]: "test-value" } });
expect(companyBody.budgetMonthlyCents).toBe(1_000);
expect(agentBody.budgetMonthlyCents).toBe(1_000);
await fixtures.teardown();
},
);
it.each(["runner-codex", "runner-acpx-claude", "runner-opencode"])(
"sets both connection-guidance budget stops for %s before execution", async profile => {
const execution = runnerMatrix.find(row => row.suite.id === "native-connection-guidance" && row.profile.id === profile)!;
let companyBudget: unknown;
let agentBudget: unknown;
const api = {
async get() { return [{ id: "local", driver: "local" }]; },
async postSensitive() { return { id: "secret" }; },
async post(url: string, data: any) {
if (url === "/api/companies") { companyBudget = data.budgetMonthlyCents; return { id: "company", name: "Test" }; }
if (url.endsWith("/agents")) { agentBudget = data.budgetMonthlyCents; return { id: "lead", ...data }; }
throw new Error("Unexpected POST " + url);
},
} as unknown as RunnerApi;
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/test",
credentials: { [execution.profile.credential]: "test-value" } });
expect(companyBudget).toBe(1_000);
expect(agentBudget).toBe(1_000);
await fixtures.teardown();
},
);
it.each(["runner-codex", "legacy-codex", "runner-acpx-claude", "legacy-opencode"])(
"creates a production-default %s hire with company and agent budget stops", async (profile) => {
const execution = runnerMatrix.find(row => row.suite.id === "stock-harness" && row.profile.id === profile)!;
let agentBody: any;
const api = {
async get() { return [{ id: "local", driver: "local" }]; },
async postSensitive() { return { id: "secret" }; },
async post(url: string, data: any) {
if (url === "/api/companies") {
expect(data.budgetMonthlyCents).toBe(1_000);
return { id: "company", name: "Fixture" };
}
if (url.endsWith("/agents")) { agentBody = data; return { id: "agent", ...data }; }
throw new Error(`Unexpected POST ${url}`);
},
} as unknown as RunnerApi;
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/fixture",
credentials: { [execution.profile.credential]: "fixture-key" } });
expect(agentBody).not.toHaveProperty("instructionsBundle");
expect(agentBody.budgetMonthlyCents).toBe(1_000);
expect(agentBody.adapterConfig.env[execution.profile.credential]).toEqual({ type: "secret_ref", secretId: "secret", version: "latest" });
await fixtures.teardown();
},
);
it("anchors extended file validation to a public project workspace for local and remote copy-back", async () => {
const execution = runnerMatrix.find(e => e.id === "extended-harnesses.runner-acpx-pi.local.file-edit-validate")!;
let projectBody: any;
const api = {
async get() { return [{ id: "local", driver: "local" }]; },
async postSensitive() { return { id: "secret" }; },
async post(url: string, data: any) {
if (url === "/api/companies") return { id: "company", name: "Test" };
if (url.endsWith("/agents")) return { id: "agent", ...data };
if (url.endsWith("/projects")) { projectBody = data; return { id: "project", ...data }; }
throw new Error(`Unexpected POST ${url}`);
},
} as unknown as RunnerApi;
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/fixture-workspace", credentials: { OPENROUTER_API_KEY: "fixture-key" } });
expect(fixtures.project?.id).toBe("project");
expect(projectBody).toMatchObject({ executionWorkspacePolicy: { environmentId: "local", workspaceStrategy: { type: "project_primary" } }, workspace: { cwd: "/tmp/fixture-workspace", sourceType: "local_path" } });
});
it.each([
["runner-codex", "hiring-templates", "hire-coder-template-reuse"],
["runner-acpx-claude", "hiring-templates", "hire-coder-template-reuse"],
["runner-codex", "everyday-workflows", "hire-reuse"],
["runner-acpx-claude", "everyday-workflows", "hire-reuse"],
["runner-opencode", "everyday-workflows", "hire-reuse"],
["runner-codex", "agent-chat-hardening", "hire-delegate-reuse"],
["runner-acpx-claude", "agent-chat-hardening", "hire-delegate-reuse"],
])(
"gives %s %s/%s a personal managed account without env overrides",
async (profile, suite, task) => {
const execution = runnerMatrix.find(
(e) =>
e.suite.id === suite &&
e.task.id === task &&
e.profile.id === profile &&
e.environment.id === "local",
)!;
const provider =
profile === "runner-acpx-claude" ? "anthropic" : profile === "runner-opencode" ? "openrouter" : "openai";
let connected = false;
let agentBody: any;
const api = {
async post(url: string, data: any) {
if (url === "/api/companies") return { id: "company", name: "Test" };
if (url.endsWith("/agents")) {
agentBody = data;
return { id: "lead", ...data };
}
throw new Error(`Unexpected POST ${url}`);
},
async postSensitive(url: string, data: any) {
if (url.endsWith("/ai-connections")) {
expect(data).toMatchObject({
provider,
method: "api_key",
ownership: "personal",
apiKey: "test-value",
agentIds: [],
allAgents: false,
});
connected = true;
return { connectionId: "managed-account" };
}
return { id: "secret" };
},
async get() {
return [{ id: "local", driver: "local" }];
},
} as unknown as RunnerApi;
const fixtures = await setupLiveFixtures({
api,
execution,
executionNonce: "nonce",
workspacePath: "/tmp/test",
credentials: {
[execution.profile.credential]: "test-value",
},
});
expect(connected).toBe(true);
expect(agentBody.adapterConfig.env).toBeUndefined();
expect(agentBody.runtimeConfig.aiConnection).toEqual({
provider,
method: "api_key",
mode: "responsible_user",
});
expect((fixtures as any).aiConnection.connectionId).toBe(
"managed-account",
);
},
);
it("installs the Daytona provider through the public API before creating its environment", async () => {
const calls: string[] = [];
const api = {
async post(path: string, data?: Record<string, unknown>) {
calls.push(`POST ${path}`);
if (path === "/api/plugins/install") {
expect(data).toMatchObject({ isLocalPath: true });
expect(data?.packageName).toEqual(
expect.stringContaining(
"packages/plugins/sandbox-providers/daytona",
),
);
return {
id: "plugin-daytona",
pluginKey: "paperclip.daytona-sandbox-provider",
status: "ready",
};
}
if (path === "/api/companies") {
return { id: "company-1", name: "Runner E2E" };
}
if (path.endsWith("/environments")) {
expect(calls).toContain("POST /api/plugins/install");
return { id: "environment-1", driver: "sandbox" };
}
if (path.endsWith("/agents")) {
return { id: "agent-1", name: "Agent", companyId: "company-1" };
}
throw new Error(`Unexpected POST ${path}`);
},
async postSensitive(path: string, data?: Record<string, unknown>) {
calls.push(`POST ${path}`);
return { id: `secret-${String(data?.key).toLowerCase()}` };
},
async delete(path: string) {
calls.push(`DELETE ${path}`);
},
} as unknown as RunnerApi;
const execution = runnerMatrix.find(
(candidate) =>
candidate.id ===
"core-compatibility.legacy-codex.daytona.message-marker",
);
expect(execution).toBeDefined();
const fixtures = await setupLiveFixtures({
api,
execution: execution!,
executionNonce: "nonce",
workspacePath: "/tmp/workspace",
credentials: {
OPENAI_API_KEY: "openai-test-value",
DAYTONA_API_KEY: "daytona-test-value",
},
daytonaImage:
"ghcr.io/paperclip/image@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
});
expect(calls.indexOf("POST /api/plugins/install")).toBeLessThan(
calls.indexOf("POST /api/companies/company-1/environments"),
);
await fixtures.teardown();
expect(calls).toContain(
"DELETE /api/environments/environment-1?destroyReusableSandboxLeases=true",
);
});
it("creates a primary project workspace for reusable Daytona scope", async () => {
const calls: string[] = [];
const api = {
async post(path: string, data?: Record<string, unknown>) {
calls.push(`POST ${path}`);
if (path === "/api/plugins/install") {
return {
id: "plugin-daytona",
pluginKey: "paperclip.daytona-sandbox-provider",
status: "ready",
};
}
if (path === "/api/companies") {
return { id: "company-1", name: "Runner E2E" };
}
if (path.endsWith("/environments")) {
return { id: "environment-1", driver: "sandbox" };
}
if (path.endsWith("/agents")) {
return { id: "agent-1", name: "Agent", companyId: "company-1" };
}
if (path.endsWith("/projects")) {
expect(data).toMatchObject({
executionWorkspacePolicy: {
enabled: true,
defaultMode: "shared_workspace",
environmentId: "environment-1",
},
workspace: {
sourceType: "local_path",
cwd: "/tmp/workspace",
isPrimary: true,
},
});
return {
id: "project-1",
name: data?.name,
primaryWorkspace: {
id: "project-workspace-1",
cwd: "/tmp/workspace",
},
};
}
throw new Error(`Unexpected POST ${path}`);
},
async postSensitive(_path: string, data?: Record<string, unknown>) {
return { id: `secret-${String(data?.key).toLowerCase()}` };
},
async delete() {},
} as unknown as RunnerApi;
const execution = runnerMatrix.find(
(candidate) =>
candidate.id ===
"daytona-warm-continuity.runner-codex.daytona.warm-three-turn",
)!;
const fixtures = await setupLiveFixtures({
api,
execution,
executionNonce: "nonce",
workspacePath: "/tmp/workspace",
credentials: {
OPENAI_API_KEY: "openai-test-value",
DAYTONA_API_KEY: "daytona-test-value",
},
daytonaImage:
"ghcr.io/paperclip/image@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
});
expect(fixtures.project?.primaryWorkspace?.id).toBe("project-workspace-1");
expect(
calls.indexOf("POST /api/companies/company-1/environments"),
).toBeLessThan(calls.indexOf("POST /api/companies/company-1/projects"));
await fixtures.teardown();
});
});