mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 20:50:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need to discover connections, obtain consent, and continue from saved decisions. > - We want to reduce repeated instructions only when measured behavior supports the change. > - The existing decline tasks tell the model not to retry. One provider-decline check can pass without an explanation or an observed service counter. > - This PR adds neutral tasks and stricter saved-evidence checks before any connection instruction reduction. > - Production instructions remain unchanged. The new cells are configured, not live-qualified. ## Linked Issues or Issue Description Refs #15218. Refs #15389. **What existing behavior does this improve?** The Product E2E connection workflow evaluation and its instruction measurement provenance. **Current behavior** Some decline prompts supply the policy they intend to test. The provider-decline workflow does not require a saved post-decision explanation. Its old no-call check can use a missing fixture counter as zero. OpenCode has no connection cases in the original Everyday matrix. **Proposed behavior** Add an explicit-only suite with five connection stories on native Codex, ACPX Claude, and OpenCode. Require an explanation attributed by exact run ID after a saved decline. Observe the provider fixture counter. Preserve the original cases and grades. ## What Changed - Add fifteen configured cells with one attempt, twelve-minute deadlines, and verified 1,000-cent company and agent budget stops. - Remove procedure hints from the three new decline prompts. Keep a user-permitted explanation fallback and the existing positive controls. - Require saved decline state, one decision, unchanged connections, observed zero service calls where applicable, and a post-decision explanation from a successful run on the same task. - Add negative grader calibration and test the actual fixture budget payloads. Exclude the suite from default and generic selection. - Extend the existing full-catalog measurement source manifest with connection descriptions and schemas. Add an audit of fixed text, tool descriptions, returned instructions, and unqualified behavior. - Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and preserve its new Cursor suites. No production, credential, workflow, or lockfile change. ## Verification - Before rebase: Product E2E support passed 1,424 TypeScript tests and 128 Node checks. Six catalog measurement tests, repository typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery passed. - The full pre-rebase repository test run was stopped when master advanced. Its partial result is not a pass. - After rebase and the review correction: repository build/typecheck, Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128 Node checks, six measurement tests, and exact fifteen-cell discovery pass. The duplicate local full-suite run was stopped incomplete after about 20 minutes once complete CI passed; no local full-suite pass is claimed. - Review found that the initial grader read `runId` instead of public `createdByRunId`. A regression calibration reproduced both rejection of valid public comments and acceptance of the wrong alias. The fix uses the actual field and binds the evidence type to the shared `IssueComment` contract. A subsequent type-only import path correction passes Product E2E typecheck. - Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes [complete CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545): 51 successful checks and two intentional Storybook skips, plus separate Snyk success. Fresh Greptile review is 5/5 with the single review thread resolved and no new findings. The PR is clean and mergeable. - Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm test:e2e:runner -- --list --suite native-connection-guidance`. The measurement uses `PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json pnpm exec vitest run --project @paperclipai/server server/src/__tests__/native-procedure-measurement.test.ts`. Validation used pinned pnpm 9.15.4. - No paid provider campaign was started. There is no baseline/candidate behavior result for these new cells. - The audit records 654 UTF-8 bytes of fixed connection guidance. A clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41 supplied tools, 53,341 normalized bytes at start/resume, 50,949 at compact continuation, and a 48,195-byte authenticated OpenCode MCP catalog. These are byte counts, not tokens, bills, vendor-private prompt sizes, or savings from this PR. ## Risks - This is eval preparation. Passing support tests do not establish live model behavior or qualify an instruction reduction. - The explanation oracle checks attributed saved output. It does not prove cognition or arbitrary prose truthfulness. One saved interaction also does not prove the absence of repeated idempotent tool calls. - Successful new authentication and tool refresh, existing-connection agent grants, independent work while waiting, explicit retry after decline, and blocking when mandatory work remains still need separate coverage. - Notion setup decline does not execute a real Notion service. Positive service approval uses an already installed deterministic service; it does not qualify new connection creation. - The original historical failures remain unchanged. Future comparisons must freeze source, fixture, model, input, and grading controls and retain every actual attempt. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code execution. The exact serving model ID and context-window size were not exposed in this session; they are not inferred. No model provider was invoked by the eval suite in this PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
309 lines
13 KiB
TypeScript
309 lines
13 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import type { RunnerApi } from "./api.js";
|
|
import { runnerMatrix } from "./catalog.js";
|
|
import { setupLiveFixtures } from "./live-fixtures.js";
|
|
|
|
describe("live runner fixtures", () => {
|
|
it.each(["runner-codex", "runner-acpx-claude", "runner-opencode"].flatMap(profile =>
|
|
["hire-reuse", "delegate-feedback"].map(task => [profile, task]))) (
|
|
"applies %s %s budget stops in actual setup with a nonce and only its selected credential", async (profile, task) => {
|
|
const execution = runnerMatrix.find(row => row.id === `everyday-workflows.${profile}.local.${task}`)!;
|
|
let agentBody: any;
|
|
let companyBody: any;
|
|
const api = {
|
|
async get() { return [{ id: "local", driver: "local" }]; },
|
|
async postSensitive(url: string) { return url.endsWith("/ai-connections") ? { connectionId: "account" } : { id: "secret" }; },
|
|
async post(url: string, data: any) {
|
|
if (url === "/api/companies") { companyBody = data; return { id: "company", name: "Test" }; }
|
|
if (url.endsWith("/agents")) { agentBody = data; return { id: "lead", ...data }; }
|
|
throw new Error(`Unexpected POST ${url}`);
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "random-nonce", workspacePath: "/tmp/test",
|
|
credentials: { [execution.profile.credential]: "test-value" } });
|
|
expect(companyBody.budgetMonthlyCents).toBe(1_000);
|
|
expect(agentBody.budgetMonthlyCents).toBe(1_000);
|
|
await fixtures.teardown();
|
|
},
|
|
);
|
|
|
|
it.each(["runner-codex", "runner-acpx-claude", "runner-opencode"])(
|
|
"sets both connection-guidance budget stops for %s before execution", async profile => {
|
|
const execution = runnerMatrix.find(row => row.suite.id === "native-connection-guidance" && row.profile.id === profile)!;
|
|
let companyBudget: unknown;
|
|
let agentBudget: unknown;
|
|
const api = {
|
|
async get() { return [{ id: "local", driver: "local" }]; },
|
|
async postSensitive() { return { id: "secret" }; },
|
|
async post(url: string, data: any) {
|
|
if (url === "/api/companies") { companyBudget = data.budgetMonthlyCents; return { id: "company", name: "Test" }; }
|
|
if (url.endsWith("/agents")) { agentBudget = data.budgetMonthlyCents; return { id: "lead", ...data }; }
|
|
throw new Error("Unexpected POST " + url);
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/test",
|
|
credentials: { [execution.profile.credential]: "test-value" } });
|
|
expect(companyBudget).toBe(1_000);
|
|
expect(agentBudget).toBe(1_000);
|
|
await fixtures.teardown();
|
|
},
|
|
);
|
|
|
|
it.each(["runner-codex", "legacy-codex", "runner-acpx-claude", "legacy-opencode"])(
|
|
"creates a production-default %s hire with company and agent budget stops", async (profile) => {
|
|
const execution = runnerMatrix.find(row => row.suite.id === "stock-harness" && row.profile.id === profile)!;
|
|
let agentBody: any;
|
|
const api = {
|
|
async get() { return [{ id: "local", driver: "local" }]; },
|
|
async postSensitive() { return { id: "secret" }; },
|
|
async post(url: string, data: any) {
|
|
if (url === "/api/companies") {
|
|
expect(data.budgetMonthlyCents).toBe(1_000);
|
|
return { id: "company", name: "Fixture" };
|
|
}
|
|
if (url.endsWith("/agents")) { agentBody = data; return { id: "agent", ...data }; }
|
|
throw new Error(`Unexpected POST ${url}`);
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/fixture",
|
|
credentials: { [execution.profile.credential]: "fixture-key" } });
|
|
expect(agentBody).not.toHaveProperty("instructionsBundle");
|
|
expect(agentBody.budgetMonthlyCents).toBe(1_000);
|
|
expect(agentBody.adapterConfig.env[execution.profile.credential]).toEqual({ type: "secret_ref", secretId: "secret", version: "latest" });
|
|
await fixtures.teardown();
|
|
},
|
|
);
|
|
it("anchors extended file validation to a public project workspace for local and remote copy-back", async () => {
|
|
const execution = runnerMatrix.find(e => e.id === "extended-harnesses.runner-acpx-pi.local.file-edit-validate")!;
|
|
let projectBody: any;
|
|
const api = {
|
|
async get() { return [{ id: "local", driver: "local" }]; },
|
|
async postSensitive() { return { id: "secret" }; },
|
|
async post(url: string, data: any) {
|
|
if (url === "/api/companies") return { id: "company", name: "Test" };
|
|
if (url.endsWith("/agents")) return { id: "agent", ...data };
|
|
if (url.endsWith("/projects")) { projectBody = data; return { id: "project", ...data }; }
|
|
throw new Error(`Unexpected POST ${url}`);
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const fixtures = await setupLiveFixtures({ api, execution, executionNonce: "nonce", workspacePath: "/tmp/fixture-workspace", credentials: { OPENROUTER_API_KEY: "fixture-key" } });
|
|
expect(fixtures.project?.id).toBe("project");
|
|
expect(projectBody).toMatchObject({ executionWorkspacePolicy: { environmentId: "local", workspaceStrategy: { type: "project_primary" } }, workspace: { cwd: "/tmp/fixture-workspace", sourceType: "local_path" } });
|
|
});
|
|
|
|
it.each([
|
|
["runner-codex", "hiring-templates", "hire-coder-template-reuse"],
|
|
["runner-acpx-claude", "hiring-templates", "hire-coder-template-reuse"],
|
|
["runner-codex", "everyday-workflows", "hire-reuse"],
|
|
["runner-acpx-claude", "everyday-workflows", "hire-reuse"],
|
|
["runner-opencode", "everyday-workflows", "hire-reuse"],
|
|
["runner-codex", "agent-chat-hardening", "hire-delegate-reuse"],
|
|
["runner-acpx-claude", "agent-chat-hardening", "hire-delegate-reuse"],
|
|
])(
|
|
"gives %s %s/%s a personal managed account without env overrides",
|
|
async (profile, suite, task) => {
|
|
const execution = runnerMatrix.find(
|
|
(e) =>
|
|
e.suite.id === suite &&
|
|
e.task.id === task &&
|
|
e.profile.id === profile &&
|
|
e.environment.id === "local",
|
|
)!;
|
|
const provider =
|
|
profile === "runner-acpx-claude" ? "anthropic" : profile === "runner-opencode" ? "openrouter" : "openai";
|
|
let connected = false;
|
|
let agentBody: any;
|
|
const api = {
|
|
async post(url: string, data: any) {
|
|
if (url === "/api/companies") return { id: "company", name: "Test" };
|
|
if (url.endsWith("/agents")) {
|
|
agentBody = data;
|
|
return { id: "lead", ...data };
|
|
}
|
|
throw new Error(`Unexpected POST ${url}`);
|
|
},
|
|
async postSensitive(url: string, data: any) {
|
|
if (url.endsWith("/ai-connections")) {
|
|
expect(data).toMatchObject({
|
|
provider,
|
|
method: "api_key",
|
|
ownership: "personal",
|
|
apiKey: "test-value",
|
|
agentIds: [],
|
|
allAgents: false,
|
|
});
|
|
connected = true;
|
|
return { connectionId: "managed-account" };
|
|
}
|
|
return { id: "secret" };
|
|
},
|
|
async get() {
|
|
return [{ id: "local", driver: "local" }];
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const fixtures = await setupLiveFixtures({
|
|
api,
|
|
execution,
|
|
executionNonce: "nonce",
|
|
workspacePath: "/tmp/test",
|
|
credentials: {
|
|
[execution.profile.credential]: "test-value",
|
|
},
|
|
});
|
|
expect(connected).toBe(true);
|
|
expect(agentBody.adapterConfig.env).toBeUndefined();
|
|
expect(agentBody.runtimeConfig.aiConnection).toEqual({
|
|
provider,
|
|
method: "api_key",
|
|
mode: "responsible_user",
|
|
});
|
|
expect((fixtures as any).aiConnection.connectionId).toBe(
|
|
"managed-account",
|
|
);
|
|
},
|
|
);
|
|
|
|
it("installs the Daytona provider through the public API before creating its environment", async () => {
|
|
const calls: string[] = [];
|
|
const api = {
|
|
async post(path: string, data?: Record<string, unknown>) {
|
|
calls.push(`POST ${path}`);
|
|
if (path === "/api/plugins/install") {
|
|
expect(data).toMatchObject({ isLocalPath: true });
|
|
expect(data?.packageName).toEqual(
|
|
expect.stringContaining(
|
|
"packages/plugins/sandbox-providers/daytona",
|
|
),
|
|
);
|
|
return {
|
|
id: "plugin-daytona",
|
|
pluginKey: "paperclip.daytona-sandbox-provider",
|
|
status: "ready",
|
|
};
|
|
}
|
|
if (path === "/api/companies") {
|
|
return { id: "company-1", name: "Runner E2E" };
|
|
}
|
|
if (path.endsWith("/environments")) {
|
|
expect(calls).toContain("POST /api/plugins/install");
|
|
return { id: "environment-1", driver: "sandbox" };
|
|
}
|
|
if (path.endsWith("/agents")) {
|
|
return { id: "agent-1", name: "Agent", companyId: "company-1" };
|
|
}
|
|
throw new Error(`Unexpected POST ${path}`);
|
|
},
|
|
async postSensitive(path: string, data?: Record<string, unknown>) {
|
|
calls.push(`POST ${path}`);
|
|
return { id: `secret-${String(data?.key).toLowerCase()}` };
|
|
},
|
|
async delete(path: string) {
|
|
calls.push(`DELETE ${path}`);
|
|
},
|
|
} as unknown as RunnerApi;
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.id ===
|
|
"core-compatibility.legacy-codex.daytona.message-marker",
|
|
);
|
|
expect(execution).toBeDefined();
|
|
|
|
const fixtures = await setupLiveFixtures({
|
|
api,
|
|
execution: execution!,
|
|
executionNonce: "nonce",
|
|
workspacePath: "/tmp/workspace",
|
|
credentials: {
|
|
OPENAI_API_KEY: "openai-test-value",
|
|
DAYTONA_API_KEY: "daytona-test-value",
|
|
},
|
|
daytonaImage:
|
|
"ghcr.io/paperclip/image@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
|
});
|
|
|
|
expect(calls.indexOf("POST /api/plugins/install")).toBeLessThan(
|
|
calls.indexOf("POST /api/companies/company-1/environments"),
|
|
);
|
|
await fixtures.teardown();
|
|
expect(calls).toContain(
|
|
"DELETE /api/environments/environment-1?destroyReusableSandboxLeases=true",
|
|
);
|
|
});
|
|
|
|
it("creates a primary project workspace for reusable Daytona scope", async () => {
|
|
const calls: string[] = [];
|
|
const api = {
|
|
async post(path: string, data?: Record<string, unknown>) {
|
|
calls.push(`POST ${path}`);
|
|
if (path === "/api/plugins/install") {
|
|
return {
|
|
id: "plugin-daytona",
|
|
pluginKey: "paperclip.daytona-sandbox-provider",
|
|
status: "ready",
|
|
};
|
|
}
|
|
if (path === "/api/companies") {
|
|
return { id: "company-1", name: "Runner E2E" };
|
|
}
|
|
if (path.endsWith("/environments")) {
|
|
return { id: "environment-1", driver: "sandbox" };
|
|
}
|
|
if (path.endsWith("/agents")) {
|
|
return { id: "agent-1", name: "Agent", companyId: "company-1" };
|
|
}
|
|
if (path.endsWith("/projects")) {
|
|
expect(data).toMatchObject({
|
|
executionWorkspacePolicy: {
|
|
enabled: true,
|
|
defaultMode: "shared_workspace",
|
|
environmentId: "environment-1",
|
|
},
|
|
workspace: {
|
|
sourceType: "local_path",
|
|
cwd: "/tmp/workspace",
|
|
isPrimary: true,
|
|
},
|
|
});
|
|
return {
|
|
id: "project-1",
|
|
name: data?.name,
|
|
primaryWorkspace: {
|
|
id: "project-workspace-1",
|
|
cwd: "/tmp/workspace",
|
|
},
|
|
};
|
|
}
|
|
throw new Error(`Unexpected POST ${path}`);
|
|
},
|
|
async postSensitive(_path: string, data?: Record<string, unknown>) {
|
|
return { id: `secret-${String(data?.key).toLowerCase()}` };
|
|
},
|
|
async delete() {},
|
|
} as unknown as RunnerApi;
|
|
const execution = runnerMatrix.find(
|
|
(candidate) =>
|
|
candidate.id ===
|
|
"daytona-warm-continuity.runner-codex.daytona.warm-three-turn",
|
|
)!;
|
|
|
|
const fixtures = await setupLiveFixtures({
|
|
api,
|
|
execution,
|
|
executionNonce: "nonce",
|
|
workspacePath: "/tmp/workspace",
|
|
credentials: {
|
|
OPENAI_API_KEY: "openai-test-value",
|
|
DAYTONA_API_KEY: "daytona-test-value",
|
|
},
|
|
daytonaImage:
|
|
"ghcr.io/paperclip/image@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
|
|
});
|
|
|
|
expect(fixtures.project?.primaryWorkspace?.id).toBe("project-workspace-1");
|
|
expect(
|
|
calls.indexOf("POST /api/companies/company-1/environments"),
|
|
).toBeLessThan(calls.indexOf("POST /api/companies/company-1/projects"));
|
|
await fixtures.teardown();
|
|
});
|
|
});
|