mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need to discover connections, obtain consent, and continue from saved decisions. > - We want to reduce repeated instructions only when measured behavior supports the change. > - The existing decline tasks tell the model not to retry. One provider-decline check can pass without an explanation or an observed service counter. > - This PR adds neutral tasks and stricter saved-evidence checks before any connection instruction reduction. > - Production instructions remain unchanged. The new cells are configured, not live-qualified. ## Linked Issues or Issue Description Refs #15218. Refs #15389. **What existing behavior does this improve?** The Product E2E connection workflow evaluation and its instruction measurement provenance. **Current behavior** Some decline prompts supply the policy they intend to test. The provider-decline workflow does not require a saved post-decision explanation. Its old no-call check can use a missing fixture counter as zero. OpenCode has no connection cases in the original Everyday matrix. **Proposed behavior** Add an explicit-only suite with five connection stories on native Codex, ACPX Claude, and OpenCode. Require an explanation attributed by exact run ID after a saved decline. Observe the provider fixture counter. Preserve the original cases and grades. ## What Changed - Add fifteen configured cells with one attempt, twelve-minute deadlines, and verified 1,000-cent company and agent budget stops. - Remove procedure hints from the three new decline prompts. Keep a user-permitted explanation fallback and the existing positive controls. - Require saved decline state, one decision, unchanged connections, observed zero service calls where applicable, and a post-decision explanation from a successful run on the same task. - Add negative grader calibration and test the actual fixture budget payloads. Exclude the suite from default and generic selection. - Extend the existing full-catalog measurement source manifest with connection descriptions and schemas. Add an audit of fixed text, tool descriptions, returned instructions, and unqualified behavior. - Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and preserve its new Cursor suites. No production, credential, workflow, or lockfile change. ## Verification - Before rebase: Product E2E support passed 1,424 TypeScript tests and 128 Node checks. Six catalog measurement tests, repository typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery passed. - The full pre-rebase repository test run was stopped when master advanced. Its partial result is not a pass. - After rebase and the review correction: repository build/typecheck, Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128 Node checks, six measurement tests, and exact fifteen-cell discovery pass. The duplicate local full-suite run was stopped incomplete after about 20 minutes once complete CI passed; no local full-suite pass is claimed. - Review found that the initial grader read `runId` instead of public `createdByRunId`. A regression calibration reproduced both rejection of valid public comments and acceptance of the wrong alias. The fix uses the actual field and binds the evidence type to the shared `IssueComment` contract. A subsequent type-only import path correction passes Product E2E typecheck. - Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes [complete CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545): 51 successful checks and two intentional Storybook skips, plus separate Snyk success. Fresh Greptile review is 5/5 with the single review thread resolved and no new findings. The PR is clean and mergeable. - Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm test:e2e:runner -- --list --suite native-connection-guidance`. The measurement uses `PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json pnpm exec vitest run --project @paperclipai/server server/src/__tests__/native-procedure-measurement.test.ts`. Validation used pinned pnpm 9.15.4. - No paid provider campaign was started. There is no baseline/candidate behavior result for these new cells. - The audit records 654 UTF-8 bytes of fixed connection guidance. A clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41 supplied tools, 53,341 normalized bytes at start/resume, 50,949 at compact continuation, and a 48,195-byte authenticated OpenCode MCP catalog. These are byte counts, not tokens, bills, vendor-private prompt sizes, or savings from this PR. ## Risks - This is eval preparation. Passing support tests do not establish live model behavior or qualify an instruction reduction. - The explanation oracle checks attributed saved output. It does not prove cognition or arbitrary prose truthfulness. One saved interaction also does not prove the absence of repeated idempotent tool calls. - Successful new authentication and tool refresh, existing-connection agent grants, independent work while waiting, explicit retry after decline, and blocking when mandatory work remains still need separate coverage. - Notion setup decline does not execute a real Notion service. Positive service approval uses an already installed deterministic service; it does not qualify new connection creation. - The original historical failures remain unchanged. Future comparisons must freeze source, fixture, model, input, and grading controls and retain every actual attempt. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code execution. The exact serving model ID and context-window size were not exposed in this session; they are not inferred. No model provider was invoked by the eval suite in this PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
379 lines
13 KiB
TypeScript
379 lines
13 KiB
TypeScript
import { NATIVE_COMPLETION_BUDGET_CENTS } from "./native-completion-defaults.js";
|
|
import path from "node:path";
|
|
import { installedReleaseDaytonaPlugin } from "./installed-release.js";
|
|
import { isManagedHiringCase } from "./chat-cases.js";
|
|
import { FixtureRegistry } from "./fixture-registry.js";
|
|
import { TASK_TITLE_BUDGET_CENTS } from "./task-titles.js";
|
|
import { stageGrokSubscriptionFixture } from "./grok-subscription-fixture.js";
|
|
import type { RunnerApi } from "./api.js";
|
|
import type {
|
|
CredentialName,
|
|
MatrixExecution,
|
|
SecretReference,
|
|
SecretReferenceMap,
|
|
} from "./types.js";
|
|
|
|
interface CompanyRecord {
|
|
id: string;
|
|
issuePrefix?: string | null;
|
|
name: string;
|
|
}
|
|
|
|
interface SecretRecord {
|
|
id: string;
|
|
}
|
|
interface PluginRecord {
|
|
id: string;
|
|
pluginKey: string;
|
|
status: string;
|
|
}
|
|
interface EnvironmentRecord {
|
|
id: string;
|
|
driver: string;
|
|
config?: Record<string, unknown>;
|
|
}
|
|
interface AgentRecord {
|
|
id: string;
|
|
name: string;
|
|
companyId: string;
|
|
}
|
|
interface ManagedAccountFixture {
|
|
connectionId: string;
|
|
binding: {
|
|
provider: "openai" | "anthropic" | "openrouter";
|
|
method: "api_key";
|
|
mode: "responsible_user";
|
|
};
|
|
}
|
|
|
|
interface ProjectRecord {
|
|
id: string;
|
|
name: string;
|
|
primaryWorkspace?: {
|
|
id: string;
|
|
cwd?: string | null;
|
|
} | null;
|
|
}
|
|
|
|
export interface LiveFixtureValues {
|
|
company: CompanyRecord;
|
|
secretRefs: SecretReferenceMap;
|
|
environment: EnvironmentRecord;
|
|
agent: AgentRecord;
|
|
project?: ProjectRecord;
|
|
aiConnection?: ManagedAccountFixture;
|
|
|
|
onboardingRuntime?: {
|
|
mode: "production-wizard" | "post-onboarding-runtime-switch";
|
|
originalAdapterType: string;
|
|
originalModel: string | null;
|
|
testedAdapterType: string;
|
|
};
|
|
teardown(): Promise<void>;
|
|
}
|
|
|
|
function value<T>(resolved: ReadonlyMap<string, unknown>, id: string): T {
|
|
const result = resolved.get(id);
|
|
if (!result) throw new Error(`Missing resolved fixture ${id}`);
|
|
return result as T;
|
|
}
|
|
|
|
async function deleteDaytonaEnvironment(api: RunnerApi, environmentId: string) {
|
|
const deadlineAt = Date.now() + 120_000;
|
|
let lastError: unknown;
|
|
while (Date.now() < deadlineAt) {
|
|
try {
|
|
await api.delete(
|
|
`/api/environments/${environmentId}?destroyReusableSandboxLeases=true`,
|
|
{
|
|
allowNotFound: true,
|
|
},
|
|
);
|
|
return;
|
|
} catch (error) {
|
|
lastError = error;
|
|
await new Promise((resolve) => setTimeout(resolve, 3_000));
|
|
}
|
|
}
|
|
throw new Error(
|
|
`Daytona lease cleanup failed: ${lastError instanceof Error ? lastError.message : String(lastError)}`,
|
|
);
|
|
}
|
|
|
|
export async function setupLiveFixtures(input: {
|
|
api: RunnerApi;
|
|
execution: MatrixExecution;
|
|
executionNonce: string;
|
|
workspacePath: string;
|
|
credentials: Partial<Record<CredentialName, string>>;
|
|
daytonaImage?: string;
|
|
}): Promise<LiveFixtureValues> {
|
|
const { api, execution } = input;
|
|
const registry = new FixtureRegistry();
|
|
|
|
if (execution.environment.id === "daytona") {
|
|
registry.register<PluginRecord>({
|
|
id: "sandbox-provider",
|
|
async setup() {
|
|
return api.post<PluginRecord>("/api/plugins/install", {
|
|
packageName: process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_CLI
|
|
? installedReleaseDaytonaPlugin(process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_CLI, process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_DAYTONA_PLUGIN, process.env.PAPERCLIP_RUNNER_E2E_INSTALLED_DAYTONA_PLUGIN_VERSION)
|
|
: path.resolve(
|
|
import.meta.dirname,
|
|
"../../packages/plugins/sandbox-providers/daytona",
|
|
),
|
|
isLocalPath: true,
|
|
});
|
|
},
|
|
async teardown() {
|
|
// The plugin is installed only in the isolated instance/database. The
|
|
// launcher removes that complete instance after the environment lease
|
|
// has been destroyed, so no global uninstall mutation is necessary.
|
|
},
|
|
});
|
|
}
|
|
|
|
registry.register<CompanyRecord>({
|
|
id: "company",
|
|
async setup() {
|
|
return api.post<CompanyRecord>("/api/companies", {
|
|
name: `Runner E2E ${execution.id} ${input.executionNonce}`,
|
|
description: "Ephemeral paid full-stack runner acceptance fixture",
|
|
budgetMonthlyCents: ["native-completion", "native-instruction-consolidation", "native-connection-guidance"].includes(execution.suite.id)
|
|
|| (execution.suite.id === "everyday-workflows" && ["hire-reuse", "delegate-feedback"].includes(execution.task.id)) ? NATIVE_COMPLETION_BUDGET_CENTS
|
|
: execution.suite.id === "task-titles" ? TASK_TITLE_BUDGET_CENTS
|
|
: execution.suite.id === "stock-harness" ? 1_000 : 0,
|
|
});
|
|
},
|
|
async teardown() {
|
|
// The isolated instance/database is removed by the launcher. Do not call
|
|
// company deletion here: metered runs intentionally retain cost-event
|
|
// references until that instance-wide teardown.
|
|
},
|
|
});
|
|
|
|
registry.register<SecretReferenceMap>({
|
|
id: "secrets",
|
|
dependencies: ["company"],
|
|
async setup(resolved) {
|
|
const company = value<CompanyRecord>(resolved, "company");
|
|
const refs: SecretReferenceMap = {};
|
|
for (const credentialName of execution.requiredCredentials) {
|
|
if (credentialName === "GROK_AUTH_JSON") continue;
|
|
const rawValue = input.credentials[credentialName];
|
|
if (!rawValue) throw new Error(`Missing credential ${credentialName}`);
|
|
const secret = await api.postSensitive<SecretRecord>(
|
|
`/api/companies/${company.id}/secrets`,
|
|
{
|
|
name: `Runner E2E ${credentialName} ${input.executionNonce}`,
|
|
key: credentialName,
|
|
value: rawValue,
|
|
description: `Ephemeral credential for ${execution.id}`,
|
|
},
|
|
);
|
|
refs[credentialName] = {
|
|
type: "secret_ref",
|
|
secretId: secret.id,
|
|
version: "latest",
|
|
} satisfies SecretReference;
|
|
}
|
|
return refs;
|
|
},
|
|
});
|
|
|
|
const grokSubscription = execution.profile.credential === "GROK_AUTH_JSON";
|
|
if (grokSubscription) {
|
|
registry.register<() => Promise<void>>({
|
|
id: "subscription-login",
|
|
dependencies: ["company"],
|
|
async setup(resolved) {
|
|
const raw = input.credentials.GROK_AUTH_JSON;
|
|
if (!raw) throw new Error("Missing credential GROK_AUTH_JSON");
|
|
return stageGrokSubscriptionFixture({
|
|
raw, companyId: value<CompanyRecord>(resolved, "company").id,
|
|
environment: process.env,
|
|
});
|
|
},
|
|
async teardown(remove) { await remove(); },
|
|
});
|
|
}
|
|
|
|
registry.register<EnvironmentRecord>({
|
|
id: "environment",
|
|
dependencies: [
|
|
"company",
|
|
"secrets",
|
|
...(grokSubscription ? ["subscription-login"] : []),
|
|
...(execution.environment.id === "daytona" ? ["sandbox-provider"] : []),
|
|
],
|
|
async setup(resolved) {
|
|
const company = value<CompanyRecord>(resolved, "company");
|
|
const secretRefs = value<SecretReferenceMap>(resolved, "secrets");
|
|
if (execution.environment.id === "local") {
|
|
// Paperclip has one instance-managed local environment. The company
|
|
// creation API ensures it exists; creating a second local environment
|
|
// is intentionally rejected by the public API.
|
|
const environments = await api.get<EnvironmentRecord[]>(
|
|
`/api/companies/${company.id}/environments?driver=local`,
|
|
);
|
|
const local = environments.find(
|
|
(candidate) => candidate.driver === "local",
|
|
);
|
|
if (!local)
|
|
throw new Error(
|
|
"Isolated Paperclip instance did not create its managed local environment",
|
|
);
|
|
return local;
|
|
}
|
|
return api.post<EnvironmentRecord>(
|
|
`/api/companies/${company.id}/environments`,
|
|
execution.environment.buildEnvironment({
|
|
secretRefs,
|
|
daytonaImage: input.daytonaImage,
|
|
executionId: input.executionNonce,
|
|
}),
|
|
);
|
|
},
|
|
async teardown(environment) {
|
|
if (execution.environment.id === "daytona") {
|
|
await deleteDaytonaEnvironment(api, environment.id);
|
|
}
|
|
},
|
|
});
|
|
|
|
const managedHiring = isManagedHiringCase(execution.suite.id, execution.task.id);
|
|
if (managedHiring) {
|
|
registry.register<ManagedAccountFixture>({
|
|
id: "ai-connection",
|
|
dependencies: ["company"],
|
|
async setup(resolved) {
|
|
const company = value<CompanyRecord>(resolved, "company");
|
|
const key = execution.profile.credential;
|
|
const provider = key === "ANTHROPIC_API_KEY" ? "anthropic"
|
|
: key === "OPENROUTER_API_KEY" ? "openrouter"
|
|
: key === "OPENAI_API_KEY" ? "openai" : null;
|
|
if (!provider) throw new Error(`Unsupported managed hiring credential ${key}`);
|
|
const apiKey = input.credentials[key];
|
|
if (!apiKey) throw new Error(`Missing credential ${key}`);
|
|
const account = await api.postSensitive<{ connectionId: string }>(
|
|
`/api/companies/${company.id}/ai-connections`,
|
|
{
|
|
provider,
|
|
method: "api_key",
|
|
name: `Runner E2E account ${input.executionNonce}`,
|
|
ownership: "personal",
|
|
apiKey,
|
|
agentIds: [],
|
|
allAgents: false,
|
|
},
|
|
);
|
|
return {
|
|
connectionId: account.connectionId,
|
|
binding: { provider, method: "api_key", mode: "responsible_user" },
|
|
};
|
|
},
|
|
});
|
|
}
|
|
|
|
registry.register<AgentRecord>({
|
|
id: "agent",
|
|
dependencies: [
|
|
"company",
|
|
"secrets",
|
|
"environment",
|
|
...(grokSubscription ? ["subscription-login"] : []),
|
|
...(managedHiring ? ["ai-connection"] : []),
|
|
],
|
|
async setup(resolved) {
|
|
const company = value<CompanyRecord>(resolved, "company");
|
|
const environment = value<EnvironmentRecord>(resolved, "environment");
|
|
const secretRefs = value<SecretReferenceMap>(resolved, "secrets");
|
|
const agent = execution.profile.buildAgent({
|
|
environmentId: environment.id,
|
|
environmentFixtureId: execution.environment.id,
|
|
workspacePath: input.workspacePath,
|
|
secretRefs,
|
|
executionId: input.executionNonce,
|
|
});
|
|
if (["stock-harness", "native-connection-guidance"].includes(execution.suite.id)
|
|
|| (execution.suite.id === "everyday-workflows" && ["hire-reuse", "delegate-feedback"].includes(execution.task.id))) {
|
|
agent.budgetMonthlyCents = 1_000;
|
|
}
|
|
if (managedHiring) {
|
|
const account = value<ManagedAccountFixture>(resolved, "ai-connection");
|
|
const config = agent.adapterConfig as Record<string, unknown>;
|
|
delete config.env;
|
|
agent.runtimeConfig = {
|
|
...(agent.runtimeConfig as Record<string, unknown>),
|
|
aiConnection: account.binding,
|
|
};
|
|
}
|
|
return api.post<AgentRecord>(
|
|
`/api/companies/${company.id}/agents`,
|
|
agent,
|
|
);
|
|
},
|
|
async teardown() {
|
|
// Agent state is instance-local. Daytona environment teardown below is
|
|
// the only fixture cleanup that must reach an external provider.
|
|
},
|
|
});
|
|
|
|
if (execution.task.flow === "warm_three_turn"
|
|
|| (execution.suite.id === "extended-harnesses" && execution.task.id === "file-edit-validate")) {
|
|
registry.register<ProjectRecord>({
|
|
id: "project",
|
|
dependencies: ["company", "environment"],
|
|
async setup(resolved) {
|
|
const company = value<CompanyRecord>(resolved, "company");
|
|
const environment = value<EnvironmentRecord>(resolved, "environment");
|
|
return api.post<ProjectRecord>(
|
|
`/api/companies/${company.id}/projects`,
|
|
{
|
|
name: `Runner E2E workspace project ${input.executionNonce}`,
|
|
description:
|
|
"Ephemeral project anchoring the fixture execution workspace and file copy-back",
|
|
executionWorkspacePolicy: {
|
|
enabled: true,
|
|
defaultMode: "shared_workspace",
|
|
sharedWorkspaceConcurrency: "serialize",
|
|
allowIssueOverride: false,
|
|
environmentId: environment.id,
|
|
workspaceStrategy: { type: "project_primary" },
|
|
},
|
|
workspace: {
|
|
name: "Primary",
|
|
sourceType: "local_path",
|
|
cwd: input.workspacePath,
|
|
isPrimary: true,
|
|
},
|
|
},
|
|
);
|
|
},
|
|
async teardown() {
|
|
// The isolated instance is deleted after provider resources are gone.
|
|
},
|
|
});
|
|
}
|
|
|
|
const setup = await registry.setupAll();
|
|
return {
|
|
company: value<CompanyRecord>(setup.values, "company"),
|
|
secretRefs: value<SecretReferenceMap>(setup.values, "secrets"),
|
|
environment: value<EnvironmentRecord>(setup.values, "environment"),
|
|
agent: value<AgentRecord>(setup.values, "agent"),
|
|
...(setup.values.has("project")
|
|
? { project: value<ProjectRecord>(setup.values, "project") }
|
|
: {}),
|
|
...(managedHiring
|
|
? {
|
|
aiConnection: value<ManagedAccountFixture>(
|
|
setup.values,
|
|
"ai-connection",
|
|
),
|
|
}
|
|
: {}),
|
|
teardown: setup.teardown,
|
|
};
|
|
}
|