Files
PaperClipAI/tests/runner-e2e/catalog.ts
T
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00

1727 lines
95 KiB
TypeScript

import { planTaskCases, planTaskProfile, planDefinitionDigest } from "./plan-task-cases.js";
import { nativeCompletionTasks, nativeCompletionDefinitionDigest } from "./native-completion-cases.js";
import { NATIVE_INSTRUCTION_SUITE, NATIVE_INSTRUCTION_BASE_SHA, nativeInstructionDefinitionDigest } from "./native-instruction-consolidation.js";
import { nativeCompletionProfile, NATIVE_COMPLETION_BUDGET_CENTS } from "./native-completion-defaults.js";
import { chatConfirmationTasks } from "./chat-cases.js";
import { buildConnectionSuite } from "./connection-cases.js";
import { hiringTemplateTasks, hiringTemplateProfile, hiringTemplateDefinitionDigest } from "./hiring-template-cases.js";
import { instructionPersistenceTask } from "./instruction-persistence.js";
import { apiResponseReadingTask } from "./api-response-reading.js";
import { taskTitleTasks, taskTitleDefinitionDigest, TASK_TITLE_BUDGET_CENTS } from "./task-titles.js";
import { blockerTasks, blockerProfile } from "./blocker-cases.js";
import { accountingTasks } from "./accounting-cases.js";
import { continuationTasks } from "./continuation-cases.js";
import { contextIntegrityTasks, paperclipDocumentTask } from "./context-integrity-cases.js";
import { productionDefaultHireProfile, stockHarnessSourceDigest, stockHarnessSkillSources } from "./stock-harness.js";
import { lifecycleLiveTasks, lifecycleLiveDefinitionDigest } from "./lifecycle-live-cases.js";
import { publicMcpTasks, publicMcpSetupDigest, publicMcpWorkflowDigest, publicMcpWorkerInstructions, publicMcpWorkerSkillDigest } from "./public-mcp-cases.js";
import { graderVersion as publicMcpGraderVersion } from "./public-mcp-grading.js";
import { everydayTasks, productionStoryProfile } from "./everyday-cases.js";
import { firstTaskTasks } from "./first-task-cases.js";
import { chatTasks, chatHardeningTasks, chatStoryTasks, chatQualificationTasks, chatCompletionTasks } from "./chat-cases.js";
import { createHash } from "node:crypto";
import { fileURLToPath } from "node:url";
import { createAgentSchema } from "../../packages/shared/src/validators/agent.js";
import { createEnvironmentSchema } from "../../packages/shared/src/validators/environment.js";
import { DEFAULT_CODEX_LOCAL_MODEL } from "../../packages/adapters/codex-local/src/index.js";
import { models as claudeModels } from "../../packages/adapters/claude-local/src/index.js";
import { DEFAULT_KIMI_LOCAL_MODEL } from "../../packages/adapters/kimi-local/src/index.js";
import { DEFAULT_GROK_LOCAL_MODEL } from "../../packages/adapters/grok-local/src/index.js";
import { QUALIFIED_ACPX_PROFILES } from "../../packages/paperclip-runner/src/drivers/acpx/qualified-profiles.js";
import { QUALIFIED_OPENCODE_MODEL } from "../../packages/paperclip-runner/src/drivers/opencode/opencode-server-driver.js";
import { CREDENTIAL_NAMES } from "./types.js";
import { createGitStreamingTask } from "./daytona-git-streaming.js";
import { PENDING_PROFILE_PREREQUISITES } from "./prerequisites.js";
import {
openRouterProfileId,
openRouterRankingSnapshot,
} from "./openrouter-ranking.js";
import type {
AgentFixtureBuildInput,
EnvironmentFixture,
EnvironmentFixtureBuildInput,
MatrixExecution,
RunnerProfileFixture,
RunnerTaskFixture,
RunnerSuiteFixture,
SecretReference,
} from "./types.js";
const ENVIRONMENT_IDS = ["local", "daytona"] as const;
const SELECTABLE_GROUPS = [
"legacy",
"native",
"local",
"daytona",
"warm",
"core",
"breadth",
"chat",
"onboarding",
"context-integrity",
"stock-harness",
] as const;
const SAMPLE_UUID = "11111111-1111-4111-8111-111111111111";
export function isImmutableDaytonaImage(value: string | undefined) {
return /^.+@sha256:[0-9a-f]{64}$/i.test(value ?? "");
}
function requiredSecret(
input: AgentFixtureBuildInput,
name: RunnerProfileFixture["credential"],
): SecretReference {
const value = input.secretRefs[name];
if (!value) throw new Error(`Missing fixture secret reference ${name}`);
return value;
}
function commonAgent(
input: AgentFixtureBuildInput,
fixtureId: string,
adapterType: string,
adapterConfig: Record<string, unknown>,
additionalInstructions: readonly string[] = [],
) {
return {
name: `Runner E2E ${fixtureId} ${input.executionId}`,
role: "qa",
title: "Paid full-stack runner acceptance fixture",
capabilities:
"Completes deterministic standard, planning, and ask-mode runner acceptance tasks.",
adapterType,
adapterConfig,
defaultEnvironmentId: input.environmentId,
budgetMonthlyCents: 0,
instructionsBundle: {
entryFile: "AGENTS.md",
files: {
"AGENTS.md": [
"You are running a paid Paperclip end-to-end acceptance fixture.",
"Follow the assigned task and its Paperclip work mode literally.",
"In ongoing agent chats, follow the injected production chat directive; keep the conversation available after replying. The completion and implementation instructions below apply only to ordinary execution tasks.",
"For ordinary standard and ask tasks, publish the requested visible answer and mark the task done.",
"For ordinary planning tasks, publish or revise the canonical Plan document and its revision-bound request_confirmation, then wait. Only implement after that exact plan is accepted.",
"Invoke assigned tools only through the runtime's real tool-call channel. Never print XML, DSML, JSON, or other tool-call markup as assistant text.",
"Legacy adapters must use the public Paperclip API and the injected PAPERCLIP_API_URL, PAPERCLIP_API_KEY, PAPERCLIP_TASK_ID, and PAPERCLIP_RUN_ID values for comments, documents, interactions, and status changes.",
...(adapterType === "paperclip_runner"
? []
: [
"Read the assigned paperclip skill before using the Paperclip API. Use its documented request payloads rather than guessing fields from memory.",
'For a planning task, do not inspect the OpenAPI schema. PUT /api/issues/$PAPERCLIP_TASK_ID/documents/plan with {title:"Plan",format:"markdown",body,changeSummary}; read latestRevisionId and latestRevisionNumber from that response. Then POST /api/issues/$PAPERCLIP_TASK_ID/interactions with {kind:"request_confirmation",continuationPolicy:"wake_assignee",payload:{version:1,prompt,acceptLabel:"Approve",rejectLabel:"Reject",rejectRequiresReason:true,target:{type:"issue_document",key:"plan",revisionId,revisionNumber}}}, and PATCH the issue to {status:"in_review"}. Include Authorization and X-Paperclip-Run-Id on every write.',
]),
"Never print, persist, or expose credential values, and never create unrelated work.",
...additionalInstructions,
].join("\n"),
},
},
runtimeConfig: {},
};
}
function legacyProfile(input: {
id: string;
label: string;
adapterType: "codex_local" | "claude_local" | "opencode_local" | "kimi_local" | "grok_local";
provider: string;
model: string;
credential: RunnerProfileFixture["credential"];
extraConfig?: Record<string, unknown>;
instructions?: readonly string[];
modelQualification?: RunnerProfileFixture["modelQualification"];
}): RunnerProfileFixture {
return {
...input,
modelQualification: input.modelQualification ?? {
source: "adapter_constant",
qualificationId: `${input.adapterType}:default-model`,
},
generation: "legacy",
groups: ["legacy"],
supportedEnvironments: ENVIRONMENT_IDS,
expectedRuntimeMode: "legacy",
expectedRuntimeMetadata: {
adapterType: input.adapterType,
provider: input.provider,
},
buildAgent(buildInput) {
return commonAgent(buildInput, input.id, input.adapterType, {
// Remote adapters must inherit the lease's provider-owned remoteCwd.
// A host path here would override that mapping inside the sandbox.
...(buildInput.environmentFixtureId === "local"
? { cwd: buildInput.workspacePath }
: {}),
model: input.model,
timeoutSec: buildInput.environmentFixtureId === "daytona" ? 780 : 360,
dangerouslySkipPermissions: true,
...input.extraConfig,
env: {
[input.credential]: requiredSecret(buildInput, input.credential),
},
}, input.instructions);
},
};
}
function nativeProfile(input: {
id: string;
label: string;
provider: "codex" | "opencode" | "acpx";
model: string;
credential: RunnerProfileFixture["credential"];
acpxAgent?: "claude" | "codex" | "grok" | "cursor" | "copilot" | "pi";
qualificationCandidate?: RunnerProfileFixture["qualificationCandidate"];
supportedEnvironments?: readonly (typeof ENVIRONMENT_IDS)[number][];
modelQualification?: RunnerProfileFixture["modelQualification"];
ranking?: RunnerProfileFixture["ranking"];
}): RunnerProfileFixture {
return {
...input,
adapterType: "paperclip_runner",
generation: "native",
groups: ["native"],
supportedEnvironments: input.supportedEnvironments ?? ENVIRONMENT_IDS,
expectedRuntimeMode: "native",
modelQualification: input.modelQualification ?? {
source:
input.provider === "acpx"
? "qualified_runner_profile"
: "adapter_constant",
qualificationId:
input.provider === "acpx"
? `acpx:${input.acpxAgent}`
: `${input.provider}:qualified-model`,
},
...(input.ranking ? { ranking: input.ranking } : {}),
expectedRuntimeMetadata: {
adapterType: "paperclip_runner",
provider: input.provider,
},
buildAgent(buildInput) {
const credentialRef = input.credential === "GROK_AUTH_JSON"
? null
: requiredSecret(buildInput, input.credential);
const permissionConfig =
input.provider === "codex"
? { codexPermissionMode: "never" }
: input.provider === "opencode"
? { opencodePermissionMode: "allow" }
: { acpxPermissionMode: "approve-all", acpxAgent: input.acpxAgent };
return commonAgent(buildInput, input.id, "paperclip_runner", {
provider: input.provider,
model: input.model,
...(input.qualificationCandidate ? { timeoutSec: 120 } : {}),
lifecycleMode: "per_turn",
idleTimeoutMs: 300_000,
...permissionConfig,
env: {
...(credentialRef ? { [input.credential]: credentialRef } : {}),
// Codex's supported automation credential is CODEX_API_KEY. Keep
// OPENAI_API_KEY as the operator-facing fixture secret name and bind
// the same encrypted reference to the runtime-specific alias.
...(input.provider === "codex"
? { CODEX_API_KEY: credentialRef }
: {}),
},
});
},
};
}
/** Chat acceptance exercises the shipped provider defaults, not full-auto fixtures. */
function defaultPermissionProfile(profile: RunnerProfileFixture): RunnerProfileFixture {
if (profile.generation !== "native") return profile;
return {
...profile,
buildAgent(input) {
const payload = profile.buildAgent(input);
const { codexPermissionMode: _codex, acpxPermissionMode: _acpx, ...adapterConfig } = payload.adapterConfig as Record<string, unknown>;
return { ...payload, adapterConfig };
},
};
}
const claudeLegacyModel = "claude-sonnet-4-6";
if (!claudeModels.some((model) => model.id === claudeLegacyModel)) {
throw new Error(
`Claude adapter does not expose the qualified ${claudeLegacyModel} model`,
);
}
export const runnerProfiles: readonly RunnerProfileFixture[] = [
legacyProfile({
id: "legacy-codex",
label: "Legacy Codex",
adapterType: "codex_local",
provider: "codex",
model: DEFAULT_CODEX_LOCAL_MODEL,
credential: "OPENAI_API_KEY",
// Keep this fixture on the classic adapter/CLI lane. ACP execution is
// covered independently by the native runner ACPX profiles below.
extraConfig: {
engine: "cli",
// Shell snapshots serialize inherited environment values into CODEX_HOME.
// These disposable runs carry short-lived API credentials; keep that
// optional optimization off rather than exempting leaked files from scans.
extraArgs: ["-c", "features.shell_snapshot=false"],
},
}),
legacyProfile({
id: "legacy-claude",
label: "Legacy Claude",
adapterType: "claude_local",
provider: "claude",
model: claudeLegacyModel,
credential: "ANTHROPIC_API_KEY",
extraConfig: {
engine: "cli",
// Plan fixtures need enough tool turns to read the issue, write the
// canonical Plan document, and request confirmation. Four turns caused
// the Claude CLI to terminate correctly but prematurely with
// `max_turns_exhausted` during the revision flow.
maxTurnsPerRun: 24,
},
}),
legacyProfile({
id: "legacy-opencode",
label: "Legacy OpenCode",
adapterType: "opencode_local",
provider: "opencode",
model: QUALIFIED_OPENCODE_MODEL,
credential: "OPENROUTER_API_KEY",
}),
nativeProfile({
id: "runner-codex",
label: "Runner Codex",
provider: "codex",
model: DEFAULT_CODEX_LOCAL_MODEL,
credential: "OPENAI_API_KEY",
}),
nativeProfile({
id: "runner-opencode",
label: "Runner OpenCode",
provider: "opencode",
model: QUALIFIED_OPENCODE_MODEL,
credential: "OPENROUTER_API_KEY",
}),
nativeProfile({
id: "runner-acpx-claude",
label: "Runner ACPX Claude",
provider: "acpx",
acpxAgent: "claude",
model: QUALIFIED_ACPX_PROFILES.claude.qualificationModel,
credential: "ANTHROPIC_API_KEY",
}),
nativeProfile({
id: "runner-acpx-grok",
label: "Runner Grok Build",
provider: "acpx",
acpxAgent: "grok",
model: QUALIFIED_ACPX_PROFILES.grok.qualificationModel,
credential: "XAI_API_KEY",
}),
nativeProfile({
id: "runner-acpx-codex",
label: "Runner ACPX Codex",
provider: "acpx",
acpxAgent: "codex",
model: QUALIFIED_ACPX_PROFILES.codex.qualificationModel,
credential: "OPENAI_API_KEY",
}),
] as const;
// Explicit qualification choices from authenticated model discovery, not shipped
// defaults or claims that these candidates have passed inference qualification.
export const extendedHarnessProfiles: readonly RunnerProfileFixture[] = [
nativeProfile({
id: "runner-acpx-cursor", label: "Runner Cursor (candidate)", provider: "acpx", acpxAgent: "cursor",
qualificationCandidate: "cursor", credential: "CURSOR_AUTH_TOKEN",
model: "gpt-5.6-luna[context=272k,reasoning=medium,fast=false]",
modelQualification: { source: "candidate_runner_profile", qualificationId: "cursor:2026.09.26-dd393fe:discovery-2026-09-28" },
}),
nativeProfile({
id: "runner-acpx-copilot", label: "Runner Copilot (candidate)", provider: "acpx", acpxAgent: "copilot",
qualificationCandidate: "copilot", credential: "COPILOT_GITHUB_TOKEN", model: "gpt-5.6-luna",
modelQualification: { source: "candidate_runner_profile", qualificationId: "copilot:1.0.88:discovery-2026-09-28" },
}),
nativeProfile({
id: "runner-acpx-pi", label: "Runner Pi (candidate)", provider: "acpx", acpxAgent: "pi",
qualificationCandidate: "pi", credential: "OPENROUTER_API_KEY", model: QUALIFIED_ACPX_PROFILES.pi.qualificationModel,
modelQualification: { source: "candidate_runner_profile", qualificationId: "pi:0.0.33:0.84.2:openrouter" },
}),
];
/** Narrow legacy ACP lanes used only by the explicit context-integrity matrix. */
export const legacyAcpxProfiles: readonly RunnerProfileFixture[] = [
legacyProfile({
id: "legacy-acp-codex",
label: "Legacy ACP Codex",
adapterType: "codex_local",
provider: "codex",
model: DEFAULT_CODEX_LOCAL_MODEL,
credential: "OPENAI_API_KEY",
extraConfig: { engine: "acp", mode: "oneshot" },
}),
legacyProfile({
id: "legacy-acp-claude",
label: "Legacy ACP Claude",
adapterType: "claude_local",
provider: "claude",
model: claudeLegacyModel,
credential: "ANTHROPIC_API_KEY",
extraConfig: { engine: "acp", mode: "oneshot" },
}),
] as const;
/** Explicit-only context-integrity profiles; admission is blocked until qualification is complete. */
export const pendingContextIntegrityProfiles: readonly RunnerProfileFixture[] = [
legacyProfile({
id: "legacy-kimi-cli",
label: "Legacy Kimi CLI (pending qualification)",
adapterType: "kimi_local",
provider: "kimi",
model: DEFAULT_KIMI_LOCAL_MODEL,
credential: "KIMI_MODEL_API_KEY",
extraConfig: { engine: "cli" },
}),
legacyProfile({
id: "legacy-kimi-acp",
label: "Legacy Kimi ACP (pending qualification)",
adapterType: "kimi_local",
provider: "kimi",
model: DEFAULT_KIMI_LOCAL_MODEL,
credential: "KIMI_MODEL_API_KEY",
extraConfig: { engine: "acp", mode: "oneshot" },
}),
legacyProfile({
id: "legacy-grok",
label: "Legacy Grok (pending qualification)",
adapterType: "grok_local",
provider: "grok",
model: DEFAULT_GROK_LOCAL_MODEL,
credential: "XAI_API_KEY",
}),
] as const;
export const contextIntegrityProfiles: readonly RunnerProfileFixture[] = [
...runnerProfiles.filter((profile) => ["runner-codex", "runner-acpx-claude", "runner-opencode", "legacy-codex", "legacy-claude"].includes(profile.id)),
...legacyAcpxProfiles,
...pendingContextIntegrityProfiles,
];
export const openRouterBreadthExcludedModelIds = ["xiaomi/mimo-v2.5"] as const;
export const openRouterBreadthExcludedExecutionIds = [
"openrouter-model-breadth.openrouter-deepseek-deepseek-v4-flash-0731.local.plan-approve-complete",
"openrouter-model-breadth.openrouter-tencent-hy3.local.plan-approve-complete",
] as const;
const openRouterBreadthExcludedModelIdSet = new Set<string>(
openRouterBreadthExcludedModelIds,
);
export const openRouterBreadthProfiles: readonly RunnerProfileFixture[] =
openRouterRankingSnapshot.models
.filter(
(rankedModel) => !openRouterBreadthExcludedModelIdSet.has(rankedModel.id),
)
.map((rankedModel) =>
nativeProfile({
id: openRouterProfileId(rankedModel.id),
label: `#${rankedModel.rank} ${rankedModel.name}`,
provider: "opencode",
model: `openrouter/${rankedModel.id}`,
credential: "OPENROUTER_API_KEY",
supportedEnvironments: ["local"],
modelQualification: {
source: "openrouter_rankings_snapshot",
qualificationId: `${openRouterRankingSnapshot.snapshotId}:${rankedModel.rank}`,
},
ranking: {
rank: rankedModel.rank,
canonicalModelId: rankedModel.id,
snapshotId: openRouterRankingSnapshot.snapshotId,
capturedAt: openRouterRankingSnapshot.capturedAt,
sourceUrl: openRouterRankingSnapshot.sourceUrl,
},
}),
);
function requiredDaytonaSecret(input: EnvironmentFixtureBuildInput) {
const apiKey = input.secretRefs.DAYTONA_API_KEY;
if (!apiKey)
throw new Error("Missing fixture secret reference DAYTONA_API_KEY");
return apiKey;
}
export const runnerEnvironments: readonly EnvironmentFixture[] = [
{
id: "local",
label: "Isolated local",
groups: ["local"],
driver: "local",
provider: "local",
lifecycle: {
setup: "instance_managed",
probe: "run_context_via_api",
cleanup: "instance_shutdown",
},
expectedExecutionTarget: { kind: "local" },
buildEnvironment(input) {
return {
name: `Runner E2E local ${input.executionId}`,
description: "Ephemeral local runner E2E environment",
driver: "local",
config: {},
envVars: {},
};
},
},
{
id: "daytona",
label: "Daytona sandbox",
groups: ["daytona"],
driver: "sandbox",
provider: "daytona",
credential: "DAYTONA_API_KEY",
lifecycle: {
setup: "create_via_api",
probe: "run_context_via_api",
cleanup: "delete_via_api_and_destroy_leases",
},
expectedExecutionTarget: { kind: "remote", transport: "sandbox" },
buildEnvironment(input) {
if (!isImmutableDaytonaImage(input.daytonaImage)) {
throw new Error(
"PAPERCLIP_E2E_DAYTONA_IMAGE must be an immutable image digest",
);
}
return {
name: `Runner E2E Daytona ${input.executionId}`,
description: "Ephemeral Daytona runner E2E environment",
driver: "sandbox",
config: {
provider: "daytona",
apiKey: requiredDaytonaSecret(input),
image: input.daytonaImage,
// Pin the billable resource shape so per-test runtime list-price
// estimates remain reproducible when provider defaults change.
cpu: 4,
memory: 4,
disk: 10,
reuseLease: false,
runnerLifecycleMode: "per_turn",
autoStopInterval: 5,
autoArchiveInterval: 15,
autoDeleteInterval: 60,
timeoutMs: 300_000,
livenessTimeoutMs: 30_000,
},
envVars: {},
};
},
},
] as const;
export const daytonaWarmEnvironment: EnvironmentFixture = {
id: "daytona",
configurationKey: "warm-reuse-v1",
label: "Daytona warm reusable sandbox",
groups: ["daytona", "warm"],
driver: "sandbox",
provider: "daytona",
credential: "DAYTONA_API_KEY",
lifecycle: {
setup: "create_via_api",
probe: "run_context_via_api",
cleanup: "delete_via_api_and_destroy_leases",
},
expectedExecutionTarget: { kind: "remote", transport: "sandbox" },
buildEnvironment(input) {
if (!isImmutableDaytonaImage(input.daytonaImage)) {
throw new Error(
"PAPERCLIP_E2E_DAYTONA_IMAGE must be an immutable image digest",
);
}
return {
name: `Runner E2E Daytona warm ${input.executionId}`,
description: "Ephemeral reusable Daytona runner E2E environment",
driver: "sandbox",
config: {
provider: "daytona",
apiKey: requiredDaytonaSecret(input),
image: input.daytonaImage,
cpu: 4,
memory: 4,
disk: 10,
reuseLease: true,
runnerLifecycleMode: "warm",
autoStopInterval: 5,
autoArchiveInterval: 15,
autoDeleteInterval: 60,
timeoutMs: 300_000,
livenessTimeoutMs: 30_000,
},
envVars: {},
};
},
};
export const runnerTasks: readonly RunnerTaskFixture[] = [
{
id: "message-marker",
label: "Basic response",
groups: [],
workMode: "standard",
flow: "single_turn",
expectedRunCount: 1,
attemptTimeoutMs: {
local: 8 * 60_000,
daytona: 15 * 60_000,
},
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `Runner E2E PAPERCLIP_E2E_OK_${nonce}`,
buildVisibleMarker: (nonce) => `PAPERCLIP_E2E_OK_${nonce}`,
buildPrompt: (nonce) =>
[
"Complete this task in a single run.",
`The exact marker also appears unescaped in the task title: PAPERCLIP_E2E_OK_${nonce}`,
`Your final visible task-thread response must be exactly this marker: PAPERCLIP_E2E_OK_${nonce}`,
`In a native runner, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"PAPERCLIP_E2E_OK_${nonce}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit exactly PAPERCLIP_E2E_OK_${nonce} once as the complete user-facing final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
`In a legacy runner, make exactly one public-API write containing the marker: PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"PAPERCLIP_E2E_OK_${nonce}"}. Do not POST to /comments, and do not include the marker in any other write.`,
"The visible task-thread response is asserted; hidden reasoning or provider terminal output alone does not count.",
"Use underscore characters exactly as shown and do not insert backslashes.",
"Do not create files, ask questions, start additional tasks, or include any credentials.",
].join("\n"),
buildMatchers(nonce, execution) {
return [
{ kind: "message_exact", expected: `PAPERCLIP_E2E_OK_${nonce}` },
{
kind: "message_occurrences",
expected: `PAPERCLIP_E2E_OK_${nonce}`,
count: 1,
},
{
kind: "issue_status",
expected: execution.task.expectedTerminalState.issue,
},
{
kind: "run_status",
expected: execution.task.expectedTerminalState.run,
},
{
kind: "runtime_mode",
expected: execution.profile.expectedRuntimeMode,
},
{ kind: "environment", expected: execution.environment.id },
];
},
},
{
id: "plan-revise-accept",
label: "Plan, revise, accept, implement",
groups: [],
workMode: "planning",
flow: "plan_revision_acceptance",
expectedRunCount: 3,
attemptTimeoutMs: {
local: 8 * 60_000,
daytona: 12 * 60_000,
},
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `Runner E2E plan lifecycle ${nonce}`,
buildVisibleMarker: (nonce) => `PAPERCLIP_E2E_PLAN_DONE_${nonce}`,
buildPlanMarkers: (nonce) => ({
draft: `PAPERCLIP_E2E_PLAN_DRAFT_${nonce}`,
revised: `PAPERCLIP_E2E_PLAN_REVISED_${nonce}`,
}),
buildRevisionRequest: (nonce) =>
[
"Revise this same plan; do not implement it yet.",
`Remove PAPERCLIP_E2E_PLAN_DRAFT_${nonce} and include PAPERCLIP_E2E_PLAN_REVISED_${nonce}.`,
"Change the plan from two steps to exactly three numbered steps, with verification as step 3.",
"Publish the revised canonical Plan revision and request confirmation for that new revision.",
"In a native runner, call write_document for key `plan`, then call request_human_input exactly once with interactionKind `confirmation`, targetRevisionId set to the returned latest Plan revision, and continuationPolicy `wake_assignee`; do not call paperclip_finish while waiting. After write_document succeeds, request_human_input must be your immediate next action using that call's returned latestRevisionId; do not emit assistant text, end the response or heartbeat, or stop after write_document alone.",
"In a legacy runner, first GET the current `plan` issue document, then PUT the revised Plan with `baseRevisionId` set to that response's `latestRevisionId`; after the update succeeds, create the equivalent request_confirmation targeting the newly returned `latestRevisionId` through the public Paperclip API.",
].join(" "),
buildPrompt: (nonce) =>
[
"This is a planning-mode lifecycle acceptance task.",
"First, create a small canonical Plan with exactly two numbered steps and request approval; do not implement it.",
`The initial Plan body must contain PAPERCLIP_E2E_PLAN_DRAFT_${nonce}.`,
"If the plan is sent back, revise that same Plan document according to the revision note, publish a new revision-bound confirmation, and still do not implement.",
`Before the revised Plan is accepted, do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_PLAN_DONE_${nonce} in any visible response, comment, or tool payload; refer to it only as “the terminal marker.”`,
"Only after the revised plan is accepted, implement it by posting one final visible task-thread response containing exactly " +
`PAPERCLIP_E2E_PLAN_DONE_${nonce}` +
" and mark the task Done.",
`For a native runner, remain in the requested planning collaboration mode. Call write_document for key \`plan\`, then call request_human_input exactly once with interactionKind \`confirmation\`, targetRevisionId set to the returned latest Plan revision, and continuationPolicy \`wake_assignee\`. For both the initial Plan and the revised Plan, those two tool calls form one indivisible response sequence: immediately after write_document succeeds, request_human_input must be your next action using that call's returned latestRevisionId. Do not emit assistant text, end the response or heartbeat, or stop after write_document alone before the matching confirmation request succeeds. Do not call paperclip_finish while waiting for either Plan confirmation. When an acceptance wake arrives, first call get_task_context. Treat the wake as valid only when that control-plane result is for the current task and identifies the exact revised Plan revision used as the confirmation target as accepted; otherwise do not finish and continue waiting for the matching revision-bound confirmation. After that verification succeeds, your immediate next action must be the paperclip_finish tool call. Do not call list_documents or any other tool, and do not emit any assistant text, acknowledgement, progress note, or preamble between verification and paperclip_finish. Use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal). Call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"PAPERCLIP_E2E_PLAN_DONE_${nonce}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit only PAPERCLIP_E2E_PLAN_DONE_${nonce} as the complete final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
`For a legacy runner, use the public Paperclip API. The first PUT of the \`plan\` issue document creates it. For every later PUT, first GET the current document and set \`baseRevisionId\` to its \`latestRevisionId\`; a 409 means you must GET again and retry with the new latest revision. Create a \`request_confirmation\` targeting the successful PUT response's \`latestRevisionId\` with \`continuationPolicy: wake_assignee\`, and move the issue to \`in_review\` while waiting. After the revised Plan is accepted, write PAPERCLIP_E2E_PLAN_DONE_${nonce} exactly once through one atomic issue PATCH with status \`done\` and that exact comment; do not POST a separate comment or perform a second write.`,
"Do not create files, child tasks, or unrelated work, and do not expose credentials.",
].join("\n"),
buildMatchers(nonce, execution) {
return [
{
kind: "message_exact",
expected: `PAPERCLIP_E2E_PLAN_DONE_${nonce}`,
},
{
kind: "message_occurrences",
expected: `PAPERCLIP_E2E_PLAN_DONE_${nonce}`,
count: 1,
},
{
kind: "issue_status",
expected: execution.task.expectedTerminalState.issue,
},
{
kind: "run_status",
expected: execution.task.expectedTerminalState.run,
},
{
kind: "runtime_mode",
expected: execution.profile.expectedRuntimeMode,
},
{ kind: "environment", expected: execution.environment.id },
];
},
},
{
id: "ask-question",
label: "Ask mode question",
groups: [],
workMode: "ask",
flow: "single_turn",
expectedRunCount: 1,
attemptTimeoutMs: {
local: 8 * 60_000,
daytona: 15 * 60_000,
},
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `Runner E2E ask mode ${nonce}`,
buildVisibleMarker: (nonce) => `E2E_ASK_12_${nonce}`,
buildPrompt: (nonce) =>
[
"Answer this question directly in Ask mode: what is seven plus five?",
`Your final visible task-thread answer must be exactly this complete marker, including its final suffix: E2E_ASK_12_${nonce}.`,
"Do not create or modify files, do not create a plan or additional work, and do not expose credentials.",
"After posting the direct answer, mark this task Done.",
`In a native runner, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"E2E_ASK_12_${nonce}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. For the entire run, paperclip_finish must be your only tool call: never call report_progress or any other tool before or after it. Wait for that tool call to succeed, then emit exactly E2E_ASK_12_${nonce} as the complete final response. Do not write a user-facing final response before paperclip_finish succeeds.`,
`In a legacy runner, make exactly one public-API write containing the marker: PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"E2E_ASK_12_${nonce}"}. Do not POST to /comments, do not PATCH the status separately, and do not include the marker in any other API write.`,
].join("\n"),
buildMatchers(nonce, execution) {
return [
{
kind: "message_exact",
expected: `E2E_ASK_12_${nonce}`,
},
{
kind: "message_occurrences",
expected: `E2E_ASK_12_${nonce}`,
count: 1,
},
{
kind: "issue_status",
expected: execution.task.expectedTerminalState.issue,
},
{
kind: "run_status",
expected: execution.task.expectedTerminalState.run,
},
{
kind: "runtime_mode",
expected: execution.profile.expectedRuntimeMode,
},
{ kind: "environment", expected: execution.environment.id },
];
},
},
] as const;
function terminalMatchers(
nonceMarker: string,
execution: MatrixExecution,
): readonly ReturnType<RunnerTaskFixture["buildMatchers"]>[number][] {
return [
{ kind: "message_exact", expected: nonceMarker },
{ kind: "message_occurrences", expected: nonceMarker, count: 1 },
{
kind: "issue_status",
expected: execution.task.expectedTerminalState.issue,
},
{ kind: "run_status", expected: execution.task.expectedTerminalState.run },
{ kind: "runtime_mode", expected: execution.profile.expectedRuntimeMode },
{ kind: "environment", expected: execution.environment.id },
];
}
function breadthMarker(phase: "H" | "Q_C" | "P_READY" | "P_OK", nonce: string) {
// Keep the complete visible marker comfortably below the shortest output
// fragments observed across ranked models while retaining the attempt nonce.
return `PC_${phase}_${nonce}`;
}
const structuredQuestionResumeTask = {
id: "structured-question-resume",
label: "Structured question, answer, resume",
groups: [],
workMode: "standard",
flow: "question_resume_completion",
expectedRunCount: 2,
attemptTimeoutMs: { local: 12 * 60_000, daytona: 12 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `Runner E2E structured question ${nonce}`,
buildVisibleMarker: (nonce) => `PAPERCLIP_E2E_QUESTION_DONE_${nonce}`,
buildQuestionAnswer: (nonce) => ({
optionLabel: "Cobalt",
expectedMarker: `PAPERCLIP_E2E_QUESTION_DONE_${nonce}`,
}),
buildPrompt: (nonce) =>
[
"Ask the user one structured question before completing this task.",
"The question must be required, single-select, and offer Cobalt and Amber. Do not publish a final answer or mark the task Done while it is pending.",
`Before the answer arrives, do not spell, quote, repeat, announce, or include PAPERCLIP_E2E_QUESTION_DONE_${nonce} in any visible response, comment, or tool payload; refer to it only as “the terminal marker.”`,
`In a native runner, call request_human_input exactly once with idempotencyKey \`question-${nonce}\`, interactionKind \`questions\`, title \`Verification word\`, prompt \`Choose the verification word\`, continuationPolicy \`wake_assignee\`, and payload {version:1,questions:[{id:\`verification-word\`,prompt:\`Choose the verification word.\`,selectionMode:\`single\`,required:true,options:[{id:\`cobalt\`,label:\`Cobalt\`},{id:\`amber\`,label:\`Amber\`}]}]}.`,
'In a legacy runner, derive `API_ORIGIN` exactly once with `API_ORIGIN="${PAPERCLIP_API_URL%/}"; API_ORIGIN="${API_ORIGIN%/api}"`. Build every endpoint as `$API_ORIGIN/api/...`; never append `/api` to a base that already ends in `/api`.',
`In a legacy runner, create exactly one question interaction: POST $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID/interactions once with {"kind":"ask_user_questions","idempotencyKey":"question-${nonce}","continuationPolicy":"wake_assignee","payload":{"version":1,"questions":[{"id":"verification-word","prompt":"Choose the verification word.","selectionMode":"single","required":true,"options":[{"id":"cobalt","label":"Cobalt"},{"id":"amber","label":"Amber"}]}]}} using Authorization and X-Paperclip-Run-Id. Do not create a replacement interaction if a later write fails.`,
'In a legacy runner, after that POST returns 2xx, PATCH $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID with exactly {"status":"in_review"}. Do not include `reviewInteractionId`: it only designates confirmation interactions, not `ask_user_questions`. If the PATCH fails, retry only that PATCH and never POST the interaction again.',
"In a legacy runner, after those two writes succeed, end the current response and heartbeat immediately. Do not wait, sleep, poll, or fetch the interaction; `wake_assignee` will start a new heartbeat after the user answers.",
`After the answer arrives, if it is Cobalt, publish exactly PAPERCLIP_E2E_QUESTION_DONE_${nonce} once as the complete visible response and mark the task Done.`,
`In a native runner, after the answer arrives, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"PAPERCLIP_E2E_QUESTION_DONE_${nonce}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit exactly PAPERCLIP_E2E_QUESTION_DONE_${nonce} as the complete final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
`In a legacy runner, make exactly one completion write: PATCH $API_ORIGIN/api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"PAPERCLIP_E2E_QUESTION_DONE_${nonce}"}. Do not POST a separate comment or perform a second write containing the marker.`,
"Do not create files, plans, child tasks, or unrelated work, and do not expose credentials.",
].join("\n"),
buildMatchers(nonce, execution) {
return [
{
kind: "message_exact",
expected: `PAPERCLIP_E2E_QUESTION_DONE_${nonce}`,
},
{
kind: "message_occurrences",
expected: `PAPERCLIP_E2E_QUESTION_DONE_${nonce}`,
count: 1,
},
{
kind: "issue_status",
expected: execution.task.expectedTerminalState.issue,
},
{
kind: "run_status",
expected: execution.task.expectedTerminalState.run,
},
{
kind: "runtime_mode",
expected: execution.profile.expectedRuntimeMode,
},
{ kind: "environment", expected: execution.environment.id },
];
},
} satisfies RunnerTaskFixture;
export const localIntegrityTasks: readonly RunnerTaskFixture[] = [
structuredQuestionResumeTask,
{
...structuredQuestionResumeTask,
id: "structured-question-restart-resume",
label: "Structured question, server restart, answer, resume",
restartServerBeforeQuestionAnswer: true,
buildTitle: (nonce) => `Runner E2E restart question ${nonce}`,
},
] as const;
export const openRouterBreadthTasks: readonly RunnerTaskFixture[] = [
{
id: "hello-complete",
label: "Hello and complete",
groups: [],
workMode: "standard",
flow: "single_turn",
expectedRunCount: 1,
attemptTimeoutMs: { local: 8 * 60_000, daytona: 8 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `OpenRouter breadth hello ${nonce}`,
buildVisibleMarker: (nonce) => breadthMarker("H", nonce),
buildPrompt: (nonce) =>
[
"Complete this deterministic hello task in one turn.",
"Your first response action must be the paperclip_finish tool call. Do not emit any assistant text, acknowledgement, or preamble before calling it.",
`Use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal). Call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"${breadthMarker("H", nonce)}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit exactly ${breadthMarker("H", nonce)} as the complete user-facing final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
"Do not create files, plans, interactions, or additional work.",
].join("\n"),
buildMatchers: (nonce, execution) =>
terminalMatchers(breadthMarker("H", nonce), execution),
},
{
id: "question-resume-complete",
label: "Ask, answer, resume",
groups: [],
workMode: "standard",
flow: "question_resume_completion",
expectedRunCount: 2,
attemptTimeoutMs: { local: 12 * 60_000, daytona: 12 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `OpenRouter breadth question ${nonce}`,
buildVisibleMarker: (nonce) => breadthMarker("Q_C", nonce),
buildQuestionAnswer: (nonce) => ({
optionLabel: "Cobalt",
expectedMarker: breadthMarker("Q_C", nonce),
}),
buildPrompt: (nonce) =>
[
"Ask the user one structured question before completing this task.",
`Call request_human_input exactly once with idempotencyKey \`question-${nonce}\`, interactionKind \`questions\`, title \`Verification word\`, prompt \`Choose the verification word\`, continuationPolicy \`wake_assignee\`, and payload {version:1,questions:[{id:\`verification-word\`,prompt:\`Choose the verification word.\`,selectionMode:\`single\`,required:true,options:[{id:\`cobalt\`,label:\`Cobalt\`},{id:\`amber\`,label:\`Amber\`}]}]}.`,
`Before the answer arrives, do not spell, quote, repeat, announce, or include ${breadthMarker("Q_C", nonce)} in any visible response, comment, or tool payload; refer to it only as “the terminal marker.”`,
"Do not call paperclip_finish while the question is pending.",
`After the answer arrives, if it is Cobalt, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"${breadthMarker("Q_C", nonce)}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit exactly ${breadthMarker("Q_C", nonce)} as the complete user-facing final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
"Do not create files, plans, or additional work.",
].join("\n"),
buildMatchers: (nonce, execution) =>
terminalMatchers(breadthMarker("Q_C", nonce), execution),
},
{
id: "plan-approve-complete",
label: "Plan, approve, complete",
groups: [],
workMode: "planning",
flow: "plan_approval_completion",
expectedRunCount: 2,
attemptTimeoutMs: { local: 15 * 60_000, daytona: 15 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `OpenRouter breadth plan ${nonce}`,
buildVisibleMarker: (nonce) => breadthMarker("P_OK", nonce),
buildPlanMarkers: (nonce) => ({
draft: breadthMarker("P_READY", nonce),
revised: breadthMarker("P_OK", nonce),
}),
buildPrompt: (nonce) =>
[
"Create a canonical Plan with exactly two numbered steps and request approval; do not implement before approval.",
`The Plan body must contain ${breadthMarker("P_READY", nonce)}.`,
"Call write_document for key `plan`, then call request_human_input exactly once with interactionKind `confirmation`, targetRevisionId set to the returned latest Plan revision, and continuationPolicy `wake_assignee`.",
`Before that exact Plan revision is accepted, do not spell, quote, repeat, announce, or include ${breadthMarker("P_OK", nonce)} in any visible response, comment, or tool payload; refer to it only as “the terminal marker.”`,
"Do not call paperclip_finish while confirmation is pending.",
`After that exact Plan revision is accepted, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"done",summary:"${breadthMarker("P_OK", nonce)}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},evidence:[],verification:[]}. Wait for that tool call to succeed, then emit exactly ${breadthMarker("P_OK", nonce)} as the complete user-facing final response. Do not write a user-facing final response before paperclip_finish succeeds, and do not call another tool.`,
"Do not create files, child tasks, or unrelated work.",
].join("\n"),
buildMatchers: (nonce, execution) =>
terminalMatchers(breadthMarker("P_OK", nonce), execution),
},
] as const;
const localEnvironment = runnerEnvironments.find(
(environment) => environment.id === "local",
)!;
function warmTurnMarker(turn: 1 | 2 | 3, nonce: string) {
return `PAPERCLIP_E2E_WARM_T${turn}_${nonce}`;
}
function warmWorkspaceLine(turn: 1 | 2 | 3, nonce: string) {
// Issue descriptions are rendered through the native prompt's Markdown
// boundary, which escapes underscores. Keep the byte-level workspace
// sentinel Markdown-inert so both legacy and native providers receive the
// same literal content.
return `T${turn}-${nonce}`;
}
function warmTurnInstructions(turn: 1 | 2 | 3, nonce: string) {
const file = `daytona-warm-${nonce}.txt`;
const lines = Array.from({ length: turn }, (_, index) =>
warmWorkspaceLine((index + 1) as 1 | 2 | 3, nonce),
);
const marker = warmTurnMarker(turn, nonce);
const finalTurn = turn === 3;
const reviewRequests = finalTurn ? [] : [{
kind: "review", ownerClass: "human",
summary: `Review the verified turn ${turn} workspace file and decide whether this task is ready to complete or needs another revision.`,
}];
const legacyCompletion = finalTurn
? `In a legacy runner, make exactly one public-API completion write after verification: PATCH /api/issues/$PAPERCLIP_TASK_ID with {"status":"done","comment":"${marker}"}. Include Authorization and X-Paperclip-Run-Id. Do not POST a separate comment.`
: `In a legacy runner, after verification POST exactly one request_confirmation to /api/issues/$PAPERCLIP_TASK_ID/interactions with {"kind":"request_confirmation","idempotencyKey":"daytona-warm-review-T${turn}-${nonce}","resolverPolicy":"human_only","title":"Warm continuity turn ${turn}","summary":"Review completed warm continuity turn ${turn}.","continuationPolicy":"wake_assignee","payload":{"version":1,"prompt":"Is this warm continuity task ready to complete after turn ${turn}?","acceptLabel":"Approve completion","rejectLabel":"Continue work","rejectRequiresReason":true,"allowDeclineReason":true,"supersedeOnUserComment":false,"target":{"type":"custom","key":"daytona_warm_turn_${turn}","revisionId":"${nonce}-T${turn}","label":"Warm continuity turn ${turn}"}}}. Capture the returned interaction id. Then make exactly one issue PATCH with {"status":"in_review","comment":"${marker}","reviewInteractionId":"<returned interaction id>"}. Include Authorization and X-Paperclip-Run-Id on both writes. If the issue PATCH fails, retry only that PATCH and never create another interaction. Do not POST a separate comment. After both writes succeed, end the response and heartbeat immediately; do not wait or poll because the reviewer action will start the next turn.`;
return [
`This is warm Daytona continuity turn ${turn} of 3. Work only in the current execution workspace.`,
turn === 1
? `Create ${file} with exactly this one line followed by a newline: ${lines[0]}`
: `Before changing anything, read ${file} and verify its content is exactly ${lines.slice(0, -1).join("\\n")} followed by a newline. Then append exactly ${lines.at(-1)} followed by a newline.`,
`After the write, verify ${file} contains exactly these lines, once each and in order: ${lines.join(" | ")}.`,
`In a native runner, use the current turn completion revision and criterion IDs (CURRENT_TURN_REVISION below is a placeholder, not a literal), then call paperclip_finish exactly once with {reportedWorkDisposition:"${finalTurn ? "done" : "needs_review"}",summary:"${marker}",completionClaim:{contractRevision:CURRENT_TURN_REVISION,objectiveSatisfied:true,criteria:[{criterionId:"objective",status:"satisfied",evidenceRefs:[]}],remainingWork:[]},attentionRequests:${JSON.stringify(reviewRequests)},evidence:[],verification:[{commandOrCheck:"read ${file}",status:"passed"}]}. Wait for that tool call to succeed, then emit exactly ${marker} once as the complete user-facing final response.`,
legacyCompletion,
`In a legacy runner, the PATCH comment is the complete visible response. After its 2xx response, finish silently: do not print, echo, or emit ${marker} again as assistant text.`,
`Do not include ${marker} in any other visible response or write. Do not recreate, truncate, reorder, or duplicate prior lines.`,
].join("\n");
}
function managedWarmTurnInstructions(turn: 1 | 2 | 3, nonce: string) {
return [warmTurnInstructions(turn, nonce),
"Also update your personal AGENT_HOME with ordinary filesystem tools. Do not edit the loaded AGENTS.md instructions and do not use an API to save these files.",
turn === 1
? `Create notes/warm-memory.txt containing exactly T1-${nonce} followed by a newline. Create notes/unchanged.bin with exactly 8388608 bytes, each byte equal to 93. Create notes/delete-me.txt containing temporary.`
: `Read notes/warm-memory.txt under AGENT_HOME and verify it contains exactly the prior turn lines ${Array.from({ length: turn - 1 }, (_, i) => `T${i + 1}-${nonce}`).join(" | ")}, each followed by a newline. Append exactly T${turn}-${nonce} and a newline. Verify notes/unchanged.bin still has 8388608 bytes, each equal to 93, and leave it unchanged. ${turn === 2 ? "Delete notes/delete-me.txt." : "Verify notes/delete-me.txt is absent."}`,
"Perform these personal-file edits and verification before calling paperclip_finish. Paperclip saves them at the turn boundary.",
].join("\n");
}
export const daytonaWarmContinuityTask: RunnerTaskFixture = {
id: "warm-three-turn",
label: "Warm three-turn workspace continuity",
groups: ["warm"],
workMode: "standard",
flow: "warm_three_turn",
expectedRunCount: 3,
attemptTimeoutMs: { local: 30 * 60_000, daytona: 30 * 60_000 },
turnTimeoutMs: 10 * 60_000,
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `Runner E2E warm Daytona continuity ${nonce}`,
buildVisibleMarker: (nonce) => warmTurnMarker(3, nonce),
buildPrompt: (nonce) => managedWarmTurnInstructions(1, nonce),
buildFollowupMessages: (nonce) => [
managedWarmTurnInstructions(2, nonce),
managedWarmTurnInstructions(3, nonce),
],
buildMatchers(nonce, execution) {
// Workspace persistence is the oracle for this story. Exact response text
// formatting must not mask a valid workspace, but every warm turn still
// needs one visible marker in chronological order. Surrounding provider
// prose is allowed; the occurrence and order matchers grade only markers.
const markers = ([1, 2, 3] as const).map((turn) => warmTurnMarker(turn, nonce));
return [
...markers.map((marker) => ({ kind: "message_occurrences" as const, expected: marker, count: 1 })),
{ kind: "message_ordered" as const, expected: markers },
{
kind: "file_exact",
path: `daytona-warm-${nonce}.txt`,
expected: `${([1, 2, 3] as const)
.map((turn) => warmWorkspaceLine(turn, nonce))
.join("\n")}\n`,
},
{ kind: "issue_status", expected: "done" },
{ kind: "run_status", expected: "succeeded" },
{ kind: "runtime_mode", expected: execution.profile.expectedRuntimeMode },
{ kind: "environment", expected: "daytona" },
];
},
};
export const daytonaLargeJournalTask: RunnerTaskFixture = {
...daytonaWarmContinuityTask,
id: "large-journal-three-turn",
label: "Large journal three-turn workspace continuity",
buildFollowupMessages: nonce => [warmTurnInstructions(2, nonce), warmTurnInstructions(3, nonce)],
buildTitle: (nonce) => `Runner E2E large journal continuity ${nonce}`,
buildPrompt: (nonce) => [
"First exercise ordinary execution history with 240 separate execution-tool calls. In each call, run the Python command below exactly once. Issue the calls one by one. Do not combine them into a shell loop, script, parallel wrapper, or a single tool call: each command must be a separate ordinary execution-tool invocation. Keep a count from 1 through 240. The output is synthetic fixture data and needs no analysis.",
`python3 -c 'print("journal-continuity-" + "x" * 65000)'`,
"Wait for all 240 execution-tool calls to exit successfully. Do not redirect, suppress, or pipe their stdout. Then do the workspace task below and submit its completion report.",
warmTurnInstructions(1, nonce),
].join("\n"),
};
export const daytonaGitStreamingTask = createGitStreamingTask({ ...daytonaWarmContinuityTask, buildPrompt: nonce => warmTurnInstructions(1, nonce), buildFollowupMessages: nonce => [warmTurnInstructions(2, nonce), warmTurnInstructions(3, nonce)] });
const codexContinuityProfiles = runnerProfiles.filter((profile) =>
["legacy-codex", "runner-codex"].includes(profile.id),
);
// The journal stress fixture keeps a fixed external bundle as a control.
// Ordinary warm continuity exercises managed files and incremental checkpoints.
const warmCodexContinuityProfiles = codexContinuityProfiles.map((profile) =>
profile.generation !== "native" ? profile : {
...profile,
buildAgent(input: AgentFixtureBuildInput) {
const agent = profile.buildAgent(input);
return { ...agent, adapterConfig: {
...(agent.adapterConfig as Record<string, unknown>),
instructionsBundleMode: "external",
instructionsRootPath: fileURLToPath(new URL("./fixtures/warm-continuity/", import.meta.url)),
instructionsEntryFile: "AGENTS.md",
instructionsFilePath: fileURLToPath(new URL("./fixtures/warm-continuity/AGENTS.md", import.meta.url)),
} };
},
},
);
export const connectionReviewSuite: RunnerSuiteFixture = {
id: "connection-reviews",
label: "Governed Connection Reviews",
description: "Opt-in local review, decision, stored execution, and continuation acceptance.",
groups: ["local"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude", "legacy-codex", "legacy-claude"].includes(profile.id)),
environments: [localEnvironment],
expectedMatrixSize: 16,
tasks: (["approve", "decline", "always", "restart"] as const).map(decision => ({
id: `tool-review-${decision}`, label: `Connection review: ${decision}`, groups: [],
workMode: "standard", flow: "governed_tool_review", toolReviewDecision: decision,
expectedRunCount: 2, attemptTimeoutMs: { local: 12 * 60_000, daytona: 12 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: nonce => `Runner E2E connection review ${decision} ${nonce}`,
buildVisibleMarker: nonce => `PAPERCLIP_E2E_REVIEW_DONE_${nonce}`,
buildPrompt: nonce => [
'Use the installed connection action "List fixture pages" with query "recent". Discover it through the runtime tool gateway. Do not simulate or invent its result.',
'It requires human approval. When approval_required is returned, do not retry and do not create another interaction. End the turn waiting for that existing server-owned review; leave this task in_review. Do not claim completion before the decision arrives.',
'Approval executes the stored call on the server. Your continuation receives its result; do not repeat that approved call. A decline executes nothing and must be respected.',
...(decision === "always" ? ['After the approved result arrives, call the same action once more with query "changed arguments" to verify the remembered permission.'] : []),
`After consuming the real result (or the human decline), finish the task and publish exactly PAPERCLIP_E2E_REVIEW_DONE_${nonce} once as your final task response.`,
'For native runners use paperclip_finish. Legacy runners PATCH the task with status done and the final comment. Do not create unrelated work.',
].join("\n"),
buildMatchers: (nonce, execution) => terminalMatchers(`PAPERCLIP_E2E_REVIEW_DONE_${nonce}`, execution),
})),
};
const everydayProfiles = [
...runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id)),
nativeProfile({ id: "runner-codex-mini", label: "Runner Codex Mini", provider: "codex", model: "gpt-5.4-mini", modelQualification: {source:"qualified_runner_profile",qualificationId:"everyday-codex-mini-pilot"}, credential: "OPENAI_API_KEY", supportedEnvironments: ["local"] }),
].map(productionStoryProfile);
export const extendedHarnessFileTask: RunnerTaskFixture = {
id: "file-edit-validate", label: "Edit a file and validate its contents", groups: [],
workMode: "standard", flow: "single_turn", expectedRunCount: 1,
attemptTimeoutMs: { local: 180_000, daytona: 300_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: nonce => `Extended harness file validation ${nonce}`,
buildVisibleMarker: nonce => `EXTENDED-FILE-${nonce}`,
buildPrompt: nonce => [
`Create extended-${nonce}.txt in the current execution workspace with exactly ready-${nonce} followed by a newline.`,
`Edit that file to contain exactly verified-${nonce} followed by a newline, then run a command that reads it and fails unless those exact bytes match.`,
`After successful validation, call paperclip_finish with reportedWorkDisposition done, summary EXTENDED-FILE-${nonce}, the current completion contract revision and satisfied objective criterion, no remaining work, and the actual validation command and result.`,
`Wait for paperclip_finish to succeed, then emit exactly EXTENDED-FILE-${nonce} as your final response. Do not create unrelated files or work.`,
].join("\n"),
buildMatchers: (nonce, execution) => [
...terminalMatchers(`EXTENDED-FILE-${nonce}`, execution),
{ kind: "file_exact", path: `extended-${nonce}.txt`, expected: `verified-${nonce}\n` },
],
};
export const runnerSuites: readonly RunnerSuiteFixture[] = [
{
id: "public-mcp", label: "Paperclip through an assistant", manualOnly: true,
description: "Paid assistant tool use plus actual team execution, browser OAuth consent, durable outcomes and authorization boundaries.",
groups: ["local"], environments: [localEnvironment], tasks: publicMcpTasks, expectedMatrixSize: 39,
profiles: [
...["mini"].map(size => legacyProfile({
id: `assistant-codex-${size}`, label: `Assistant + Codex ${size}`, adapterType: "codex_local", provider: "codex",
model: `gpt-5.4-${size}`, credential: "OPENAI_API_KEY",
instructions: publicMcpWorkerInstructions,
modelQualification: { source: "qualified_runner_profile", qualificationId: `public-mcp-gpt-5.4-${size}-pilot` },
extraConfig: { engine: "cli", extraArgs: ["-c", "features.shell_snapshot=false", "-c", "features.apps=false", "-c", "features.plugins=false"] },
})),
...[["haiku", "claude-haiku-4-5-20251001"], ["sonnet", "claude-sonnet-4-6"]].map(([name, model]) => legacyProfile({
id: `assistant-claude-${name}`, label: `Assistant + Claude ${name}`, adapterType: "claude_local", provider: "claude",
model: model!, credential: "ANTHROPIC_API_KEY",
instructions: publicMcpWorkerInstructions,
modelQualification: { source: "qualified_runner_profile", qualificationId: `public-mcp-${name}-pilot` },
extraConfig: { engine: "cli", maxTurnsPerRun: 16 },
})),
],
definitionMetadata: { version: 9, setupDigest: publicMcpSetupDigest, grader: publicMcpGraderVersion, workflowDigest: publicMcpWorkflowDigest, workerSkillDigest: publicMcpWorkerSkillDigest, workerInstructions: publicMcpWorkerInstructions, assistantMaxRequests: 16, assistantMaxEstimatedUsd: 2, instructions: "shipped-plugin-skills", scheduling: "explicit-only" },
},
{
id: "plan-task-guidance", label: "Planning guidance utility", manualOnly: true,
description: "Current, short and disabled planning skills on cohesive, parallel, dependent and independently reviewed work.",
groups: ["native"], profiles: runnerProfiles.filter(p => p.id === "runner-codex").map(planTaskProfile),
environments: [localEnvironment], tasks: planTaskCases, expectedMatrixSize: 12,
definitionMetadata: { version: 1, definitionDigest: planDefinitionDigest, scheduling: "explicit-only", comparison: "public-skill-availability-ablation", scope: "one-profile-single-attempt" },
},
{
id: "blocker-guidance", label: "Direct blocker handling", manualOnly: true,
description: "Human authority, hiring permissions, and requester scope under the production coordination skill.",
groups: ["legacy"], profiles: runnerProfiles.filter(p => ["legacy-codex", "legacy-claude"].includes(p.id)).map(blockerProfile),
environments: [localEnvironment], tasks: blockerTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 1, instructions: "production-coordination-skill", grading: "saved-human-decision-ownership-and-resume", scheduling: "explicit-only" },
},
{
id: "extended-harnesses", label: "Extended ACP harnesses", manualOnly: true,
description: "Explicit candidate qualification through real Paperclip tools, browser interactions, file edits and restart recovery.",
groups: ["native"], profiles: extendedHarnessProfiles, environments: runnerEnvironments,
tasks: [...openRouterBreadthTasks, localIntegrityTasks[1]!, extendedHarnessFileTask],
expectedMatrixSize: 30,
definitionMetadata: { version: 1, qualification: "pending", scheduling: "explicit-only", admission: "host-exact-candidate-and-model", authenticatedDiscoveryDate: "2026-09-28" },
},
{
id: "instruction-persistence", label: "Instruction Persistence",
description: "Agent-owned text and binary files round trip through the editor, survive a server restart and fresh task, and synchronize concurrent edits per file with last-sync-wins.",
groups: [], profiles: codexContinuityProfiles,
environments: [localEnvironment, runnerEnvironments.find(environment => environment.id === "daytona")!], tasks: [instructionPersistenceTask],
excludedExecutionIds: ["instruction-persistence.legacy-codex.daytona.private-copy-persists"],
expectedMatrixSize: 3, manualOnly: true,
definitionMetadata: { version: 6, oracle: "current-directory-independent-nonce-binary-last-sync-wins-and-nonblocking-storage-quota", providerTurns: 6, restart: "between-tasks", instructions: "production" },
},
{
id: "grok-subscription-qualification", label: "Grok Build Subscription Qualification", manualOnly: true,
description: "Explicit company subscription login across Grok browser workflows in local and Daytona environments.",
groups: ["native"],
profiles: [nativeProfile({
id: "runner-acpx-grok-subscription", label: "Grok Build Subscription",
provider: "acpx", acpxAgent: "grok",
model: QUALIFIED_ACPX_PROFILES.grok.qualificationModel,
credential: "GROK_AUTH_JSON",
})],
environments: [localEnvironment, daytonaWarmEnvironment],
tasks: [
...runnerTasks, ...localIntegrityTasks,
...everydayTasks.filter(task => task.id === "build-revise"),
...chatHardeningTasks.filter(task => ["stop-new-resume", "continuity-restart"].includes(task.id)),
], expectedMatrixSize: 16,
definitionMetadata: { version: 1, authentication: "company-subscription", binary: "1.0.13", model: "grok-4.7", scheduling: "explicit-only", repetitionsRequired: 3, artifactOracle: "independent-python-contract", stopBoundary: "provider-turn-started" },
},
{
id: "grok-qualification", label: "Grok Build Qualification", manualOnly: true,
description: "Grok replies, planning approval, questions, downloadable artifacts, stop/resume and controller restart in local and Daytona environments.",
groups: ["native"],
profiles: runnerProfiles.filter(profile => profile.id === "runner-acpx-grok"),
environments: [localEnvironment, daytonaWarmEnvironment],
tasks: [
...runnerTasks, ...localIntegrityTasks,
...everydayTasks.filter(task => task.id === "build-revise"),
...chatHardeningTasks.filter(task => ["stop-new-resume", "continuity-restart"].includes(task.id)),
], expectedMatrixSize: 16,
definitionMetadata: { version: 2, binary: "1.0.13", model: "grok-4.7", scheduling: "explicit-only", repetitionsRequired: 3, artifactOracle: "independent-python-contract", stopBoundary: "provider-turn-started" },
},
{
id: "api-response-reading", label: "Bounded API response reading", manualOnly: true,
description: "Retrieve evidence beyond a saved API preview through authorized bounded text windows.",
groups: [], environments: runnerEnvironments,
profiles: runnerProfiles.filter(profile => profile.id === "runner-codex"),
tasks: [apiResponseReadingTask], expectedMatrixSize: 2,
definitionMetadata: { version: 1, grading: "hidden-evidence-exact-copy-and-api-tool-events", scheduling: "explicit-only" },
},
{
id: "task-titles", label: "Automatic task titles", manualOnly: true,
description: "A real agent names a prompt-only task early using production guidance and preserves a supplied title.",
groups: ["native", "local"], environments: [localEnvironment],
profiles: everydayProfiles.filter(profile => profile.provider === "codex").map(profile => ({
...profile,
buildAgent(input) { return { ...profile.buildAgent(input), budgetMonthlyCents: TASK_TITLE_BUDGET_CENTS }; },
})),
tasks: taskTitleTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 1, gradingDigest: taskTitleDefinitionDigest, instructions: "production", grading: "initial-response-early-correlated-tool-and-agent-audit", providerRuns: 1, budgetMonthlyCents: TASK_TITLE_BUDGET_CENTS, scheduling: "explicit-only" },
},
{
id: "continuation-accounting", label: "Continuation accounting baseline", manualOnly: true,
description: "Structured productive steps, bounded repair, restart and late gates; comments cannot buy more attempts.",
groups: ["local"], environments: [localEnvironment], profiles: codexContinuityProfiles.map(productionStoryProfile),
tasks: accountingTasks, expectedMatrixSize: 8,
excludedExecutionIds: accountingTasks.filter(t => !t.id.includes("productive")).map(t => `continuation-accounting.runner-codex.local.${t.id}`),
definitionMetadata: { version: 4, grading: "accounting-v4-cancellation-evidence", scheduling: "explicit-only", providerTurns: "five productive, three repair, two executed plus one cancelled for Stop" },
},
{
id: "lifecycle-baseline", label: "Lifecycle authority baseline", manualOnly: true,
description: "Paired narrative probes plus real stop/resume and governed-action controls; live browser/server/database/provider execution.",
groups: ["local"], environments: [localEnvironment],
profiles: codexContinuityProfiles.map(productionStoryProfile),
tasks: [...lifecycleLiveTasks,
...chatTasks.filter(task => ["clarify-reuse", "stop-new-resume"].includes(task.id)),
...connectionReviewSuite.tasks],
expectedMatrixSize: 46,
excludedExecutionIds: ["neutral", "challenge"].map(variant => `lifecycle-baseline.runner-codex.local.lifecycle-repair-${variant}`),
definitionMetadata: { version: 4, clarificationGrading: "wrapped-information-fields-list", narrativeDigest: lifecycleLiveDefinitionDigest, grading: "durable-state-and-attributed-narrative", scheduling: "explicit-only" },
},
{
id: "continuation", label: "Task continuation",
description: "Human direction, approval boundaries, untrusted evidence, and completed actions across turns.",
groups: ["local"], environments: [localEnvironment],
profiles: runnerProfiles.filter(profile => ["legacy-codex", "legacy-claude", "runner-codex", "runner-acpx-claude"].includes(profile.id)).map(productionStoryProfile),
tasks: continuationTasks, expectedMatrixSize: 23,
excludedExecutionIds: [
...["legacy-codex", "legacy-claude"].map(profile => `continuation.${profile}.local.question-tool-documentation`),
...["legacy-codex", "legacy-claude", "runner-codex"].map(profile => `continuation.${profile}.local.provider-question-bridge`),
],
definitionMetadata: { version: 4, grading: "durable-state-and-approval-boundaries", instructions: "production" },
},
{
id: "everyday-workflows", label: "Everyday Paperclip Work", manualOnly: true,
description: "Real user requests, useful downloaded work, and durable continuation using production instructions.",
groups: ["native"], profiles: [...everydayProfiles,
productionStoryProfile(runnerProfiles.find(profile => profile.id === "runner-opencode")!),
], environments: [localEnvironment, daytonaWarmEnvironment],
tasks: everydayTasks.map(task => ["hire-reuse", "delegate-feedback"].includes(task.id)
? { ...task, automaticRetryPolicy: "single_attempt" as const } : task), expectedMatrixSize: 52,
excludedExecutionIds: [...everydayProfiles.flatMap(profile => everydayTasks
.filter(task => !["build-revise", "delegate-feedback", "recover-controller", "create-skill-studio"].includes(task.id))
.map(task => `everyday-workflows.${profile.id}.daytona.${task.id}`)),
...everydayTasks.flatMap(task => [
`everyday-workflows.runner-opencode.daytona.${task.id}`,
...(["hire-reuse", "delegate-feedback"].includes(task.id) ? [] : [`everyday-workflows.runner-opencode.local.${task.id}`]),
]),
],
definitionMetadata: { version: 6, instructions: "production", grading: "outcome-and-invariants", scheduling: "explicit-only",
procedureCases: { ids: ["hire-reuse", "delegate-feedback"], maximumAttemptsPerCell: 1, companyAndLeadBudgetCents: 1_000 } },
},
{
id: NATIVE_INSTRUCTION_SUITE, label: "Native completion instruction consolidation", manualOnly: true,
description: "Matched production-default durable-document and concrete-blocker checks for the completion constraint reduction.",
groups: ["native", "local"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude", "runner-opencode"].includes(profile.id)).map(nativeCompletionProfile),
environments: [localEnvironment], tasks: nativeCompletionTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 1, fixtureDigest: nativeInstructionDefinitionDigest(), baseSha: NATIVE_INSTRUCTION_BASE_SHA,
instructions: "production-default", maximumAttemptsPerCell: 1, automaticRetryPolicy: "single_attempt",
budgetMonthlyCents: NATIVE_COMPLETION_BUDGET_CENTS, scheduling: "explicit-only",
grading: "original-assigned-skill-and-strict-native-completion", },
},
{
id: "native-completion", label: "Native completion guidance", manualOnly: true,
description: "Production-default assigned-skill document completion and concrete whole-task blocking, with independent native result/final ordering.",
groups: ["native", "local"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude", "runner-opencode"].includes(profile.id)).map(nativeCompletionProfile),
environments: [localEnvironment], tasks: nativeCompletionTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 1, fixtureDigest: nativeCompletionDefinitionDigest(), instructions: "unchanged-master-production-default",
providerTurns: 6, maximumAttemptsPerCell: 1, automaticRetryPolicy: "single_attempt", budgetMonthlyCents: NATIVE_COMPLETION_BUDGET_CENTS,
grading: "original-assigned-skill-durable-document-plus-native-disposition-and-observable-final-order", scheduling: "explicit-only" },
},
{
id: "context-integrity",
label: "Context Integrity",
manualOnly: true,
description: "Explicit-only proof that ordered user comments and assigned skills stay bound to the current task context.",
groups: ["context-integrity", "native", "legacy"],
profiles: contextIntegrityProfiles,
environments: [localEnvironment],
tasks: contextIntegrityTasks,
expectedMatrixSize: contextIntegrityProfiles.length * contextIntegrityTasks.length,
definitionMetadata: {
version: 1,
instructions: "production",
grading: "ordered-public-context-and-explicit-skill-invocation",
scheduling: "explicit-only",
paidCalls: "one provider run per skill case; two bounded turns per comment case",
prerequisiteGate: PENDING_PROFILE_PREREQUISITES,
},
},
{
id: "stock-harness",
label: "Stock harness with Paperclip",
manualOnly: true,
description: "Production-default hires, reduced shared prompts, assigned skills, continuation, and persistent chat.",
groups: ["stock-harness", "native", "legacy"],
profiles: [...contextIntegrityProfiles.filter(profile => !pendingContextIntegrityProfiles.includes(profile)),
runnerProfiles.find(profile => profile.id === "legacy-opencode")!].map(productionDefaultHireProfile),
environments: [localEnvironment],
tasks: [...contextIntegrityTasks, chatTasks.find(task => task.id === "continuity-restart")!, paperclipDocumentTask]
.map(task => ({ ...task, automaticRetryPolicy: "single_attempt" as const })),
expectedMatrixSize: 26,
excludedExecutionIds: ["legacy-codex", "legacy-acp-codex", "legacy-acp-claude", "runner-codex", "runner-acpx-claude", "runner-opencode"]
.map(profile => `stock-harness.${profile}.local.${paperclipDocumentTask.id}`),
definitionMetadata: {
version: 2, instructions: "production-default-hire", scheduling: "explicit-only",
automaticRetryPolicy: "single_attempt", maximumAttemptsPerCell: 1,
sourceDigest: stockHarnessSourceDigest(),
operationalSkillSources: stockHarnessSkillSources(),
grading: "public-default-bundle-and-delivered-prompts-plus-independent-lifecycle-oracles",
vendorBaseEvidence: "required deterministic Codex driver/runnerd/Rust gate; task success is not vendor-base proof",
paidCalls: "one skill turn, two ordered-comment turns, three chat turns per profile",
},
},
{
id: "first-task", label: "First-task onboarding",
description: "Production onboarding, first replies, approval, and durable task execution.",
groups: ["onboarding"],
profiles: runnerProfiles.filter(profile => ["legacy-codex", "legacy-claude", "runner-codex", "runner-acpx-claude"].includes(profile.id)),
environments: [localEnvironment], tasks: firstTaskTasks, expectedMatrixSize: 52,
definitionMetadata: { version: 4, runGrading: "evidenced-nonexecution-and-refusal", credentialPersistenceCheck: false, questionChoiceMinimum: 2, nativeSetup: "post-onboarding-runtime-switch", productionInstructions: true, qualityGrading: "informational" },
},
{
id: "agent-chat", label: "Persistent Agent Chat",
description: "Task-backed conversations, session resets, and project plan handoff.",
groups: ["chat"],
profiles: runnerProfiles.filter(profile => ["legacy-codex", "legacy-claude", "runner-codex", "runner-acpx-claude"].includes(profile.id)).map(defaultPermissionProfile),
environments: [localEnvironment], tasks: chatTasks, expectedMatrixSize: 28,
excludedExecutionIds: ["legacy-codex", "legacy-claude"].flatMap(profile => ["reassign-task", "create-backlog"].map(task => `agent-chat.${profile}.local.${task}`)),
definitionMetadata: { version: 9, clarificationGrading: "wrapped-information-fields-list", resetRunsCountedSeparately: true, permissions: "production-defaults", stopBoundary: "provider-turn-started", restartMemory: "required-after-restart" },
},
{
id: "agent-chat-hardening", label: "Agent Chat Recovery and Coordination", manualOnly: true,
description: "Native chat startup cancellation, committed sends, hiring, grounded status, and remote continuity.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment, daytonaWarmEnvironment], tasks: chatHardeningTasks, expectedMatrixSize: 18,
excludedExecutionIds: ["runner-codex", "runner-acpx-claude"].flatMap(profile =>
["stop-startup-new-resume", "hire-delegate-reuse", "blocked-status-review"].map(task => `agent-chat-hardening.${profile}.daytona.${task}`)),
definitionMetadata: { version: 5, permissions: "production-defaults", instructions: "production", grading: "durable-state-and-source-evidence", scheduling: "explicit-only", restartMemory: "required-after-restart", statusEvidence: "structured-current-blocker-and-active-run-count", readOnlyState: "public-mutation-contract-and-relations", hiringReference: "neutral-document-reference-line" },
},
{
id: "hiring-templates", label: "Production Hiring Templates", manualOnly: true,
description: "Production CEO and hiring skill/reference discovery, one coder hire, independently checked JSON artifacts and worker reuse.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => hiringTemplateProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment], tasks: hiringTemplateTasks, expectedMatrixSize: 2,
definitionMetadata: { version: 1, definitionDigest: hiringTemplateDefinitionDigest, instructions: "source-revision-default-ceo", scheduling: "explicit-only",
grading: "independent-json-and-read-receipts", baselineComparison: "same-fixture-source-derived-bundles", providerTurns: 5 },
},
{
id: "agent-chat-stories", label: "Agent Chat Setup and Interruptions", manualOnly: true,
description: "Experimental settings lifecycle and user follow-ups during active native work.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment], tasks: chatStoryTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 3, setup: "configured-native-agent", permissions: "production-defaults", interruptionBoundary: "provider-file-wait-in-agent-workspace", grading: "persisted-comments-and-plan-run-attributed", scheduling: "explicit-only" },
},
{
id: "agent-chat-qualification", label: "Agent Chat Remaining Qualification", manualOnly: true,
description: "Active ownership transfer, user recovery after worker loss, and grounded answer quality.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment], tasks: chatQualificationTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 10, instructionSetup: "read-before-write-base-hash", permissions: "production-defaults", instructions: "production", crashBoundary: "verified-native-worker-pid-at-file-wait", recovery: "new-user-message-after-verified-cleanup", answerGrading: "exact-grounded-propositions-plus-separate-semantic-review", scheduling: "explicit-only" },
},
{
id: "confirmation-replies", label: "Conversational Approval Cards", manualOnly: true,
description: "Persist approval/refusal from chat before execution, preserve card clicks, clarify ambiguous proposals, and answer historical questions after moving on.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment],
tasks: [...firstTaskTasks.filter(task => ["task-reply-accept", "interview-plan-accept", "reject-no-execution", "task-card-accept"].includes(task.id)), ...chatConfirmationTasks],
expectedMatrixSize: 12,
definitionMetadata: { version: 6, instructions: "production", grading: "card-message-provenance-before-child-creation", unansweredQuestionComposer: "history-card-only-including-fresh-dismissal", completionObservation: "120-seconds", scheduling: "explicit-only" },
},
{
id: "completion-updates", label: "Delegated Completion Updates", manualOnly: true,
description: "Qualify completion delivery in onboarding and idle, busy, multiple-task, and restart Agent Chat handoffs.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment],
tasks: [...firstTaskTasks.filter(task => task.id === "interview-plan-accept"), ...chatCompletionTasks],
expectedMatrixSize: 10,
definitionMetadata: { version: 27, runGrading: "evidenced-nonexecution-and-refusal", resultNavigation: "loaded-task-header", instructionSetup: "read-before-write-base-hash", requirementEvidence: "recorded-user-comments-and-resolved-answers", busyReferenceWait: "committed-conversation-document-response-held", busyBoundary: "source-tool-in-flight-and-public-deferred-wake", workerReference: "rsvp-code-in-saved-note", judge: "completion-quality-v14-observed-rendered-result-access", judgeMaxDollarsPerRequest: 0.5, instructions: "production", correlation: "authoritative-task-facts-required-in-reply-run", restartBoundary: "done-and-source-provider-at-reference-gate", idleBoundaryTimeoutMs: 180_000, workerBriefTimeoutMs: 240_000, workerBriefWorkspace: "managed-project", workerBriefEvidence: "released-start-time", grading: "post-completion-reply-and-result-access", semanticReview: "required-separately", chatBoundary: "worker-gated-until-source-idle", observationWindowMs: 120_000, scheduling: "explicit-only" },
},
...(process.env.PAPERCLIP_RUNNER_E2E_CONNECTION_REVIEWS === "1" ? [connectionReviewSuite] : []),
{
id: "core-compatibility",
label: "Core Runner Compatibility",
description:
"Major provider, runtime generation, and execution-environment compatibility.",
groups: ["core"],
profiles: runnerProfiles,
environments: runnerEnvironments,
tasks: runnerTasks,
expectedMatrixSize: 48,
},
{
id: "local-session-integrity",
label: "Local Session Integrity",
description:
"Structured interaction and continuation qualification for every supported local profile.",
groups: ["core"],
profiles: runnerProfiles,
environments: [localEnvironment],
tasks: localIntegrityTasks,
expectedMatrixSize: 16,
},
{
id: "openrouter-model-breadth",
label: "OpenRouter Model Breadth",
description:
"Weekly-ranked tool-capable OpenRouter models through native OpenCode on isolated local workspaces.",
groups: ["breadth"],
profiles: openRouterBreadthProfiles,
environments: [localEnvironment],
tasks: openRouterBreadthTasks,
excludedExecutionIds: openRouterBreadthExcludedExecutionIds,
expectedMatrixSize: 10,
definitionMetadata: {
rankingSnapshotId: openRouterRankingSnapshot.snapshotId,
rankingContentHash: openRouterRankingSnapshot.contentHash,
rankingCapturedAt: openRouterRankingSnapshot.capturedAt,
rankingSourceUrl: openRouterRankingSnapshot.sourceUrl,
excludedModelIds: openRouterBreadthExcludedModelIds,
excludedExecutionIds: openRouterBreadthExcludedExecutionIds,
},
},
{
id: "daytona-warm-continuity",
label: "Daytona Warm Continuity",
description:
"Three browser-driven turns on one reusable Daytona sandbox for legacy and native Codex.",
groups: ["daytona", "warm"],
profiles: codexContinuityProfiles,
environments: [daytonaWarmEnvironment],
tasks: [daytonaWarmContinuityTask],
expectedMatrixSize: 2,
definitionMetadata: { version: 3, nativeInstructions: "managed-incremental", managedFileBytes: 8 * 1024 * 1024 },
},
{
id: "daytona-journal-continuity",
label: "Daytona Large Journal Continuity",
manualOnly: true,
description: "Continue the same native session after separate ordinary tool invocations and their output grow its durable journal beyond 2 MiB.",
groups: ["daytona", "warm"],
profiles: warmCodexContinuityProfiles.filter((profile) => profile.id === "runner-codex"),
environments: [daytonaWarmEnvironment],
tasks: [daytonaLargeJournalTask],
expectedMatrixSize: 1,
definitionMetadata: { version: 5, nativeInstructions: "fixed-external", journalMinimumBytes: 2 * 1024 * 1024, toolInvocations: 240, outputBytesPerInvocation: 65020, scheduling: "explicit-only" },
},
{
id: "daytona-git-streaming",
label: "Daytona Git Streaming",
manualOnly: true,
description: "Copy back 60,000 real untracked files and continue twice with a Git filename manifest above 32 MiB.",
groups: ["daytona", "warm"],
// Copyback plus the next preparation can exceed the ordinary five-minute
// idle window for this 60,000-file workload. Keep the PID oracle strict
// while explicitly retaining both runner and sandbox for the workload.
profiles: codexContinuityProfiles.filter(profile => profile.id === "runner-codex").map(profile => ({
...profile,
buildAgent(input: AgentFixtureBuildInput) {
const agent = profile.buildAgent(input);
// Managed agent-folder collection intentionally stops the provider at
// every turn. Fixed external instructions exercise retained processes.
return { ...agent, adapterConfig: {
...(agent.adapterConfig as Record<string, unknown>),
idleTimeoutMs: 1_200_000,
instructionsBundleMode: "external",
instructionsRootPath: fileURLToPath(new URL("./fixtures/git-streaming/", import.meta.url)),
instructionsEntryFile: "AGENTS.md",
instructionsFilePath: fileURLToPath(new URL("./fixtures/git-streaming/AGENTS.md", import.meta.url)),
} };
},
})),
environments: [{
...daytonaWarmEnvironment,
buildEnvironment(input: EnvironmentFixtureBuildInput) {
const environment = daytonaWarmEnvironment.buildEnvironment(input);
return { ...environment, config: { ...(environment.config as Record<string, unknown>), runnerIdleTimeoutMs: 1_200_000, autoStopInterval: 25, autoArchiveInterval: 30 } };
},
}],
tasks: [daytonaGitStreamingTask],
expectedMatrixSize: 1,
definitionMetadata: { version: 8, instructions: "fixed-external", nativeIdleTimeoutMs: 1_200_000, autoStopIntervalMinutes: 25, generatedFileCount: 60_000, filenameBytes: 39_828_890, scheduling: "explicit-only", finalization: "committed-without-active-sync-or-retry", copyback: "all-generated-file-contents-change-each-turn" },
},
buildConnectionSuite(runnerProfiles, runnerEnvironments),
] as const;
export function suiteDefinitionHash(suite: RunnerSuiteFixture) {
return createHash("sha256")
.update(
JSON.stringify({
id: suite.id,
profiles: suite.profiles.map((profile) => ({
id: profile.id,
model: profile.model,
qualification: profile.modelQualification,
})),
environments: suite.environments.map((environment) => ({
id: environment.id,
configurationKey: environment.configurationKey ?? "default",
})),
tasks: suite.tasks.map((task) => ({
id: task.id,
flow: task.flow,
expectedRunCount: task.expectedRunCount,
...(task.minimumExpectedRunCount === undefined ? {} : { minimumExpectedRunCount: task.minimumExpectedRunCount }),
restartServerBeforeQuestionAnswer:
task.restartServerBeforeQuestionAnswer ?? false,
})),
excludedExecutionIds: [...(suite.excludedExecutionIds ?? [])].sort(),
metadata: suite.definitionMetadata ?? null,
}),
)
.digest("hex");
}
export function buildRunnerMatrix(
suites: readonly RunnerSuiteFixture[] = runnerSuites,
): MatrixExecution[] {
return suites.flatMap((suite) => {
const excludedExecutionIds = new Set(suite.excludedExecutionIds ?? []);
return suite.profiles.flatMap((profile) =>
suite.environments
.filter((environment) =>
profile.supportedEnvironments.includes(environment.id),
)
.flatMap((environment) =>
suite.tasks
.map((task) => ({
id: `${suite.id}.${profile.id}.${environment.id}.${task.id}`,
suite,
suiteDefinitionHash: suiteDefinitionHash(suite),
profile,
environment,
task,
groups: [
...new Set([
...suite.groups,
...profile.groups,
...environment.groups,
...task.groups,
]),
],
requiredCredentials: [
...(task.flow === "provider_connection" ? [] : [profile.credential]),
...(task.flow !== "provider_connection" && environment.credential ? [environment.credential] : []),
],
}))
.filter((execution) => !excludedExecutionIds.has(execution.id)),
),
);
});
}
function duplicateIds(values: readonly { id: string }[]) {
const seen = new Set<string>();
return values
.map((value) => value.id)
.filter((id) => {
if (seen.has(id)) return true;
seen.add(id);
return false;
});
}
function assertNoRawSecretValues(value: unknown, label: string) {
if (typeof value === "string") {
if (/\b(?:sk-(?:proj-)?|sk-ant-)[A-Za-z0-9_-]{12,}\b/.test(value)) {
throw new Error(`${label} contains a raw secret-looking value`);
}
return;
}
if (Array.isArray(value)) {
value.forEach((entry) => assertNoRawSecretValues(entry, label));
return;
}
if (value && typeof value === "object") {
Object.entries(value).forEach(([key, entry]) => {
if (
typeof entry === "string" &&
/(?:api.?key|access.?token|credential|secret)$/i.test(key) &&
entry.trim()
) {
throw new Error(`${label} contains a raw credential at ${key}`);
}
assertNoRawSecretValues(entry, label);
});
}
}
export function validateRunnerCatalog(): MatrixExecution[] {
const connectionSuite = runnerSuites.find(suite => suite.id === "provider-connections")!;
const allProfiles = [...connectionSuite.profiles, ...extendedHarnessProfiles, ...runnerProfiles, ...legacyAcpxProfiles, ...pendingContextIntegrityProfiles, ...openRouterBreadthProfiles, ...everydayProfiles.filter(p => !runnerProfiles.some(existing => existing.id === p.id))];
const allTasks = [
...connectionSuite.tasks,
extendedHarnessFileTask,
...contextIntegrityTasks,
...blockerTasks,
...accountingTasks,
...lifecycleLiveTasks,
...continuationTasks,
...everydayTasks,
...taskTitleTasks,
...runnerTasks,
...localIntegrityTasks,
...openRouterBreadthTasks,
daytonaWarmContinuityTask,
daytonaGitStreamingTask,
instructionPersistenceTask,
daytonaLargeJournalTask,
];
for (const [label, values] of [
["suite", runnerSuites],
["profile", allProfiles],
["environment", runnerEnvironments],
["task", allTasks],
] as const) {
const duplicates = duplicateIds(values);
if (duplicates.length > 0)
throw new Error(
`Duplicate ${label} fixture ids: ${duplicates.join(", ")}`,
);
}
const selectableGroups = new Set<string>(SELECTABLE_GROUPS);
for (const fixture of [
...runnerSuites,
...allProfiles,
...runnerEnvironments,
daytonaWarmEnvironment,
...allTasks,
]) {
const unknownGroups = fixture.groups.filter(
(group) => !selectableGroups.has(group),
);
if (unknownGroups.length > 0) {
throw new Error(
`Fixture ${fixture.id} declares unknown groups: ${unknownGroups.join(", ")}`,
);
}
}
const sampleRefs = Object.fromEntries(
CREDENTIAL_NAMES.map((name, index) => [
name,
{
type: "secret_ref" as const,
secretId: `${String(index + 1).padStart(8, "0")}-1111-4111-8111-111111111111`,
version: "latest" as const,
},
]),
);
for (const environment of [...runnerEnvironments, daytonaWarmEnvironment]) {
const payload = environment.buildEnvironment({
secretRefs: sampleRefs,
daytonaImage:
"ghcr.io/paperclipai/paperclip-daytona-runner@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
executionId: "schema-validation",
});
createEnvironmentSchema.parse(payload);
assertNoRawSecretValues(payload, `environment ${environment.id}`);
}
for (const profile of [...allProfiles, ...everydayProfiles]) {
if (!CREDENTIAL_NAMES.includes(profile.credential)) {
throw new Error(
`Profile ${profile.id} declares unknown credential ${profile.credential}`,
);
}
const unsupportedEnvironmentIds = profile.supportedEnvironments.filter(
(environmentId) => !ENVIRONMENT_IDS.includes(environmentId),
);
if (unsupportedEnvironmentIds.length > 0) {
throw new Error(
`Profile ${profile.id} declares unknown environments: ${unsupportedEnvironmentIds.join(", ")}`,
);
}
const payload = profile.buildAgent({
environmentId: SAMPLE_UUID,
environmentFixtureId: "local",
workspacePath: "/tmp/paperclip-runner-e2e-schema",
secretRefs: sampleRefs,
executionId: "schema-validation",
});
createAgentSchema.parse(payload);
assertNoRawSecretValues(payload, `profile ${profile.id}`);
}
const matrix = buildRunnerMatrix();
const duplicateMatrixIds = duplicateIds(matrix);
if (duplicateMatrixIds.length > 0) {
throw new Error(
`Duplicate matrix execution ids: ${duplicateMatrixIds.join(", ")}`,
);
}
for (const suite of runnerSuites) {
const suiteSize = matrix.filter(
(execution) => execution.suite.id === suite.id,
).length;
if (suiteSize !== suite.expectedMatrixSize) {
throw new Error(
`Expected ${suite.expectedMatrixSize} ${suite.id} executions; received ${suiteSize}`,
);
}
}
const expectedTotal = runnerSuites.reduce((total, suite) => total + suite.expectedMatrixSize, 0);
if (matrix.length !== expectedTotal)
throw new Error(`Expected ${expectedTotal} runner executions; received ${matrix.length}`);
return matrix;
}
export const runnerMatrix = validateRunnerCatalog();
export function runnerExecutionById(id: string): MatrixExecution {
const execution = runnerMatrix.find((candidate) => candidate.id === id);
if (!execution) throw new Error(`Unknown runner E2E execution id: ${id}`);
return execution;
}