Files
PaperClipAI/tests/runner-e2e/types.ts
T
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00

425 lines
12 KiB
TypeScript

export const CREDENTIAL_NAMES = [
"OPENAI_API_KEY",
"ANTHROPIC_API_KEY",
"OPENROUTER_API_KEY",
"DAYTONA_API_KEY",
] as const;
export type CredentialName = (typeof CREDENTIAL_NAMES)[number];
export type RunnerGeneration = "legacy" | "native";
export type RunnerEnvironmentId = "local" | "daytona";
export type RunnerTaskWorkMode = "standard" | "planning" | "ask";
export type RunnerTaskFlow =
| "everyday_workflow"
| "continuation"
| "first_task"
| "agent_chat"
| "governed_tool_review"
| "single_turn"
| "plan_revision_acceptance"
| "question_resume_completion"
| "plan_approval_completion"
| "warm_three_turn";
export interface SecretReference {
type: "secret_ref";
secretId: string;
version: "latest";
}
export type SecretReferenceMap = Partial<
Record<CredentialName, SecretReference>
>;
export interface AgentFixtureBuildInput {
environmentId: string;
environmentFixtureId: RunnerEnvironmentId;
workspacePath: string;
secretRefs: SecretReferenceMap;
executionId: string;
}
export interface EnvironmentFixtureBuildInput {
secretRefs: SecretReferenceMap;
daytonaImage?: string;
executionId: string;
}
export interface RunnerProfileFixture {
id: string;
label: string;
generation: RunnerGeneration;
groups: readonly string[];
adapterType: string;
provider: string;
model: string;
modelQualification: {
source:
| "adapter_constant"
| "qualified_runner_profile"
| "openrouter_rankings_snapshot";
qualificationId: string;
};
ranking?: {
rank: number;
canonicalModelId: string;
snapshotId: string;
capturedAt: string;
sourceUrl: string;
};
credential: Exclude<CredentialName, "DAYTONA_API_KEY">;
supportedEnvironments: readonly RunnerEnvironmentId[];
expectedRuntimeMode: RunnerGeneration;
expectedRuntimeMetadata: {
adapterType: string;
provider: string;
};
buildAgent(input: AgentFixtureBuildInput): Record<string, unknown>;
}
export interface EnvironmentFixture {
id: RunnerEnvironmentId;
/** Distinguishes materially different configurations that share a provider ID. */
configurationKey?: string;
label: string;
groups: readonly string[];
driver: "local" | "sandbox";
provider: "local" | "daytona";
credential?: "DAYTONA_API_KEY";
lifecycle: {
setup: "instance_managed" | "create_via_api";
probe: "run_context_via_api";
cleanup: "instance_shutdown" | "delete_via_api_and_destroy_leases";
};
expectedExecutionTarget: {
kind: "local" | "remote";
transport?: "sandbox";
};
buildEnvironment(
input: EnvironmentFixtureBuildInput,
): Record<string, unknown>;
}
export type Matcher =
| { kind: "message_exact"; expected: string }
| { kind: "message_contains"; expected: string }
| { kind: "message_occurrences"; expected: string; count: number }
| { kind: "message_regex"; pattern: string; flags?: string }
| { kind: "message_ordered"; expected: readonly string[] }
| { kind: "issue_status"; expected: string }
| { kind: "run_status"; expected: string }
| { kind: "runtime_mode"; expected: RunnerGeneration }
| { kind: "environment"; expected: RunnerEnvironmentId }
| { kind: "file_exists"; path: string }
| { kind: "file_exact"; path: string; expected: string }
| { kind: "file_contains"; path: string; expected: string }
| { kind: "artifact_exists"; name: string; mimeType?: string }
| { kind: "json_path"; path: string; expected: unknown }
| { kind: "json_schema"; schema: Record<string, unknown> };
export interface RunnerTaskFixture {
id: string;
label: string;
groups: readonly string[];
workMode: RunnerTaskWorkMode;
flow: RunnerTaskFlow;
expectedRunCount: number;
/** Optional lower bound; expectedRunCount remains the maximum/cost estimate. */
minimumExpectedRunCount?: number;
attemptTimeoutMs: Readonly<Record<RunnerEnvironmentId, number>>;
expectedTerminalState: {
issue: "done" | "in_review" | "blocked";
run: "succeeded" | "failed";
};
buildTitle(nonce: string): string;
buildPrompt(nonce: string): string;
buildVisibleMarker(nonce: string): string;
buildRevisionRequest?(nonce: string): string;
buildFollowupMessages?(nonce: string): readonly [string, string];
turnTimeoutMs?: number;
buildQuestionAnswer?(nonce: string): {
optionLabel: string;
expectedMarker: string;
};
/** Restart the isolated Paperclip server after the waiting turn settles. */
restartServerBeforeQuestionAnswer?: boolean;
toolReviewDecision?: "approve" | "decline" | "always" | "restart";
buildPlanMarkers?(nonce: string): {
draft: string;
revised: string;
};
buildMatchers(nonce: string, execution: MatrixExecution): readonly Matcher[];
}
export interface MatrixExecution {
id: string;
suite: RunnerSuiteFixture;
suiteDefinitionHash: string;
profile: RunnerProfileFixture;
environment: EnvironmentFixture;
task: RunnerTaskFixture;
groups: readonly string[];
requiredCredentials: readonly CredentialName[];
}
export interface RunnerSuiteFixture {
id: string;
label: string;
description: string;
groups: readonly string[];
profiles: readonly RunnerProfileFixture[];
environments: readonly EnvironmentFixture[];
tasks: readonly RunnerTaskFixture[];
excludedExecutionIds?: readonly string[];
expectedMatrixSize: number;
definitionMetadata?: Readonly<Record<string, unknown>>;
/** Requires an explicit suite or execution ID; excluded from scheduled --all. */
manualOnly?: boolean;
}
export interface MatrixJob {
executionId: string;
suiteId: string;
profileId: string;
credentialName: Exclude<CredentialName, "DAYTONA_API_KEY">;
environmentId: RunnerEnvironmentId;
caseId: string;
timeoutMinutes: number;
needsDaytona: boolean;
}
export type FailureClass =
| "candidate_failure"
| "provider_variance"
| "transient_infrastructure"
| "permanent_infrastructure"
| "secret_leak"
| "cleanup_failure";
export type RunnerE2ECostStatus =
| "reported"
| "estimated"
| "partial"
| "unpriced"
| "unavailable"
| "not_metered";
export interface RunnerE2ERuntimeUsage {
provider: RunnerEnvironmentId;
/** Sum of the selected Paperclip heartbeat-run spans. */
agentRunDurationMs: number;
/** Sum of provider lease windows when the environment exposes leases. */
leaseDurationMs: number | null;
leaseCount: number;
cpuCores?: number;
memoryGiB?: number;
diskGiB?: number;
estimatedListCostUsd?: number;
costStatus: "estimated" | "unavailable" | "not_metered";
costSource:
| "daytona_public_list_price"
| "provider_cost_unavailable"
| "local_not_metered";
pricingAsOf?: string;
pricingUrl?: string;
}
export interface RunnerE2EBillingSummary {
llm: {
runCount: number;
runsWithTokenUsage: number;
runsWithReportedCost: number;
inputTokens: number;
outputTokens: number;
cachedInputTokens: number;
totalTokens: number;
reportedCostUsd: number;
costStatus: Exclude<RunnerE2ECostStatus, "estimated" | "not_metered">;
};
runtime: RunnerE2ERuntimeUsage;
/** Provider-reported model spend only; never includes unknown/unpriced runs. */
reportedCostUsd: number;
/** Public-list-price estimate for metered execution infrastructure. */
estimatedRuntimeCostUsd: number;
/** Separately recorded post-processing judge usage; absent when not judged. */
judge?: { inputTokens: number | null; outputTokens: number | null; estimatedCostUsd: number | null; reservedCostUsd: number };
/** Reported model subtotal plus runtime and judge list-price estimates. */
observedAndEstimatedCostUsd: number | null;
complete: boolean;
}
export interface RunnerE2EResult {
schema: "paperclip.runner-e2e.result/v1" | "paperclip.runner-e2e.result/v2";
executionId: string;
suiteId?: string;
suiteDefinitionHash?: string;
source?: {
sha: string | null;
ref: string | null;
workflowRunUrl: string | null;
};
rankingSnapshot?: {
snapshotId: string;
capturedAt: string;
sourceUrl: string;
rank: number;
canonicalModelId: string;
};
attempt: number;
status: "passed" | "failed";
failureClass?: FailureClass;
error?: string;
profileId: string;
environmentId: RunnerEnvironmentId;
caseId: string;
provider: string;
model: string;
runtimeMode: RunnerGeneration;
issueId?: string;
issueIdentifier?: string | null;
runIds?: string[];
turnTimings?: Array<{
turn: number;
submittedAt: string;
runStartedAt: string | null;
runFinishedAt: string | null;
schedulerLatencyMs: number | null;
runDurationMs: number | null;
responseLatencyMs: number | null;
runId: string;
leaseAcquisitionOutcome: "created" | "resumed" | "replacement" | "unknown";
}>;
startedAt: string;
finishedAt: string;
durationMs: number;
usage?: Record<string, unknown> | null;
runtimeUsage?: RunnerE2ERuntimeUsage;
billing?: RunnerE2EBillingSummary;
matcherResults?: Array<{
matcher: Matcher;
passed: boolean;
detail: string;
}>;
screenshots?: Array<{
id: string;
label: string;
file: string;
publication?: "public-runner-fixture";
/** Absent in historical results; new captures bind the exact PNG bytes. */
sha256?: string;
}>;
firstTask?: import("./first-task-scoring.js").FirstTaskEvidence;
firstTaskQuality?: import("./first-task-quality.js").FirstTaskQuality;
cleanup: "not_started" | "passed" | "failed";
}
export interface RunnerE2ESuiteSummary {
suiteId: string;
suiteDefinitionHash: string;
expected: number;
selected: number;
executed: number;
passed: number;
failed: number;
incomplete?: number;
retries: number;
cleanupPassed: boolean;
complete: boolean;
durationMs: number;
billing: RunnerE2EAggregateBillingSummary;
}
export interface RunnerE2EJudgeBillingSummary {
attempts: number;
inputTokens: number;
outputTokens: number;
estimatedCostUsd: number | null;
reservedCostUsd: number;
attemptsWithUnknownUsage: number;
}
export interface RunnerE2EAggregateBillingSummary {
judge?: RunnerE2EJudgeBillingSummary;
testCount: number;
agentRunDurationMs: number;
leaseDurationMs: number;
llm: RunnerE2EBillingSummary["llm"];
reportedLlmCostUsd: number;
estimatedRuntimeCostUsd: number;
observedAndEstimatedCostUsd: number | null;
testsWithCompleteBilling: number;
}
export interface RunnerE2ECampaign {
schema: "paperclip.runner-e2e.campaign/v2";
campaignId: string;
generatedAt: string;
source: {
sha: string | null;
ref: string | null;
workflowRunUrl: string | null;
eventName: string | null;
};
expected: string[];
complete: boolean;
selected: number;
executed: number;
passed: number;
failed: number;
incomplete?: number;
retries: number;
cleanupPassed: boolean;
rankingSnapshots: Array<{
snapshotId: string;
capturedAt: string;
sourceUrl: string;
}>;
billing: RunnerE2EAggregateBillingSummary;
suites: RunnerE2ESuiteSummary[];
results: RunnerE2EResult[];
}
export interface RunnerE2EHistoryExecution {
executionId: string;
suiteId: string;
profileId: string;
environmentId: RunnerEnvironmentId;
caseId: string;
provider: string;
model: string;
status: "passed" | "failed" | "incomplete";
durationMs: number;
attempt: number;
cleanup: RunnerE2EResult["cleanup"];
billing: RunnerE2EBillingSummary;
}
export interface RunnerE2EHistoryCampaign {
campaignId: string;
generatedAt: string;
source: RunnerE2ECampaign["source"];
complete: boolean;
selected: number;
executed: number;
passed: number;
failed: number;
incomplete?: number;
retries: number;
cleanupPassed: boolean;
publicUrl: string;
billing: RunnerE2EAggregateBillingSummary;
suites: RunnerE2ESuiteSummary[];
executions: RunnerE2EHistoryExecution[];
}
export interface RunnerE2EHistoryIndex {
schema: "paperclip.runner-e2e.history/v1";
updatedAt: string;
latestCampaignId: string | null;
latestGreenCampaignId: string | null;
latestBySuite: Record<string, string>;
latestGreenBySuite: Record<string, string>;
campaigns: RunnerE2EHistoryCampaign[];
}