mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip manages AI agents that perform work. > - Paperclip Runner connects durable task runs to local provider processes. > - The full-stack paid matrix exposed failures after the runner integrity repair. > - Verified JavaScript entrypoints lost their relative module graph when Linux executed them through descriptor paths. > - Returned provider startup errors also remained pending and became indeterminate after recovery. > - Sparse Codex tool lifecycle events lost the `write_document` identity before task transcript projection. > - This pull request repairs those three boundaries and makes the structured-question fixture deterministic. > - The benefit is repeatable provider startup, exact failure replay, and correct inline Plan placement. ## Linked Issues or Issue Description Refs #12721 and #12700. **What happened?** The paid runner matrix failed ACPX and OpenCode startup before provider session creation. The runner journal then replaced the original startup error with an indeterminate recovery result. Native Codex saved a Plan but rendered it only as a fallback card. A legacy Claude waiting reply could also echo the reserved terminal marker before the answer arrived. **Expected behavior** Verified JavaScript providers must start from immutable descriptor-backed artifacts. Returned startup failures must persist as terminal failed command results. Native tool lifecycle updates must preserve the `write_document` boundary. Pre-answer fixture output must not contain the reserved terminal marker. **Steps to reproduce** 1. Run the local provider cells in the Runner Full-Stack E2E workflow. 2. Observe ACPX and OpenCode fail during `session.open` before provider execution. 3. Observe recovery report `execution_indeterminate` instead of the original startup error. 4. Run the native Codex Plan cell and observe the fallback Plan card after the tool activity row. 5. Run the legacy Claude structured-question resume cell and observe an early marker echo in waiting prose. **Paperclip version or commit** `0f9452101740835ce0b1488a204bf48acd5bafc3` **Deployment mode** Local development with the paid GitHub Actions acceptance workflow. ## What Changed - Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM entrypoints before hashing and verified descriptor launch. - Anchor ACPX dynamic provider package resolution at a controller-derived provider-pack root and keep that root out of the provider child environment. - Persist executor-returned startup errors as redacted durable failed command results while retaining indeterminate recovery for true process death. - Coalesce sparse native tool items by stable ID so a late `write_document` name, input, and result reach the transcript boundary once. - Forbid the structured-question fixture from spelling or announcing its reserved terminal marker before the user answers. ## Verification - Rust and TypeScript regression tests cover durable failed replay, true crash ambiguity, bundle closure, package-root derivation, environment filtering, exact Codex tool lifecycle coalescing, and prompt determinism. - Local execution is intentionally limited to formatters and static diff checks. GitHub Actions will run tests, type checks, builds, and security checks. - After ordinary CI is green, scoped paid cells will validate one ACPX launch, one OpenCode launch, native Codex Plan projection, and legacy Claude structured resume before a complete matrix rerun. - Prior failing matrix: https://github.com/paperclipai/paperclip/actions/runs/33682434315 ## Risks - Bundling changes the bytes covered by provider launch hashes. Provider-pack generation already hashes the final built files. - ACPX still loads qualified provider packages dynamically. The controller supplies a normalized package root, while existing version, digest, path, and descriptor checks remain active. - Durable `failed` is terminal. Replays return the same redacted result and do not execute the provider effect twice. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex based on GPT-5 with agentic reasoning, repository inspection, code editing, Git, parallel subagents, and GitHub Actions coordination. The exact deployed snapshot and context-window size are not exposed to this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked related public work or described the bug in this PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [ ] I have run tests locally and they pass (intentionally deferred to GitHub Actions) - [x] I have added or updated tests where applicable - [x] No documentation change is required for this runtime repair - [x] I have considered and documented the risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
377 lines
10 KiB
TypeScript
377 lines
10 KiB
TypeScript
export const CREDENTIAL_NAMES = [
|
|
"OPENAI_API_KEY",
|
|
"ANTHROPIC_API_KEY",
|
|
"OPENROUTER_API_KEY",
|
|
"DAYTONA_API_KEY",
|
|
] as const;
|
|
|
|
export type CredentialName = (typeof CREDENTIAL_NAMES)[number];
|
|
export type RunnerGeneration = "legacy" | "native";
|
|
export type RunnerEnvironmentId = "local" | "daytona";
|
|
export type RunnerTaskWorkMode = "standard" | "planning" | "ask";
|
|
export type RunnerTaskFlow =
|
|
| "single_turn"
|
|
| "plan_revision_acceptance"
|
|
| "question_resume_completion"
|
|
| "plan_approval_completion";
|
|
|
|
export interface SecretReference {
|
|
type: "secret_ref";
|
|
secretId: string;
|
|
version: "latest";
|
|
}
|
|
|
|
export type SecretReferenceMap = Partial<
|
|
Record<CredentialName, SecretReference>
|
|
>;
|
|
|
|
export interface AgentFixtureBuildInput {
|
|
environmentId: string;
|
|
environmentFixtureId: RunnerEnvironmentId;
|
|
workspacePath: string;
|
|
secretRefs: SecretReferenceMap;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface EnvironmentFixtureBuildInput {
|
|
secretRefs: SecretReferenceMap;
|
|
daytonaImage?: string;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface RunnerProfileFixture {
|
|
id: string;
|
|
label: string;
|
|
generation: RunnerGeneration;
|
|
groups: readonly string[];
|
|
adapterType: string;
|
|
provider: string;
|
|
model: string;
|
|
modelQualification: {
|
|
source:
|
|
| "adapter_constant"
|
|
| "qualified_runner_profile"
|
|
| "openrouter_rankings_snapshot";
|
|
qualificationId: string;
|
|
};
|
|
ranking?: {
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
};
|
|
credential: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
supportedEnvironments: readonly RunnerEnvironmentId[];
|
|
expectedRuntimeMode: RunnerGeneration;
|
|
expectedRuntimeMetadata: {
|
|
adapterType: string;
|
|
provider: string;
|
|
};
|
|
buildAgent(input: AgentFixtureBuildInput): Record<string, unknown>;
|
|
}
|
|
|
|
export interface EnvironmentFixture {
|
|
id: RunnerEnvironmentId;
|
|
label: string;
|
|
groups: readonly string[];
|
|
driver: "local" | "sandbox";
|
|
provider: "local" | "daytona";
|
|
credential?: "DAYTONA_API_KEY";
|
|
lifecycle: {
|
|
setup: "instance_managed" | "create_via_api";
|
|
probe: "run_context_via_api";
|
|
cleanup: "instance_shutdown" | "delete_via_api_and_destroy_leases";
|
|
};
|
|
expectedExecutionTarget: {
|
|
kind: "local" | "remote";
|
|
transport?: "sandbox";
|
|
};
|
|
buildEnvironment(
|
|
input: EnvironmentFixtureBuildInput,
|
|
): Record<string, unknown>;
|
|
}
|
|
|
|
export type Matcher =
|
|
| { kind: "message_exact"; expected: string }
|
|
| { kind: "message_contains"; expected: string }
|
|
| { kind: "message_occurrences"; expected: string; count: number }
|
|
| { kind: "message_regex"; pattern: string; flags?: string }
|
|
| { kind: "message_ordered"; expected: readonly string[] }
|
|
| { kind: "issue_status"; expected: string }
|
|
| { kind: "run_status"; expected: string }
|
|
| { kind: "runtime_mode"; expected: RunnerGeneration }
|
|
| { kind: "environment"; expected: RunnerEnvironmentId }
|
|
| { kind: "file_exists"; path: string }
|
|
| { kind: "file_contains"; path: string; expected: string }
|
|
| { kind: "artifact_exists"; name: string; mimeType?: string }
|
|
| { kind: "json_path"; path: string; expected: unknown }
|
|
| { kind: "json_schema"; schema: Record<string, unknown> };
|
|
|
|
export interface RunnerTaskFixture {
|
|
id: string;
|
|
label: string;
|
|
groups: readonly string[];
|
|
workMode: RunnerTaskWorkMode;
|
|
flow: RunnerTaskFlow;
|
|
expectedRunCount: number;
|
|
attemptTimeoutMs: Readonly<Record<RunnerEnvironmentId, number>>;
|
|
expectedTerminalState: {
|
|
issue: "done";
|
|
run: "succeeded";
|
|
};
|
|
buildTitle(nonce: string): string;
|
|
buildPrompt(nonce: string): string;
|
|
buildVisibleMarker(nonce: string): string;
|
|
buildRevisionRequest?(nonce: string): string;
|
|
buildQuestionAnswer?(nonce: string): {
|
|
optionLabel: string;
|
|
expectedMarker: string;
|
|
};
|
|
/** Restart the isolated Paperclip server after the waiting turn settles. */
|
|
restartServerBeforeQuestionAnswer?: boolean;
|
|
buildPlanMarkers?(nonce: string): {
|
|
draft: string;
|
|
revised: string;
|
|
};
|
|
buildMatchers(nonce: string, execution: MatrixExecution): readonly Matcher[];
|
|
}
|
|
|
|
export interface MatrixExecution {
|
|
id: string;
|
|
suite: RunnerSuiteFixture;
|
|
suiteDefinitionHash: string;
|
|
profile: RunnerProfileFixture;
|
|
environment: EnvironmentFixture;
|
|
task: RunnerTaskFixture;
|
|
groups: readonly string[];
|
|
requiredCredentials: readonly CredentialName[];
|
|
}
|
|
|
|
export interface RunnerSuiteFixture {
|
|
id: string;
|
|
label: string;
|
|
description: string;
|
|
groups: readonly string[];
|
|
profiles: readonly RunnerProfileFixture[];
|
|
environments: readonly EnvironmentFixture[];
|
|
tasks: readonly RunnerTaskFixture[];
|
|
excludedExecutionIds?: readonly string[];
|
|
expectedMatrixSize: number;
|
|
definitionMetadata?: Readonly<Record<string, unknown>>;
|
|
}
|
|
|
|
export interface MatrixJob {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
credentialName: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
timeoutMinutes: number;
|
|
needsDaytona: boolean;
|
|
}
|
|
|
|
export type FailureClass =
|
|
| "candidate_failure"
|
|
| "provider_variance"
|
|
| "transient_infrastructure"
|
|
| "permanent_infrastructure"
|
|
| "secret_leak"
|
|
| "cleanup_failure";
|
|
|
|
export type RunnerE2ECostStatus =
|
|
| "reported"
|
|
| "estimated"
|
|
| "partial"
|
|
| "unpriced"
|
|
| "unavailable"
|
|
| "not_metered";
|
|
|
|
export interface RunnerE2ERuntimeUsage {
|
|
provider: RunnerEnvironmentId;
|
|
/** Sum of the selected Paperclip heartbeat-run spans. */
|
|
agentRunDurationMs: number;
|
|
/** Sum of provider lease windows when the environment exposes leases. */
|
|
leaseDurationMs: number | null;
|
|
leaseCount: number;
|
|
cpuCores?: number;
|
|
memoryGiB?: number;
|
|
diskGiB?: number;
|
|
estimatedListCostUsd?: number;
|
|
costStatus: "estimated" | "unavailable" | "not_metered";
|
|
costSource:
|
|
| "daytona_public_list_price"
|
|
| "provider_cost_unavailable"
|
|
| "local_not_metered";
|
|
pricingAsOf?: string;
|
|
pricingUrl?: string;
|
|
}
|
|
|
|
export interface RunnerE2EBillingSummary {
|
|
llm: {
|
|
runCount: number;
|
|
runsWithTokenUsage: number;
|
|
runsWithReportedCost: number;
|
|
inputTokens: number;
|
|
outputTokens: number;
|
|
cachedInputTokens: number;
|
|
totalTokens: number;
|
|
reportedCostUsd: number;
|
|
costStatus: Exclude<RunnerE2ECostStatus, "estimated" | "not_metered">;
|
|
};
|
|
runtime: RunnerE2ERuntimeUsage;
|
|
/** Provider-reported model spend only; never includes unknown/unpriced runs. */
|
|
reportedCostUsd: number;
|
|
/** Public-list-price estimate for metered execution infrastructure. */
|
|
estimatedRuntimeCostUsd: number;
|
|
/** Reported model subtotal plus the runtime list-price estimate. */
|
|
observedAndEstimatedCostUsd: number;
|
|
complete: boolean;
|
|
}
|
|
|
|
export interface RunnerE2EResult {
|
|
schema: "paperclip.runner-e2e.result/v1" | "paperclip.runner-e2e.result/v2";
|
|
executionId: string;
|
|
suiteId?: string;
|
|
suiteDefinitionHash?: string;
|
|
source?: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
};
|
|
rankingSnapshot?: {
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
};
|
|
attempt: number;
|
|
status: "passed" | "failed";
|
|
failureClass?: FailureClass;
|
|
error?: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
runtimeMode: RunnerGeneration;
|
|
issueId?: string;
|
|
issueIdentifier?: string | null;
|
|
runIds?: string[];
|
|
startedAt: string;
|
|
finishedAt: string;
|
|
durationMs: number;
|
|
usage?: Record<string, unknown> | null;
|
|
runtimeUsage?: RunnerE2ERuntimeUsage;
|
|
billing?: RunnerE2EBillingSummary;
|
|
matcherResults?: Array<{
|
|
matcher: Matcher;
|
|
passed: boolean;
|
|
detail: string;
|
|
}>;
|
|
screenshots?: Array<{
|
|
id: string;
|
|
label: string;
|
|
file: string;
|
|
}>;
|
|
cleanup: "not_started" | "passed" | "failed";
|
|
}
|
|
|
|
export interface RunnerE2ESuiteSummary {
|
|
suiteId: string;
|
|
suiteDefinitionHash: string;
|
|
expected: number;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
complete: boolean;
|
|
durationMs: number;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EAggregateBillingSummary {
|
|
testCount: number;
|
|
agentRunDurationMs: number;
|
|
leaseDurationMs: number;
|
|
llm: RunnerE2EBillingSummary["llm"];
|
|
reportedLlmCostUsd: number;
|
|
estimatedRuntimeCostUsd: number;
|
|
observedAndEstimatedCostUsd: number;
|
|
testsWithCompleteBilling: number;
|
|
}
|
|
|
|
export interface RunnerE2ECampaign {
|
|
schema: "paperclip.runner-e2e.campaign/v2";
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
eventName: string | null;
|
|
};
|
|
expected: string[];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
rankingSnapshots: Array<{
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
}>;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
results: RunnerE2EResult[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryExecution {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
status: "passed" | "failed";
|
|
durationMs: number;
|
|
attempt: number;
|
|
cleanup: RunnerE2EResult["cleanup"];
|
|
billing: RunnerE2EBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EHistoryCampaign {
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: RunnerE2ECampaign["source"];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
publicUrl: string;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
executions: RunnerE2EHistoryExecution[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryIndex {
|
|
schema: "paperclip.runner-e2e.history/v1";
|
|
updatedAt: string;
|
|
latestCampaignId: string | null;
|
|
latestGreenCampaignId: string | null;
|
|
latestBySuite: Record<string, string>;
|
|
latestGreenBySuite: Record<string, string>;
|
|
campaigns: RunnerE2EHistoryCampaign[];
|
|
}
|