mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip uses runner end-to-end reports to compare agent profiles and execution environments > - The report dashboard shows each reviewed final-state screenshot as a thumbnail and gallery item > - The public history publisher removed all per-attempt images before it regenerated the dashboard > - Therefore the public dashboard had the new layout but could not show the screenshots from the run > - The publisher needs a narrow rule that keeps only screenshots from the exact live fixture issue route > - This pull request keeps those trusted PNG files in every future public S3 and GitHub Pages report > - The benefit is that each future report can show its screenshot gallery without exposing logs, traces, videos, archives, arbitrary images, or generated report trees ## Linked Issues or Issue Description **What happened?** The runner E2E job captured final-state screenshots in its private artifact. The public S3 and GitHub Pages publication step removed those screenshots before it regenerated the dashboard. As a result, the public report showed the new dashboard controls but no screenshot thumbnails or gallery items. **Expected behavior** Each future public runner E2E report must include reviewed PNG screenshots from the live fixture issue. Other captures and active or unsafe evidence must stay private. **Steps to reproduce** 1. Run the runner full-stack E2E workflow on `master` before this change. 2. Open the private `runner-e2e-report-*` artifact and confirm that it contains per-attempt PNG screenshots. 3. Open the public campaign URL and confirm that the dashboard has no screenshot gallery items. **Paperclip version or commit** The issue was reproduced on commit `64d8929`, after the report design change in PR #12889. **Deployment mode** GitHub Actions with the public S3 and GitHub Pages report publishers. Related design work: Refs #12889. ## What Changed - Mark screenshots from the exact server-created live fixture issue route with `public-runner-fixture`. - Keep marked PNG files in both the S3 history bundle and the GitHub Pages bundle. - Keep captures from other issue routes, sensitive routes, and external origins private. - Bind public files to the normalized execution ID, attempt, and safe PNG base name. - Validate every retained image with the existing PNG signature and 12 MiB size checks. - Skip missing-artifact sentinel results with attempt `0` when they have no public screenshots. - Continue to remove unmarked images, videos, traces, archives, generated HTML reports, and other private evidence. - Update publisher tests, workflow checks, report copy, and the public-evidence security documentation. ## Verification - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/history.test.ts tests/runner-e2e/report.test.ts` - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts tests/runner-e2e/workflow-security.test.ts -t "uses environment-scoped OIDC"` - `pnpm test:e2e:runner:typecheck` - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/runner-e2e-dashboard.spec.ts` - `pnpm -r typecheck` - `pnpm build` - Regenerated the dashboard from retained evidence for Actions run `33968240659` without a paid matrix rerun. The public-stage proof contained 121 screenshot gallery items and thumbnail frames, with zero generated HTML report files. The trusted-fixture marker and route gate have separate focused tests. - All pull request CI checks pass on commit `ccd2b49e1`. ## Risks - This change intentionally makes marked fixture screenshots public at the campaign URL. A screenshot can show data that a raw-byte secret scan cannot detect. - The capture helper marks a screenshot only on the exact loopback issue route for the fixture that the harness created. A different issue, sensitive page, or external origin stays private. - The publisher also requires the marker, a safe normalized path, a valid PNG signature, and the size limit. - The change does not publish videos, logs, traces, archives, arbitrary images, or generated browser report trees. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5.6-sol`, with high reasoning, repository tool use, shell execution, browser inspection, and GitHub CLI access. The working context was the Codex desktop task context. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
378 lines
10 KiB
TypeScript
378 lines
10 KiB
TypeScript
export const CREDENTIAL_NAMES = [
|
|
"OPENAI_API_KEY",
|
|
"ANTHROPIC_API_KEY",
|
|
"OPENROUTER_API_KEY",
|
|
"DAYTONA_API_KEY",
|
|
] as const;
|
|
|
|
export type CredentialName = (typeof CREDENTIAL_NAMES)[number];
|
|
export type RunnerGeneration = "legacy" | "native";
|
|
export type RunnerEnvironmentId = "local" | "daytona";
|
|
export type RunnerTaskWorkMode = "standard" | "planning" | "ask";
|
|
export type RunnerTaskFlow =
|
|
| "single_turn"
|
|
| "plan_revision_acceptance"
|
|
| "question_resume_completion"
|
|
| "plan_approval_completion";
|
|
|
|
export interface SecretReference {
|
|
type: "secret_ref";
|
|
secretId: string;
|
|
version: "latest";
|
|
}
|
|
|
|
export type SecretReferenceMap = Partial<
|
|
Record<CredentialName, SecretReference>
|
|
>;
|
|
|
|
export interface AgentFixtureBuildInput {
|
|
environmentId: string;
|
|
environmentFixtureId: RunnerEnvironmentId;
|
|
workspacePath: string;
|
|
secretRefs: SecretReferenceMap;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface EnvironmentFixtureBuildInput {
|
|
secretRefs: SecretReferenceMap;
|
|
daytonaImage?: string;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface RunnerProfileFixture {
|
|
id: string;
|
|
label: string;
|
|
generation: RunnerGeneration;
|
|
groups: readonly string[];
|
|
adapterType: string;
|
|
provider: string;
|
|
model: string;
|
|
modelQualification: {
|
|
source:
|
|
| "adapter_constant"
|
|
| "qualified_runner_profile"
|
|
| "openrouter_rankings_snapshot";
|
|
qualificationId: string;
|
|
};
|
|
ranking?: {
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
};
|
|
credential: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
supportedEnvironments: readonly RunnerEnvironmentId[];
|
|
expectedRuntimeMode: RunnerGeneration;
|
|
expectedRuntimeMetadata: {
|
|
adapterType: string;
|
|
provider: string;
|
|
};
|
|
buildAgent(input: AgentFixtureBuildInput): Record<string, unknown>;
|
|
}
|
|
|
|
export interface EnvironmentFixture {
|
|
id: RunnerEnvironmentId;
|
|
label: string;
|
|
groups: readonly string[];
|
|
driver: "local" | "sandbox";
|
|
provider: "local" | "daytona";
|
|
credential?: "DAYTONA_API_KEY";
|
|
lifecycle: {
|
|
setup: "instance_managed" | "create_via_api";
|
|
probe: "run_context_via_api";
|
|
cleanup: "instance_shutdown" | "delete_via_api_and_destroy_leases";
|
|
};
|
|
expectedExecutionTarget: {
|
|
kind: "local" | "remote";
|
|
transport?: "sandbox";
|
|
};
|
|
buildEnvironment(
|
|
input: EnvironmentFixtureBuildInput,
|
|
): Record<string, unknown>;
|
|
}
|
|
|
|
export type Matcher =
|
|
| { kind: "message_exact"; expected: string }
|
|
| { kind: "message_contains"; expected: string }
|
|
| { kind: "message_occurrences"; expected: string; count: number }
|
|
| { kind: "message_regex"; pattern: string; flags?: string }
|
|
| { kind: "message_ordered"; expected: readonly string[] }
|
|
| { kind: "issue_status"; expected: string }
|
|
| { kind: "run_status"; expected: string }
|
|
| { kind: "runtime_mode"; expected: RunnerGeneration }
|
|
| { kind: "environment"; expected: RunnerEnvironmentId }
|
|
| { kind: "file_exists"; path: string }
|
|
| { kind: "file_contains"; path: string; expected: string }
|
|
| { kind: "artifact_exists"; name: string; mimeType?: string }
|
|
| { kind: "json_path"; path: string; expected: unknown }
|
|
| { kind: "json_schema"; schema: Record<string, unknown> };
|
|
|
|
export interface RunnerTaskFixture {
|
|
id: string;
|
|
label: string;
|
|
groups: readonly string[];
|
|
workMode: RunnerTaskWorkMode;
|
|
flow: RunnerTaskFlow;
|
|
expectedRunCount: number;
|
|
attemptTimeoutMs: Readonly<Record<RunnerEnvironmentId, number>>;
|
|
expectedTerminalState: {
|
|
issue: "done";
|
|
run: "succeeded";
|
|
};
|
|
buildTitle(nonce: string): string;
|
|
buildPrompt(nonce: string): string;
|
|
buildVisibleMarker(nonce: string): string;
|
|
buildRevisionRequest?(nonce: string): string;
|
|
buildQuestionAnswer?(nonce: string): {
|
|
optionLabel: string;
|
|
expectedMarker: string;
|
|
};
|
|
/** Restart the isolated Paperclip server after the waiting turn settles. */
|
|
restartServerBeforeQuestionAnswer?: boolean;
|
|
buildPlanMarkers?(nonce: string): {
|
|
draft: string;
|
|
revised: string;
|
|
};
|
|
buildMatchers(nonce: string, execution: MatrixExecution): readonly Matcher[];
|
|
}
|
|
|
|
export interface MatrixExecution {
|
|
id: string;
|
|
suite: RunnerSuiteFixture;
|
|
suiteDefinitionHash: string;
|
|
profile: RunnerProfileFixture;
|
|
environment: EnvironmentFixture;
|
|
task: RunnerTaskFixture;
|
|
groups: readonly string[];
|
|
requiredCredentials: readonly CredentialName[];
|
|
}
|
|
|
|
export interface RunnerSuiteFixture {
|
|
id: string;
|
|
label: string;
|
|
description: string;
|
|
groups: readonly string[];
|
|
profiles: readonly RunnerProfileFixture[];
|
|
environments: readonly EnvironmentFixture[];
|
|
tasks: readonly RunnerTaskFixture[];
|
|
excludedExecutionIds?: readonly string[];
|
|
expectedMatrixSize: number;
|
|
definitionMetadata?: Readonly<Record<string, unknown>>;
|
|
}
|
|
|
|
export interface MatrixJob {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
credentialName: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
timeoutMinutes: number;
|
|
needsDaytona: boolean;
|
|
}
|
|
|
|
export type FailureClass =
|
|
| "candidate_failure"
|
|
| "provider_variance"
|
|
| "transient_infrastructure"
|
|
| "permanent_infrastructure"
|
|
| "secret_leak"
|
|
| "cleanup_failure";
|
|
|
|
export type RunnerE2ECostStatus =
|
|
| "reported"
|
|
| "estimated"
|
|
| "partial"
|
|
| "unpriced"
|
|
| "unavailable"
|
|
| "not_metered";
|
|
|
|
export interface RunnerE2ERuntimeUsage {
|
|
provider: RunnerEnvironmentId;
|
|
/** Sum of the selected Paperclip heartbeat-run spans. */
|
|
agentRunDurationMs: number;
|
|
/** Sum of provider lease windows when the environment exposes leases. */
|
|
leaseDurationMs: number | null;
|
|
leaseCount: number;
|
|
cpuCores?: number;
|
|
memoryGiB?: number;
|
|
diskGiB?: number;
|
|
estimatedListCostUsd?: number;
|
|
costStatus: "estimated" | "unavailable" | "not_metered";
|
|
costSource:
|
|
| "daytona_public_list_price"
|
|
| "provider_cost_unavailable"
|
|
| "local_not_metered";
|
|
pricingAsOf?: string;
|
|
pricingUrl?: string;
|
|
}
|
|
|
|
export interface RunnerE2EBillingSummary {
|
|
llm: {
|
|
runCount: number;
|
|
runsWithTokenUsage: number;
|
|
runsWithReportedCost: number;
|
|
inputTokens: number;
|
|
outputTokens: number;
|
|
cachedInputTokens: number;
|
|
totalTokens: number;
|
|
reportedCostUsd: number;
|
|
costStatus: Exclude<RunnerE2ECostStatus, "estimated" | "not_metered">;
|
|
};
|
|
runtime: RunnerE2ERuntimeUsage;
|
|
/** Provider-reported model spend only; never includes unknown/unpriced runs. */
|
|
reportedCostUsd: number;
|
|
/** Public-list-price estimate for metered execution infrastructure. */
|
|
estimatedRuntimeCostUsd: number;
|
|
/** Reported model subtotal plus the runtime list-price estimate. */
|
|
observedAndEstimatedCostUsd: number;
|
|
complete: boolean;
|
|
}
|
|
|
|
export interface RunnerE2EResult {
|
|
schema: "paperclip.runner-e2e.result/v1" | "paperclip.runner-e2e.result/v2";
|
|
executionId: string;
|
|
suiteId?: string;
|
|
suiteDefinitionHash?: string;
|
|
source?: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
};
|
|
rankingSnapshot?: {
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
};
|
|
attempt: number;
|
|
status: "passed" | "failed";
|
|
failureClass?: FailureClass;
|
|
error?: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
runtimeMode: RunnerGeneration;
|
|
issueId?: string;
|
|
issueIdentifier?: string | null;
|
|
runIds?: string[];
|
|
startedAt: string;
|
|
finishedAt: string;
|
|
durationMs: number;
|
|
usage?: Record<string, unknown> | null;
|
|
runtimeUsage?: RunnerE2ERuntimeUsage;
|
|
billing?: RunnerE2EBillingSummary;
|
|
matcherResults?: Array<{
|
|
matcher: Matcher;
|
|
passed: boolean;
|
|
detail: string;
|
|
}>;
|
|
screenshots?: Array<{
|
|
id: string;
|
|
label: string;
|
|
file: string;
|
|
publication?: "public-runner-fixture";
|
|
}>;
|
|
cleanup: "not_started" | "passed" | "failed";
|
|
}
|
|
|
|
export interface RunnerE2ESuiteSummary {
|
|
suiteId: string;
|
|
suiteDefinitionHash: string;
|
|
expected: number;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
complete: boolean;
|
|
durationMs: number;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EAggregateBillingSummary {
|
|
testCount: number;
|
|
agentRunDurationMs: number;
|
|
leaseDurationMs: number;
|
|
llm: RunnerE2EBillingSummary["llm"];
|
|
reportedLlmCostUsd: number;
|
|
estimatedRuntimeCostUsd: number;
|
|
observedAndEstimatedCostUsd: number;
|
|
testsWithCompleteBilling: number;
|
|
}
|
|
|
|
export interface RunnerE2ECampaign {
|
|
schema: "paperclip.runner-e2e.campaign/v2";
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
eventName: string | null;
|
|
};
|
|
expected: string[];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
rankingSnapshots: Array<{
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
}>;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
results: RunnerE2EResult[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryExecution {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
status: "passed" | "failed";
|
|
durationMs: number;
|
|
attempt: number;
|
|
cleanup: RunnerE2EResult["cleanup"];
|
|
billing: RunnerE2EBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EHistoryCampaign {
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: RunnerE2ECampaign["source"];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
publicUrl: string;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
executions: RunnerE2EHistoryExecution[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryIndex {
|
|
schema: "paperclip.runner-e2e.history/v1";
|
|
updatedAt: string;
|
|
latestCampaignId: string | null;
|
|
latestGreenCampaignId: string | null;
|
|
latestBySuite: Record<string, string>;
|
|
latestGreenBySuite: Record<string, string>;
|
|
campaigns: RunnerE2EHistoryCampaign[];
|
|
}
|