mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 07:23:08 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
403 lines
11 KiB
TypeScript
403 lines
11 KiB
TypeScript
export const CREDENTIAL_NAMES = [
|
|
"OPENAI_API_KEY",
|
|
"ANTHROPIC_API_KEY",
|
|
"OPENROUTER_API_KEY",
|
|
"DAYTONA_API_KEY",
|
|
] as const;
|
|
|
|
export type CredentialName = (typeof CREDENTIAL_NAMES)[number];
|
|
export type RunnerGeneration = "legacy" | "native";
|
|
export type RunnerEnvironmentId = "local" | "daytona";
|
|
export type RunnerTaskWorkMode = "standard" | "planning" | "ask";
|
|
export type RunnerTaskFlow =
|
|
| "everyday_workflow"
|
|
| "agent_chat"
|
|
| "governed_tool_review"
|
|
| "single_turn"
|
|
| "plan_revision_acceptance"
|
|
| "question_resume_completion"
|
|
| "plan_approval_completion"
|
|
| "warm_three_turn";
|
|
|
|
export interface SecretReference {
|
|
type: "secret_ref";
|
|
secretId: string;
|
|
version: "latest";
|
|
}
|
|
|
|
export type SecretReferenceMap = Partial<
|
|
Record<CredentialName, SecretReference>
|
|
>;
|
|
|
|
export interface AgentFixtureBuildInput {
|
|
environmentId: string;
|
|
environmentFixtureId: RunnerEnvironmentId;
|
|
workspacePath: string;
|
|
secretRefs: SecretReferenceMap;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface EnvironmentFixtureBuildInput {
|
|
secretRefs: SecretReferenceMap;
|
|
daytonaImage?: string;
|
|
executionId: string;
|
|
}
|
|
|
|
export interface RunnerProfileFixture {
|
|
id: string;
|
|
label: string;
|
|
generation: RunnerGeneration;
|
|
groups: readonly string[];
|
|
adapterType: string;
|
|
provider: string;
|
|
model: string;
|
|
modelQualification: {
|
|
source:
|
|
| "adapter_constant"
|
|
| "qualified_runner_profile"
|
|
| "openrouter_rankings_snapshot";
|
|
qualificationId: string;
|
|
};
|
|
ranking?: {
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
};
|
|
credential: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
supportedEnvironments: readonly RunnerEnvironmentId[];
|
|
expectedRuntimeMode: RunnerGeneration;
|
|
expectedRuntimeMetadata: {
|
|
adapterType: string;
|
|
provider: string;
|
|
};
|
|
buildAgent(input: AgentFixtureBuildInput): Record<string, unknown>;
|
|
}
|
|
|
|
export interface EnvironmentFixture {
|
|
id: RunnerEnvironmentId;
|
|
/** Distinguishes materially different configurations that share a provider ID. */
|
|
configurationKey?: string;
|
|
label: string;
|
|
groups: readonly string[];
|
|
driver: "local" | "sandbox";
|
|
provider: "local" | "daytona";
|
|
credential?: "DAYTONA_API_KEY";
|
|
lifecycle: {
|
|
setup: "instance_managed" | "create_via_api";
|
|
probe: "run_context_via_api";
|
|
cleanup: "instance_shutdown" | "delete_via_api_and_destroy_leases";
|
|
};
|
|
expectedExecutionTarget: {
|
|
kind: "local" | "remote";
|
|
transport?: "sandbox";
|
|
};
|
|
buildEnvironment(
|
|
input: EnvironmentFixtureBuildInput,
|
|
): Record<string, unknown>;
|
|
}
|
|
|
|
export type Matcher =
|
|
| { kind: "message_exact"; expected: string }
|
|
| { kind: "message_contains"; expected: string }
|
|
| { kind: "message_occurrences"; expected: string; count: number }
|
|
| { kind: "message_regex"; pattern: string; flags?: string }
|
|
| { kind: "message_ordered"; expected: readonly string[] }
|
|
| { kind: "issue_status"; expected: string }
|
|
| { kind: "run_status"; expected: string }
|
|
| { kind: "runtime_mode"; expected: RunnerGeneration }
|
|
| { kind: "environment"; expected: RunnerEnvironmentId }
|
|
| { kind: "file_exists"; path: string }
|
|
| { kind: "file_exact"; path: string; expected: string }
|
|
| { kind: "file_contains"; path: string; expected: string }
|
|
| { kind: "artifact_exists"; name: string; mimeType?: string }
|
|
| { kind: "json_path"; path: string; expected: unknown }
|
|
| { kind: "json_schema"; schema: Record<string, unknown> };
|
|
|
|
export interface RunnerTaskFixture {
|
|
id: string;
|
|
label: string;
|
|
groups: readonly string[];
|
|
workMode: RunnerTaskWorkMode;
|
|
flow: RunnerTaskFlow;
|
|
expectedRunCount: number;
|
|
attemptTimeoutMs: Readonly<Record<RunnerEnvironmentId, number>>;
|
|
expectedTerminalState: {
|
|
issue: "done" | "in_review" | "blocked";
|
|
run: "succeeded" | "failed";
|
|
};
|
|
buildTitle(nonce: string): string;
|
|
buildPrompt(nonce: string): string;
|
|
buildVisibleMarker(nonce: string): string;
|
|
buildRevisionRequest?(nonce: string): string;
|
|
buildFollowupMessages?(nonce: string): readonly [string, string];
|
|
turnTimeoutMs?: number;
|
|
buildQuestionAnswer?(nonce: string): {
|
|
optionLabel: string;
|
|
expectedMarker: string;
|
|
};
|
|
/** Restart the isolated Paperclip server after the waiting turn settles. */
|
|
restartServerBeforeQuestionAnswer?: boolean;
|
|
toolReviewDecision?: "approve" | "decline" | "always" | "restart";
|
|
buildPlanMarkers?(nonce: string): {
|
|
draft: string;
|
|
revised: string;
|
|
};
|
|
buildMatchers(nonce: string, execution: MatrixExecution): readonly Matcher[];
|
|
}
|
|
|
|
export interface MatrixExecution {
|
|
id: string;
|
|
suite: RunnerSuiteFixture;
|
|
suiteDefinitionHash: string;
|
|
profile: RunnerProfileFixture;
|
|
environment: EnvironmentFixture;
|
|
task: RunnerTaskFixture;
|
|
groups: readonly string[];
|
|
requiredCredentials: readonly CredentialName[];
|
|
}
|
|
|
|
export interface RunnerSuiteFixture {
|
|
id: string;
|
|
label: string;
|
|
description: string;
|
|
groups: readonly string[];
|
|
profiles: readonly RunnerProfileFixture[];
|
|
environments: readonly EnvironmentFixture[];
|
|
tasks: readonly RunnerTaskFixture[];
|
|
excludedExecutionIds?: readonly string[];
|
|
expectedMatrixSize: number;
|
|
definitionMetadata?: Readonly<Record<string, unknown>>;
|
|
/** Requires an explicit suite or execution ID; excluded from scheduled --all. */
|
|
manualOnly?: boolean;
|
|
}
|
|
|
|
export interface MatrixJob {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
credentialName: Exclude<CredentialName, "DAYTONA_API_KEY">;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
timeoutMinutes: number;
|
|
needsDaytona: boolean;
|
|
}
|
|
|
|
export type FailureClass =
|
|
| "candidate_failure"
|
|
| "provider_variance"
|
|
| "transient_infrastructure"
|
|
| "permanent_infrastructure"
|
|
| "secret_leak"
|
|
| "cleanup_failure";
|
|
|
|
export type RunnerE2ECostStatus =
|
|
| "reported"
|
|
| "estimated"
|
|
| "partial"
|
|
| "unpriced"
|
|
| "unavailable"
|
|
| "not_metered";
|
|
|
|
export interface RunnerE2ERuntimeUsage {
|
|
provider: RunnerEnvironmentId;
|
|
/** Sum of the selected Paperclip heartbeat-run spans. */
|
|
agentRunDurationMs: number;
|
|
/** Sum of provider lease windows when the environment exposes leases. */
|
|
leaseDurationMs: number | null;
|
|
leaseCount: number;
|
|
cpuCores?: number;
|
|
memoryGiB?: number;
|
|
diskGiB?: number;
|
|
estimatedListCostUsd?: number;
|
|
costStatus: "estimated" | "unavailable" | "not_metered";
|
|
costSource:
|
|
| "daytona_public_list_price"
|
|
| "provider_cost_unavailable"
|
|
| "local_not_metered";
|
|
pricingAsOf?: string;
|
|
pricingUrl?: string;
|
|
}
|
|
|
|
export interface RunnerE2EBillingSummary {
|
|
llm: {
|
|
runCount: number;
|
|
runsWithTokenUsage: number;
|
|
runsWithReportedCost: number;
|
|
inputTokens: number;
|
|
outputTokens: number;
|
|
cachedInputTokens: number;
|
|
totalTokens: number;
|
|
reportedCostUsd: number;
|
|
costStatus: Exclude<RunnerE2ECostStatus, "estimated" | "not_metered">;
|
|
};
|
|
runtime: RunnerE2ERuntimeUsage;
|
|
/** Provider-reported model spend only; never includes unknown/unpriced runs. */
|
|
reportedCostUsd: number;
|
|
/** Public-list-price estimate for metered execution infrastructure. */
|
|
estimatedRuntimeCostUsd: number;
|
|
/** Reported model subtotal plus the runtime list-price estimate. */
|
|
observedAndEstimatedCostUsd: number;
|
|
complete: boolean;
|
|
}
|
|
|
|
export interface RunnerE2EResult {
|
|
schema: "paperclip.runner-e2e.result/v1" | "paperclip.runner-e2e.result/v2";
|
|
executionId: string;
|
|
suiteId?: string;
|
|
suiteDefinitionHash?: string;
|
|
source?: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
};
|
|
rankingSnapshot?: {
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
rank: number;
|
|
canonicalModelId: string;
|
|
};
|
|
attempt: number;
|
|
status: "passed" | "failed";
|
|
failureClass?: FailureClass;
|
|
error?: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
runtimeMode: RunnerGeneration;
|
|
issueId?: string;
|
|
issueIdentifier?: string | null;
|
|
runIds?: string[];
|
|
turnTimings?: Array<{
|
|
turn: number;
|
|
submittedAt: string;
|
|
runStartedAt: string | null;
|
|
runFinishedAt: string | null;
|
|
schedulerLatencyMs: number | null;
|
|
runDurationMs: number | null;
|
|
responseLatencyMs: number | null;
|
|
runId: string;
|
|
leaseAcquisitionOutcome: "created" | "resumed" | "replacement" | "unknown";
|
|
}>;
|
|
startedAt: string;
|
|
finishedAt: string;
|
|
durationMs: number;
|
|
usage?: Record<string, unknown> | null;
|
|
runtimeUsage?: RunnerE2ERuntimeUsage;
|
|
billing?: RunnerE2EBillingSummary;
|
|
matcherResults?: Array<{
|
|
matcher: Matcher;
|
|
passed: boolean;
|
|
detail: string;
|
|
}>;
|
|
screenshots?: Array<{
|
|
id: string;
|
|
label: string;
|
|
file: string;
|
|
publication?: "public-runner-fixture";
|
|
/** Absent in historical results; new captures bind the exact PNG bytes. */
|
|
sha256?: string;
|
|
}>;
|
|
cleanup: "not_started" | "passed" | "failed";
|
|
}
|
|
|
|
export interface RunnerE2ESuiteSummary {
|
|
suiteId: string;
|
|
suiteDefinitionHash: string;
|
|
expected: number;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
complete: boolean;
|
|
durationMs: number;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EAggregateBillingSummary {
|
|
testCount: number;
|
|
agentRunDurationMs: number;
|
|
leaseDurationMs: number;
|
|
llm: RunnerE2EBillingSummary["llm"];
|
|
reportedLlmCostUsd: number;
|
|
estimatedRuntimeCostUsd: number;
|
|
observedAndEstimatedCostUsd: number;
|
|
testsWithCompleteBilling: number;
|
|
}
|
|
|
|
export interface RunnerE2ECampaign {
|
|
schema: "paperclip.runner-e2e.campaign/v2";
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: {
|
|
sha: string | null;
|
|
ref: string | null;
|
|
workflowRunUrl: string | null;
|
|
eventName: string | null;
|
|
};
|
|
expected: string[];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
rankingSnapshots: Array<{
|
|
snapshotId: string;
|
|
capturedAt: string;
|
|
sourceUrl: string;
|
|
}>;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
results: RunnerE2EResult[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryExecution {
|
|
executionId: string;
|
|
suiteId: string;
|
|
profileId: string;
|
|
environmentId: RunnerEnvironmentId;
|
|
caseId: string;
|
|
provider: string;
|
|
model: string;
|
|
status: "passed" | "failed";
|
|
durationMs: number;
|
|
attempt: number;
|
|
cleanup: RunnerE2EResult["cleanup"];
|
|
billing: RunnerE2EBillingSummary;
|
|
}
|
|
|
|
export interface RunnerE2EHistoryCampaign {
|
|
campaignId: string;
|
|
generatedAt: string;
|
|
source: RunnerE2ECampaign["source"];
|
|
complete: boolean;
|
|
selected: number;
|
|
executed: number;
|
|
passed: number;
|
|
failed: number;
|
|
retries: number;
|
|
cleanupPassed: boolean;
|
|
publicUrl: string;
|
|
billing: RunnerE2EAggregateBillingSummary;
|
|
suites: RunnerE2ESuiteSummary[];
|
|
executions: RunnerE2EHistoryExecution[];
|
|
}
|
|
|
|
export interface RunnerE2EHistoryIndex {
|
|
schema: "paperclip.runner-e2e.history/v1";
|
|
updatedAt: string;
|
|
latestCampaignId: string | null;
|
|
latestGreenCampaignId: string | null;
|
|
latestBySuite: Record<string, string>;
|
|
latestGreenBySuite: Record<string, string>;
|
|
campaigns: RunnerE2EHistoryCampaign[];
|
|
}
|