mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
243 lines
6.8 KiB
TypeScript
243 lines
6.8 KiB
TypeScript
import type { RunnerE2EResult } from "./types.js";
|
|
|
|
type Rule = (value: unknown, at: string) => void;
|
|
const invalid = (at: string): never => {
|
|
throw new Error(`Invalid retained runner result field: ${at}`);
|
|
};
|
|
const string: Rule = (value, at) => {
|
|
if (typeof value !== "string") invalid(at);
|
|
};
|
|
const number: Rule = (value, at) => {
|
|
if (typeof value !== "number" || !Number.isFinite(value) || value < 0)
|
|
invalid(at);
|
|
};
|
|
const integer: Rule = (value, at) => {
|
|
number(value, at);
|
|
if (!Number.isSafeInteger(value)) invalid(at);
|
|
};
|
|
const boolean: Rule = (value, at) => {
|
|
if (typeof value !== "boolean") invalid(at);
|
|
};
|
|
const object: Rule = (value, at) => {
|
|
if (!value || typeof value !== "object" || Array.isArray(value)) invalid(at);
|
|
};
|
|
const optional =
|
|
(rule: Rule): Rule =>
|
|
(value, at) => {
|
|
if (value !== undefined) rule(value, at);
|
|
};
|
|
const nullable =
|
|
(rule: Rule): Rule =>
|
|
(value, at) => {
|
|
if (value !== null) rule(value, at);
|
|
};
|
|
const oneOf =
|
|
(...values: string[]): Rule =>
|
|
(value, at) => {
|
|
if (typeof value !== "string" || !values.includes(value)) invalid(at);
|
|
};
|
|
const array =
|
|
(rule: Rule): Rule =>
|
|
(value, at) => {
|
|
if (!Array.isArray(value)) invalid(at);
|
|
(value as unknown[]).forEach((item, index) =>
|
|
rule(item, `${at}[${index}]`),
|
|
);
|
|
};
|
|
const shape =
|
|
(fields: Record<string, Rule>): Rule =>
|
|
(value, at) => {
|
|
object(value, at);
|
|
for (const [key, rule] of Object.entries(fields))
|
|
rule((value as Record<string, unknown>)[key], `${at}.${key}`);
|
|
};
|
|
const date: Rule = (value, at) => {
|
|
string(value, at);
|
|
if (!Number.isFinite(Date.parse(value as string))) invalid(at);
|
|
};
|
|
const url: Rule = (value, at) => {
|
|
string(value, at);
|
|
try {
|
|
const parsed = new URL(value as string);
|
|
if (
|
|
!["https:", "http:"].includes(parsed.protocol) ||
|
|
parsed.username ||
|
|
parsed.password
|
|
)
|
|
invalid(at);
|
|
} catch {
|
|
invalid(at);
|
|
}
|
|
};
|
|
const relativeFile: Rule = (value, at) => {
|
|
string(value, at);
|
|
if (
|
|
!/^[A-Za-z0-9_.\-/]+$/.test(value as string) ||
|
|
(value as string).startsWith("/") ||
|
|
(value as string)
|
|
.split("/")
|
|
.some((part) => !part || part === "." || part === "..")
|
|
)
|
|
invalid(at);
|
|
};
|
|
const runtime = shape({
|
|
provider: oneOf("local", "daytona"),
|
|
agentRunDurationMs: number,
|
|
leaseDurationMs: nullable(number),
|
|
leaseCount: integer,
|
|
cpuCores: optional(number),
|
|
memoryGiB: optional(number),
|
|
diskGiB: optional(number),
|
|
estimatedListCostUsd: optional(number),
|
|
costStatus: oneOf("estimated", "unavailable", "not_metered"),
|
|
costSource: oneOf(
|
|
"daytona_public_list_price",
|
|
"provider_cost_unavailable",
|
|
"local_not_metered",
|
|
),
|
|
pricingAsOf: optional(string),
|
|
pricingUrl: optional(url),
|
|
});
|
|
const llm = shape({
|
|
runCount: integer,
|
|
runsWithTokenUsage: integer,
|
|
runsWithReportedCost: integer,
|
|
inputTokens: number,
|
|
outputTokens: number,
|
|
cachedInputTokens: number,
|
|
totalTokens: number,
|
|
reportedCostUsd: number,
|
|
costStatus: oneOf("reported", "partial", "unpriced", "unavailable"),
|
|
});
|
|
const billing = shape({
|
|
llm,
|
|
runtime,
|
|
reportedCostUsd: number,
|
|
estimatedRuntimeCostUsd: number,
|
|
observedAndEstimatedCostUsd: number,
|
|
complete: boolean,
|
|
});
|
|
const matcher: Rule = (value, at) => {
|
|
object(value, at);
|
|
const item = value as Record<string, unknown>;
|
|
string(item.kind, `${at}.kind`);
|
|
const kinds: Record<string, Rule> = {
|
|
message_exact: shape({ expected: string }),
|
|
message_contains: shape({ expected: string }),
|
|
message_occurrences: shape({ expected: string, count: integer }),
|
|
message_regex: shape({ pattern: string, flags: optional(string) }),
|
|
message_ordered: shape({ expected: array(string) }),
|
|
issue_status: shape({ expected: string }),
|
|
run_status: shape({ expected: string }),
|
|
runtime_mode: shape({ expected: oneOf("legacy", "native") }),
|
|
environment: shape({ expected: oneOf("local", "daytona") }),
|
|
file_exists: shape({ path: string }),
|
|
file_exact: shape({ path: string, expected: string }),
|
|
file_contains: shape({ path: string, expected: string }),
|
|
artifact_exists: shape({ name: string, mimeType: optional(string) }),
|
|
json_path: shape({ path: string }),
|
|
json_schema: shape({ schema: object }),
|
|
};
|
|
const rule = kinds[item.kind as string];
|
|
if (!rule) invalid(`${at}.kind`);
|
|
rule(value, at);
|
|
};
|
|
const fields = {
|
|
schema: oneOf(
|
|
"paperclip.runner-e2e.result/v1",
|
|
"paperclip.runner-e2e.result/v2",
|
|
),
|
|
executionId: string,
|
|
suiteId: optional(string),
|
|
suiteDefinitionHash: optional(string),
|
|
source: optional(
|
|
shape({
|
|
sha: nullable(string),
|
|
ref: nullable(string),
|
|
workflowRunUrl: nullable(url),
|
|
}),
|
|
),
|
|
rankingSnapshot: optional(
|
|
shape({
|
|
snapshotId: string,
|
|
capturedAt: date,
|
|
sourceUrl: url,
|
|
rank: integer,
|
|
canonicalModelId: string,
|
|
}),
|
|
),
|
|
attempt: integer,
|
|
status: oneOf("passed", "failed"),
|
|
failureClass: optional(
|
|
oneOf(
|
|
"candidate_failure",
|
|
"provider_variance",
|
|
"transient_infrastructure",
|
|
"permanent_infrastructure",
|
|
"secret_leak",
|
|
"cleanup_failure",
|
|
),
|
|
),
|
|
error: optional(string),
|
|
profileId: string,
|
|
environmentId: oneOf("local", "daytona"),
|
|
caseId: string,
|
|
provider: string,
|
|
model: string,
|
|
runtimeMode: oneOf("legacy", "native"),
|
|
issueId: optional(string),
|
|
issueIdentifier: optional(nullable(string)),
|
|
runIds: optional(array(string)),
|
|
turnTimings: optional(
|
|
array(
|
|
shape({
|
|
turn: integer,
|
|
submittedAt: date,
|
|
runStartedAt: nullable(date),
|
|
runFinishedAt: nullable(date),
|
|
schedulerLatencyMs: nullable(number),
|
|
runDurationMs: nullable(number),
|
|
responseLatencyMs: nullable(number),
|
|
runId: string,
|
|
leaseAcquisitionOutcome: oneOf(
|
|
"created",
|
|
"resumed",
|
|
"replacement",
|
|
"unknown",
|
|
),
|
|
}),
|
|
),
|
|
),
|
|
startedAt: date,
|
|
finishedAt: date,
|
|
durationMs: number,
|
|
// Provider usage is intentionally opaque. Billing reads only finite numbers;
|
|
// the dashboard shows opaque diagnostics through escaped JSON.
|
|
usage: optional(nullable(object)),
|
|
runtimeUsage: optional(runtime),
|
|
billing: optional(billing),
|
|
matcherResults: optional(
|
|
array(shape({ matcher, passed: boolean, detail: string })),
|
|
),
|
|
screenshots: optional(
|
|
array(
|
|
shape({
|
|
id: string,
|
|
label: string,
|
|
file: relativeFile,
|
|
publication: optional(oneOf("public-runner-fixture")),
|
|
sha256: optional(string),
|
|
}),
|
|
),
|
|
),
|
|
cleanup: oneOf("not_started", "passed", "failed"),
|
|
} satisfies Record<keyof RunnerE2EResult, Rule>;
|
|
const result = shape(fields);
|
|
|
|
/** Validate every typed result field before upgrading, path construction or HTML rendering. */
|
|
export function validateRetainedRunnerResult(
|
|
value: unknown,
|
|
): asserts value is RunnerE2EResult {
|
|
result(value, "result");
|
|
}
|