mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
214 lines
6.3 KiB
TypeScript
214 lines
6.3 KiB
TypeScript
import { runnerMatrix } from "./catalog.js";
|
|
import type { MatrixExecution, MatrixJob } from "./types.js";
|
|
|
|
export interface RunnerSelectorOptions {
|
|
all: boolean;
|
|
list: boolean;
|
|
matrixJson: boolean;
|
|
ids: string[];
|
|
suites: string[];
|
|
groups: string[];
|
|
profiles: string[];
|
|
environments: string[];
|
|
cases: string[];
|
|
headed: boolean;
|
|
ui: boolean;
|
|
debug: boolean;
|
|
maxParallel: number;
|
|
}
|
|
|
|
export class RunnerSelectorError extends Error {}
|
|
|
|
function valueFor(args: string[], index: number, flag: string): string {
|
|
const value = args[index + 1];
|
|
if (!value || value.startsWith("--"))
|
|
throw new RunnerSelectorError(`${flag} requires a value`);
|
|
return value;
|
|
}
|
|
|
|
export function parseRunnerSelectors(
|
|
rawArgs: readonly string[],
|
|
): RunnerSelectorOptions {
|
|
const args = rawArgs.filter((value) => value !== "--");
|
|
const options: RunnerSelectorOptions = {
|
|
all: false,
|
|
list: false,
|
|
matrixJson: false,
|
|
ids: [],
|
|
suites: [],
|
|
groups: [],
|
|
profiles: [],
|
|
environments: [],
|
|
cases: [],
|
|
headed: false,
|
|
ui: false,
|
|
debug: false,
|
|
maxParallel: Number(process.env.PAPERCLIP_E2E_MAX_PARALLEL ?? "1"),
|
|
};
|
|
for (let index = 0; index < args.length; index += 1) {
|
|
const flag = args[index];
|
|
if (flag === "--all") options.all = true;
|
|
else if (flag === "--list") options.list = true;
|
|
else if (flag === "--matrix-json") options.matrixJson = true;
|
|
else if (flag === "--headed") options.headed = true;
|
|
else if (flag === "--ui") options.ui = true;
|
|
else if (flag === "--debug") options.debug = true;
|
|
else if (flag === "--max-parallel") {
|
|
const value = valueFor(args, index, flag);
|
|
index += 1;
|
|
options.maxParallel = Number(value);
|
|
} else if (
|
|
[
|
|
"--id",
|
|
"--suite",
|
|
"--group",
|
|
"--profile",
|
|
"--environment",
|
|
"--case",
|
|
].includes(flag)
|
|
) {
|
|
const value = valueFor(args, index, flag);
|
|
index += 1;
|
|
if (flag === "--id") options.ids.push(value);
|
|
else if (flag === "--suite") options.suites.push(value);
|
|
else if (flag === "--group") options.groups.push(value);
|
|
else if (flag === "--profile") options.profiles.push(value);
|
|
else if (flag === "--environment") options.environments.push(value);
|
|
else options.cases.push(value);
|
|
} else {
|
|
throw new RunnerSelectorError(`Unknown runner E2E argument: ${flag}`);
|
|
}
|
|
}
|
|
|
|
if (!Number.isInteger(options.maxParallel) || options.maxParallel < 1) {
|
|
throw new RunnerSelectorError("--max-parallel must be a positive integer");
|
|
}
|
|
|
|
const hasDimensions =
|
|
options.suites.length +
|
|
options.groups.length +
|
|
options.profiles.length +
|
|
options.environments.length +
|
|
options.cases.length >
|
|
0;
|
|
if (options.ids.length > 0 && (hasDimensions || options.all)) {
|
|
throw new RunnerSelectorError(
|
|
"--id is exclusive with --all and dimension filters",
|
|
);
|
|
}
|
|
if (options.all && hasDimensions) {
|
|
throw new RunnerSelectorError("--all is exclusive with dimension filters");
|
|
}
|
|
if (
|
|
!options.list &&
|
|
!options.matrixJson &&
|
|
!options.all &&
|
|
options.ids.length === 0 &&
|
|
!hasDimensions
|
|
) {
|
|
throw new RunnerSelectorError(
|
|
"Billable runner E2E runs require --all or an explicit selector",
|
|
);
|
|
}
|
|
return options;
|
|
}
|
|
|
|
function assertKnown(
|
|
label: string,
|
|
selected: readonly string[],
|
|
known: Set<string>,
|
|
) {
|
|
const unknown = selected.filter((value) => !known.has(value));
|
|
if (unknown.length > 0)
|
|
throw new RunnerSelectorError(`Unknown ${label}: ${unknown.join(", ")}`);
|
|
}
|
|
|
|
export function selectRunnerExecutions(
|
|
options: RunnerSelectorOptions,
|
|
matrix: readonly MatrixExecution[] = runnerMatrix,
|
|
): MatrixExecution[] {
|
|
const knownGroups = new Set(matrix.flatMap((execution) => execution.groups));
|
|
assertKnown(
|
|
"suite",
|
|
options.suites,
|
|
new Set(matrix.map((execution) => execution.suite.id)),
|
|
);
|
|
assertKnown("group", options.groups, knownGroups);
|
|
assertKnown(
|
|
"profile",
|
|
options.profiles,
|
|
new Set(matrix.map((execution) => execution.profile.id)),
|
|
);
|
|
assertKnown(
|
|
"environment",
|
|
options.environments,
|
|
new Set(matrix.map((execution) => execution.environment.id)),
|
|
);
|
|
assertKnown(
|
|
"case",
|
|
options.cases,
|
|
new Set(matrix.map((execution) => execution.task.id)),
|
|
);
|
|
assertKnown(
|
|
"execution id",
|
|
options.ids,
|
|
new Set(matrix.map((execution) => execution.id)),
|
|
);
|
|
|
|
const selected = matrix.filter((execution) => {
|
|
if (options.ids.length > 0) return options.ids.includes(execution.id);
|
|
if (execution.suite.manualOnly && !options.suites.includes(execution.suite.id) && !options.list) return false;
|
|
if (
|
|
options.all ||
|
|
(options.list &&
|
|
options.groups.length === 0 &&
|
|
options.suites.length === 0 &&
|
|
options.profiles.length === 0 &&
|
|
options.environments.length === 0 &&
|
|
options.cases.length === 0)
|
|
)
|
|
return true;
|
|
return (
|
|
(options.suites.length === 0 ||
|
|
options.suites.includes(execution.suite.id)) &&
|
|
options.groups.every((group) => execution.groups.includes(group)) &&
|
|
(options.profiles.length === 0 ||
|
|
options.profiles.includes(execution.profile.id)) &&
|
|
(options.environments.length === 0 ||
|
|
options.environments.includes(execution.environment.id)) &&
|
|
(options.cases.length === 0 || options.cases.includes(execution.task.id))
|
|
);
|
|
});
|
|
if (selected.length === 0)
|
|
throw new RunnerSelectorError(
|
|
"Runner E2E selectors matched zero executions",
|
|
);
|
|
return selected;
|
|
}
|
|
|
|
export function buildMatrixJobs(
|
|
executions: readonly MatrixExecution[],
|
|
): MatrixJob[] {
|
|
return executions
|
|
.map((execution) => ({
|
|
executionId: execution.id,
|
|
suiteId: execution.suite.id,
|
|
profileId: execution.profile.id,
|
|
credentialName: execution.profile.credential,
|
|
environmentId: execution.environment.id,
|
|
caseId: execution.task.id,
|
|
timeoutMinutes: Math.max(
|
|
execution.environment.id === "daytona" ? 40 : 25,
|
|
Math.ceil(
|
|
(2 *
|
|
(execution.task.attemptTimeoutMs[execution.environment.id] +
|
|
90_000) +
|
|
5 * 60_000) /
|
|
60_000,
|
|
),
|
|
),
|
|
needsDaytona: execution.environment.id === "daytona",
|
|
}))
|
|
.sort((left, right) => left.executionId.localeCompare(right.executionId));
|
|
}
|