Files
PaperClipAI/scripts/run-vitest-stable.mjs
DottaandPaperclip d0f69670db fix(runner): recover saved execution prompts after upgrades (#15518)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs save an immutable execution context for restart
recovery.
> - The context includes the prompt text, revision, and content hashes.
> - The parser required that saved prompt to match the current release.
> - A server upgrade could reject a valid saved run before provider
recovery.
> - This pull request validates and preserves the saved prompt snapshot.
> - Routine prompt changes no longer need a catalog of past strings.

## Linked Issues or Issue Description

**What happened?**

A hot restart selected a dead native runner for same-run recovery.
Reading its saved v5 execution input failed with
`input.runtimeContext.prompt must match the fixed Paperclip prompt
revision`. The new controller accepted only v6.

**Expected behavior**

Recovery uses the saved prompt and validates its content hashes. It
preserves the same run and provider session without starting a duplicate
turn.

**Steps to reproduce**

1. Start a native run and save its execution input and provider
checkpoint.
2. Change the fixed execution prompt in the server release.
3. Stop the runner and recover the saved run with the new controller.
4. Observe that the old parser rejects the saved prompt before provider
recovery.

**Paperclip version or commit**

The v5-to-v6 prompt change was introduced in #15446. The defect also
reproduces on current master before this fix.

**Deployment mode**

Source-built server with the native runner.

Related work: #15446 added task-monitor guidance. The held prompt-size
experiment in #15489 changes prompt wording but does not add recovery
compatibility.

## What Changed

- Read the prompt text and revision from the saved execution snapshot.
- Treat the revision as non-empty metadata and preserve the exact saved
bytes.
- Validate the prompt SHA-256 and the aggregate context digest.
- Keep fresh-run builders on the current prompt constants.
- Test arbitrary saved prompts, malformed fields, altered text, stale
hashes, and aggregate drift.
- Test recovery parsing for Codex input versions v3-v5 and OpenCode,
ACPX Pi, and Dot v6 inputs.
- Extend the real-process restart suite with both the incident's v5 wire
fixture and a prompt unknown to this release.
- Run the restart recovery suite in the existing Rust-equipped PR lane,
where its runner and fake-provider binaries are built. Verify complete,
non-overlapping test coverage for PR, release, and local callers.
- Document recovery from saved snapshots without a historical prompt
catalog.

## Verification

- Red: the new contract regressions fail against the catalog-based
parser with the original prompt-validation error.
- Green: 53 focused contract and materialization tests pass.
- Red: the real-process unknown-prompt regression fails with master's
original parser at the saved-input recovery read after process loss.
- Green: all 15 real-process restart tests pass locally on the final
branch. The saved-prompt cases keep the run and provider session,
replace the PID, and record one `turn/start`.
- Local repository `pnpm -r typecheck` and `pnpm build` passed after
rebase on `89f09dad723766e5351953f0731b9aa5daada28d`. The 53 focused
tests also passed on that head.
- Red: the new test-roster checks fail against the old CI placement.
- Green: all 26 test-scheduling checks pass after moving the restart
suite.
- CI ran all 15 restart recovery tests with no skips on final head
`6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`. [Runner test
job](https://github.com/paperclipai/paperclip/actions/runs/37764782334/job/113271464966).
- Greptile scored 5/5 on that exact head with no actionable findings.
- The complete CI matrix passed on final head
`6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`: general/workspace tests,
serialized server suites, both runner Vitest lanes, Rust and static
checks, all browser shards, typecheck, build, and the canary dry run.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/37764782334).
- There are no unresolved review threads or merge conflicts.
- Reproduce focused tests with `pnpm --filter
@paperclipai/paperclip-runner exec vitest run
src/contracts/runtime-context.test.ts
src/contracts/native-execution.test.ts
src/drivers/runtime-context-materializer.test.ts`.
- Reproduce restart tests with `pnpm --filter
@paperclipai/paperclip-runner build:rust` followed by `pnpm exec vitest
run
server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts`.
They use temporary PostgreSQL, real runner processes, and a fake Codex
provider. They do not use paid inference.

## Risks

- The parser now accepts internally consistent saved prompt text that is
absent from the current source. Inputs must come from trusted server
persistence. Content hashes verify consistency; they do not authenticate
authorship.
- Existing execution-schema, ownership, checkpoint, provider,
permission, and session-compatibility checks still apply.
- This change validates the saved base prompt. It does not make all
additional code-generated instruction strings versioned.
- The process-level recovery proof uses Codex. Other provider coverage
verifies the shared input parser and retained provider configuration.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code editing,
and test tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 05:57:20 -05:00

629 lines
24 KiB
JavaScript

#!/usr/bin/env node
import { spawnSync } from "node:child_process";
import { mkdirSync, mkdtempSync, readdirSync, readFileSync, realpathSync, statSync } from "node:fs";
import os from "node:os";
import path from "node:path";
import { fileURLToPath } from "node:url";
import { loadShardDurations, selectGeneralServerShard } from "./general-server-shard.mjs";
import { assertSelectedTests, partitionTestLines } from "./test-line-shard.mjs";
const repoRoot = process.cwd();
const scriptsDir = path.dirname(fileURLToPath(import.meta.url));
const generalServerShardDurations = loadShardDurations(
path.join(scriptsDir, "general-server-shard-durations.json"),
);
const serializedShardDurations = loadShardDurations(
path.join(scriptsDir, "serialized-shard-durations.json"),
);
const serverRoot = path.join(repoRoot, "server");
const serverSrcDir = path.join(repoRoot, "server", "src");
const serverTestsDir = path.join(repoRoot, "server", "src", "__tests__");
const serverScriptsDir = path.join(repoRoot, "server", "scripts");
const nonServerProjects = [
"@paperclipai/shared",
"@paperclipai/skills-catalog",
"@paperclipai/db",
"@paperclipai/adapter-utils",
"@paperclipai/adapter-claude-local",
"@paperclipai/adapter-codex-local",
"@paperclipai/adapter-cursor-cloud",
"@paperclipai/adapter-cursor-local",
"@paperclipai/adapter-gemini-local",
"@paperclipai/adapter-grok-local",
"@paperclipai/hermes-paperclip-adapter",
"@paperclipai/adapter-kimi-local",
"@paperclipai/adapter-openclaw-gateway",
"@paperclipai/adapter-opencode-local",
"@paperclipai/adapter-pi-local",
"@paperclipai/plugin-daytona",
"@paperclipai/plugin-sdk",
"@paperclipai/create-paperclip-plugin",
"@paperclipai/ui",
"paperclipai",
];
const routeTestPattern = /[^/]*(?:route|routes|authz)[^/]*\.test\.ts$/;
const additionalSerializedServerTests = new Set([
"server/src/__tests__/approval-routes-idempotency.test.ts",
"server/src/__tests__/assets.test.ts",
"server/src/__tests__/authz-company-access.test.ts",
"server/src/__tests__/companies-route-path-guard.test.ts",
"server/src/__tests__/company-portability.test.ts",
"server/src/__tests__/costs-service.test.ts",
"server/src/__tests__/express5-auth-wildcard.test.ts",
"server/src/__tests__/health-dev-server-token.test.ts",
"server/src/__tests__/health.test.ts",
"server/src/__tests__/heartbeat-dependency-scheduling.test.ts",
"server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts",
"server/src/__tests__/heartbeat-process-recovery.test.ts",
"server/src/__tests__/invite-accept-existing-member.test.ts",
"server/src/__tests__/invite-accept-gateway-defaults.test.ts",
"server/src/__tests__/invite-accept-replay.test.ts",
"server/src/__tests__/invite-expiry.test.ts",
"server/src/__tests__/invite-join-manager.test.ts",
"server/src/__tests__/invite-onboarding-text.test.ts",
"server/src/__tests__/invite-url-public-base-url.test.ts",
"server/src/__tests__/issues-checkout-wakeup.test.ts",
"server/src/__tests__/issues-service.test.ts",
"server/src/__tests__/opencode-local-adapter-environment.test.ts",
"server/src/__tests__/project-routes-env.test.ts",
"server/src/__tests__/redaction.test.ts",
"server/src/__tests__/routines-e2e.test.ts",
]);
let invocationIndex = 0;
const serializedModeName = "serialized";
const generalModeName = "general";
const allModeName = "all";
const generalServerGroupName = "general-server";
const generalServerWithoutChatGroupName = "general-server-without-chat";
const generalChatGroupName = "general-chat";
const generalServerNativeRunnerGroupName = "general-server-native-runner";
const chatSuite = "server/src/__tests__/chat-channels.integration.test.ts";
// The first suite rebuilds the Runner release binaries with cargo in beforeAll.
// Inside the PR workflow's plain server shards, which carry no Rust cache,
// that build was a ~4m30s cold compile of every third-party crate on each run
// (277s of a 291s shard vitest step, actions run 35246999382, 2026-09-17).
const nativeRunnerSuites = [
"server/src/services/native-runtime/native-codex-runner.integration.test.ts",
"server/src/__tests__/dot-runner.test.ts",
// These process cases otherwise skip in server shards without Runner binaries.
"server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts",
];
// In the PR workflow (pr.yml, the caller of pr-trusted.yml — reusable
// workflows inherit the caller's GITHUB_WORKFLOW), the last Verify Paperclip
// Runner vitest shard runs the native-runner group instead, because those
// lanes restore the shared release-runner-v1 Rust cache (see
// packages/paperclip-runner/scripts/run-pr-vitest-lane.mjs). Every other
// caller — local runs, release-verify.yml under the Release and Cloud
// readiness workflows — keeps the suite in the server shards, so a renamed or
// unknown workflow degrades to today's slower-but-covered behavior rather
// than dropping the suite.
const prWorkflowName = "PR";
const nativeRunnerSuiteRunsInRustCachedLane = process.env.GITHUB_WORKFLOW === prWorkflowName;
const withoutChatExcludedSuites = nativeRunnerSuiteRunsInRustCachedLane
? [chatSuite, ...nativeRunnerSuites]
: [chatSuite];
const generalWorkspacesAGroupName = "general-workspaces-a";
const generalWorkspacesBGroupName = "general-workspaces-b";
const generalWorkspacesAProjects = ["@paperclipai/ui", "paperclipai"];
const generalWorkspacesBProjects = nonServerProjects.filter((project) => !generalWorkspacesAProjects.includes(project));
const generalGroupNames = [generalServerGroupName, generalWorkspacesAGroupName, generalWorkspacesBGroupName];
const allowedGeneralGroupNames = [
...generalGroupNames,
generalServerWithoutChatGroupName,
generalChatGroupName,
generalServerNativeRunnerGroupName,
];
const serializedServerVitestArgs = [
"--no-file-parallelism",
"--maxWorkers=1",
];
const sourceOnlyVitestArgs = ["--exclude", "**/dist/**"];
function walk(dir) {
const entries = readdirSync(dir);
const files = [];
for (const entry of entries) {
const absolute = path.join(dir, entry);
const stats = statSync(absolute);
if (stats.isDirectory()) {
files.push(...walk(absolute));
} else if (stats.isFile()) {
files.push(absolute);
}
}
return files;
}
function toRepoPath(file) {
return path.relative(repoRoot, file).split(path.sep).join("/");
}
function toServerPath(file) {
return path.relative(serverRoot, file).split(path.sep).join("/");
}
// For a multi-line `it.each([...])("name", fn)` call, Vitest 5's `vitest
// list --includeTaskLocation` and `vitest run <file>:<line>` disagree about
// which line registers the test: `list` reports (and only matches) the line
// of the wrapped "name" argument, while `run` only matches the line above
// it, where Prettier's formatting puts that call's opening `(`. Separately,
// on this file `list` also occasionally invents an entry whose "location"
// sits on an ordinary body statement that merely ends in `)(` or starts a
// line with a quote, with no `it`/`test` call anywhere nearby — a bogus
// artifact of the same collector, not a real test. Both only showed up on
// server/src/__tests__/chat-channels.integration.test.ts, a 70,000+ line
// fixture; smaller files have not reproduced either issue.
//
// chatSuiteRunLine classifies a `list`-reported line against the actual
// source and returns the line `vitest run` accepts, or null if the line
// does not belong to any real `it`/`test` call (a bogus entry to drop).
const callAnywherePattern = /(?:^|[^\w$])(?:it|test)(?:\.(?:each|skip|only))?\(/;
const inlineCallPattern = /\)\(\s*["'`]/;
function chatSuiteRunLine(sourceLines, line) {
const text = (sourceLines[line - 1] ?? "").trim();
if (callAnywherePattern.test(text) || text.endsWith(")(") || inlineCallPattern.test(text)) {
return line;
}
const previous = (sourceLines[line - 2] ?? "").trim();
const isNameArgument = text.startsWith('"') || text.startsWith("'") || text.startsWith("`");
if (isNameArgument && previous.endsWith(")(")) {
return line - 1;
}
return null;
}
function isRouteOrAuthzTest(file) {
if (routeTestPattern.test(file)) {
return true;
}
return additionalSerializedServerTests.has(file);
}
function fail(message) {
console.error(`[test:run] ${message}`);
process.exit(1);
}
function readOptionValue(argv, index, argName) {
const value = argv[index + 1];
if (value === undefined) {
fail(`Missing value for ${argName}`);
}
return value;
}
function parseNonNegativeInteger(value, argName) {
const parsed = Number(value);
if (value.trim() === "" || !Number.isInteger(parsed) || parsed < 0) {
fail(`${argName} must be a non-negative integer. Received "${value}".`);
}
return parsed;
}
function parsePositiveInteger(value, argName) {
const parsed = Number(value);
if (value.trim() === "" || !Number.isInteger(parsed) || parsed < 1) {
fail(`${argName} must be a positive integer. Received "${value}".`);
}
return parsed;
}
function parseCliOptions(argv) {
let mode = allModeName;
let shardIndex = null;
let shardCount = null;
let group = null;
let dryRun = false;
for (let index = 0; index < argv.length; index += 1) {
const arg = argv[index];
if (arg === "--") {
continue;
}
if (arg === "--mode") {
mode = readOptionValue(argv, index, arg);
index += 1;
continue;
}
if (arg.startsWith("--mode=")) {
mode = arg.slice("--mode=".length);
continue;
}
if (arg === "--shard-index") {
shardIndex = parseNonNegativeInteger(readOptionValue(argv, index, arg), arg);
index += 1;
continue;
}
if (arg.startsWith("--shard-index=")) {
shardIndex = parseNonNegativeInteger(arg.slice("--shard-index=".length), "--shard-index");
continue;
}
if (arg === "--shard-count") {
shardCount = parsePositiveInteger(readOptionValue(argv, index, arg), arg);
index += 1;
continue;
}
if (arg.startsWith("--shard-count=")) {
shardCount = parsePositiveInteger(arg.slice("--shard-count=".length), "--shard-count");
continue;
}
if (arg === "--dry-run") {
dryRun = true;
continue;
}
if (arg === "--group") {
group = readOptionValue(argv, index, arg);
index += 1;
continue;
}
if (arg.startsWith("--group=")) {
group = arg.slice("--group=".length);
continue;
}
fail(`Unknown argument "${arg}".`);
}
if (!new Set([allModeName, generalModeName, serializedModeName]).has(mode)) {
fail(`Unknown mode "${mode}". Expected one of: ${allModeName}, ${generalModeName}, ${serializedModeName}.`);
}
if ((shardIndex === null) !== (shardCount === null)) {
fail("--shard-index and --shard-count must be provided together.");
}
const shardAllowed =
mode === serializedModeName ||
(mode === generalModeName &&
([generalServerGroupName, generalServerWithoutChatGroupName, generalChatGroupName, generalWorkspacesAGroupName].includes(group)));
if (!shardAllowed && shardIndex !== null) {
fail(
"--shard-index/--shard-count are only valid with serialized mode or a shardable general server/chat/workspaces-a group.",
);
}
if (group !== null && mode !== generalModeName) {
fail("--group is only valid with --mode general.");
}
if (group !== null && !allowedGeneralGroupNames.includes(group)) {
fail(`Unknown group "${group}". Expected one of: ${allowedGeneralGroupNames.join(", ")}.`);
}
if (shardIndex !== null) {
if (shardIndex >= shardCount) {
fail(`--shard-index must be less than --shard-count. Received ${shardIndex} of ${shardCount}.`);
}
}
if (mode === serializedModeName) {
return {
mode,
shardIndex: shardIndex ?? 0,
shardCount: shardCount ?? 1,
group: null,
dryRun,
};
}
return {
mode,
shardIndex,
shardCount,
group,
dryRun,
};
}
function selectSerializedSuites(routeTests, shardIndex, shardCount) {
// Same duration-aware LPT partition as the general-server lane. Round-robin
// over the alphabetical list clustered the heavy heartbeat/issues suites on
// one shard (291s vs 170-201s test steps across the matrix in actions run
// 32012408876), which made that shard the whole PR run's slowest check.
const byRepoPath = new Map(routeTests.map((routeTest) => [routeTest.repoPath, routeTest]));
const shardFiles = selectGeneralServerShard(
routeTests.map((routeTest) => routeTest.repoPath),
shardIndex,
shardCount,
serializedShardDurations,
);
return shardFiles.map((file) => byRepoPath.get(file));
}
function runVitest(args, label, testShard = null) {
console.log(`\n[test:run] ${label}`);
invocationIndex += 1;
const tempRootParent = process.platform === "win32" ? os.tmpdir() : "/tmp";
// Production workspace/security checks reject symlink aliases. In particular
// /tmp is /private/tmp on macOS, so fixture roots must use the canonical path.
const testRoot = realpathSync(mkdtempSync(path.join(tempRootParent, "pv-")));
// Keep per-run paths compact so Unix socket fixtures stay under macOS path limits.
const env = {
...process.env,
NODE_ENV: "test",
PAPERCLIP_TEST_HOST_HOME: process.env.PAPERCLIP_TEST_HOST_HOME
?? (process.env.PAPERCLIP_HOME?.trim() || path.join(os.homedir(), ".paperclip")),
PAPERCLIP_HOME: path.join(testRoot, "h"),
// Config discovery otherwise prefers the checkout's .paperclip/config.json
// over PAPERCLIP_HOME, importing preview scheduling policy into unit tests.
PAPERCLIP_CONFIG: path.join(testRoot, "h", "config.json"),
PAPERCLIP_INSTANCE_ID: `vt-${process.pid}-${invocationIndex}`,
TMPDIR: path.join(testRoot, "t"),
};
mkdirSync(env.PAPERCLIP_HOME, { recursive: true });
mkdirSync(env.TMPDIR, { recursive: true });
if (testShard) {
const file = path.resolve(repoRoot, chatSuite);
const sourceLines = readFileSync(file, "utf8").split("\n");
const collect = (filters, name) => {
const output = path.join(testRoot, `${name}.json`);
const result = spawnSync("pnpm", ["exec", "vitest", "list", ...sourceOnlyVitestArgs,
...filters, "--allowOnly=false", "--includeTaskLocation", `--json=${output}`], {
cwd: repoRoot, env, stdio: "inherit",
});
if (result.error || result.status !== 0) fail(`Vitest collection failed: ${result.error?.message ?? result.status}`);
return JSON.parse(readFileSync(output, "utf8"));
};
const allCollected = collect(args, "all");
const collected = [];
const bogusCollected = [];
for (const test of allCollected) {
(chatSuiteRunLine(sourceLines, test.location.line) === null ? bogusCollected : collected).push(test);
}
if (bogusCollected.length > 0) {
console.warn(
`[test:run] dropping ${bogusCollected.length} bogus Vitest collection entr${bogusCollected.length === 1 ? "y" : "ies"} with no matching it()/test() call: ${bogusCollected.map((test) => `${test.location.line}:${JSON.stringify(test.name)}`).join(", ")}`,
);
}
const selected = partitionTestLines(collected, testShard.count, file)[testShard.index];
// Validate shard membership with the same `list`-reported lines that
// produced `selected`, then switch to the `run`-compatible lines only for
// the filters handed to the real `vitest run` below (see
// chatSuiteRunLine).
const validationFilters = selected.lines.map((line) => `${chatSuite}:${line}`);
const validationArgs = [...args.filter((arg) => arg !== chatSuite), ...validationFilters];
assertSelectedTests(selected.tests, collect(validationArgs, "selected"), file);
console.log(`[test:run] chat shard ${testShard.index + 1}/${testShard.count}: ${selected.tests.length}/${collected.length} tests, ${selected.lines.length} source lines; exact filter coverage verified`);
const runFilters = selected.lines.map((line) => `${chatSuite}:${chatSuiteRunLine(sourceLines, line)}`);
args = [...args.filter((arg) => arg !== chatSuite), ...runFilters, "--allowOnly=false"];
}
const result = spawnSync("pnpm", ["exec", "vitest", "run", ...sourceOnlyVitestArgs, ...args], {
cwd: repoRoot,
env,
stdio: "inherit",
});
if (result.error) {
console.error(`[test:run] Failed to start Vitest: ${result.error.message}`);
process.exit(1);
}
if (result.status !== 0) {
process.exit(result.status ?? 1);
}
}
function runGeneralSuites(routeTests) {
for (const groupName of generalGroupNames) {
runGeneralGroup(routeTests, groupName);
}
}
function runProjectGroup(projects, groupName, shardIndex = null, shardCount = null) {
// With shard args, lean on Vitest's native --shard: each matrix job runs the
// same per-project invocations but only its slice of each project's test
// files. Vitest's sharding is deterministic for an identical file list, so
// the matrix jobs form a complete, non-overlapping cover of every project.
const shardArgs =
shardCount !== null && shardCount > 1 ? [`--shard=${shardIndex + 1}/${shardCount}`] : [];
const shardSuffix = shardArgs.length > 0 ? ` shard ${shardIndex + 1}/${shardCount}` : "";
for (const project of projects) {
runVitest(["--project", project, ...shardArgs], `${groupName} project ${project}${shardSuffix}`);
}
}
function runGeneralGroup(routeTests, groupName, shardIndex = null, shardCount = null) {
if (groupName === generalChatGroupName) {
runVitest(["--project", "@paperclipai/server", ...serializedServerVitestArgs, chatSuite],
"chat integration test shard", { index: shardIndex ?? 0, count: shardCount ?? 1 });
return;
}
if (groupName === generalServerNativeRunnerGroupName) {
runVitest(
["--project", "@paperclipai/server", ...serializedServerVitestArgs, ...nativeRunnerSuites],
"native runner vertical-slice suite",
);
return;
}
if (groupName === generalServerGroupName || groupName === generalServerWithoutChatGroupName) {
// In the PR workflow the without-chat group also leaves the native-runner
// suite to the Rust-cached vitest lane; the full general-server group
// (local runs) keeps both.
const withoutChat = groupName === generalServerWithoutChatGroupName;
const files = withoutChat
? generalServerTestFiles.filter((file) => !withoutChatExcludedSuites.includes(file))
: generalServerTestFiles;
if (shardCount !== null && shardCount > 1) {
const shardFiles = selectGeneralServerShard(
files,
shardIndex,
shardCount,
generalServerShardDurations,
);
console.log(
`\n[test:run] general-server shard ${shardIndex + 1}/${shardCount} running ${shardFiles.length} of ${files.length} suites`,
);
if (shardFiles.length === 0) {
return;
}
runVitest(
[
"--project",
"@paperclipai/server",
...serializedServerVitestArgs,
...shardFiles,
],
`${groupName} shard ${shardIndex + 1}/${shardCount}`,
);
return;
}
const excludeRouteArgs = routeTests.flatMap((file) => ["--exclude", file.serverPath]);
if (withoutChat) {
for (const suite of withoutChatExcludedSuites) {
excludeRouteArgs.push("--exclude", suite.replace(/^server\//, ""));
}
}
runVitest(
[
"--project",
"@paperclipai/server",
...serializedServerVitestArgs,
...excludeRouteArgs,
],
`${groupName} server suites excluding ${routeTests.length} serialized suites`,
);
return;
}
if (groupName === generalWorkspacesAGroupName) {
// The ui project dominates this lane (~224s of a 319s job in actions run
// 31371439296, 2026-08-10, where workspaces-a was the slowest PR check).
// Its 439 test files shard cleanly with Vitest's native --shard, so the
// lane splits across runners without a duration manifest.
runProjectGroup(generalWorkspacesAProjects, groupName, shardIndex, shardCount);
return;
}
if (groupName === generalWorkspacesBGroupName) {
runProjectGroup(generalWorkspacesBProjects, groupName);
return;
}
fail(`Unknown group "${groupName}".`);
}
function runSerializedSuites(routeTests, shardIndex, shardCount) {
const shardTests = selectSerializedSuites(routeTests, shardIndex, shardCount);
console.log(
`\n[test:run] serialized shard ${shardIndex + 1}/${shardCount} running ${shardTests.length} of ${routeTests.length} suites`,
);
for (const routeTest of shardTests) {
runVitest(
[
"--project",
"@paperclipai/server",
routeTest.repoPath,
"--pool=forks",
"--isolate",
],
routeTest.repoPath,
);
}
}
const routeTests = walk(serverTestsDir)
.filter((file) => isRouteOrAuthzTest(toRepoPath(file)))
.map((file) => ({
repoPath: toRepoPath(file),
serverPath: toServerPath(file),
}))
.sort((a, b) => a.repoPath.localeCompare(b.repoPath));
// Every server test file that the general-server group is responsible for,
// i.e. the whole server project minus the route/authz suites that run in the
// dedicated serialized shards. Sharding this list across runners is what keeps
// the general-server lane from becoming the PR critical path: the server vitest
// config pins maxWorkers to 1, so the only way to parallelize is across jobs.
// Suites are partitioned by recorded duration (scripts/general-server-shard.mjs)
// rather than round-robin, so one slow suite cluster can't stretch a single shard.
const serializedRepoPaths = new Set(routeTests.map(test => test.repoPath));
const generalServerTestFiles = [
...walk(serverSrcDir).filter(file => file.endsWith(".test.ts")),
...walk(serverScriptsDir).filter(file => file.endsWith(".test.mjs")),
]
.map((file) => toRepoPath(file))
// Only exclude suites actually assigned to the serialized lane. A name
// such as services/openrouter-models.test.ts is not a serialized route.
.filter((repoPath) => !serializedRepoPaths.has(repoPath))
.sort((a, b) => a.localeCompare(b));
const options = parseCliOptions(process.argv.slice(2));
if (options.dryRun) {
const serializedSuites =
options.mode === serializedModeName
? selectSerializedSuites(routeTests, options.shardIndex, options.shardCount)
: routeTests;
console.log(
JSON.stringify(
{
mode: options.mode,
shardIndex: options.shardIndex,
shardCount: options.shardCount,
group: options.group,
availableGeneralGroups: allowedGeneralGroupNames,
serializedSuiteCount: routeTests.length,
selectedSerializedSuites: serializedSuites.map((routeTest) => routeTest.repoPath),
generalServerSuiteCount: generalServerTestFiles.length,
selectedGeneralServerSuites:
options.mode === generalModeName && options.group === generalServerNativeRunnerGroupName
? nativeRunnerSuites
: options.mode === generalModeName &&
[generalServerGroupName, generalServerWithoutChatGroupName].includes(options.group) &&
options.shardCount !== null
? selectGeneralServerShard(
options.group === generalServerWithoutChatGroupName
? generalServerTestFiles.filter((file) => !withoutChatExcludedSuites.includes(file))
: generalServerTestFiles,
options.shardIndex,
options.shardCount,
generalServerShardDurations,
)
: null,
workspaceProjects:
options.group === generalWorkspacesAGroupName
? generalWorkspacesAProjects
: options.group === generalWorkspacesBGroupName
? generalWorkspacesBProjects
: null,
workspacesVitestShard:
options.group === generalWorkspacesAGroupName &&
options.shardCount !== null &&
options.shardCount > 1
? `${options.shardIndex + 1}/${options.shardCount}`
: null,
},
null,
2,
),
);
process.exit(0);
}
if (options.mode === generalModeName || options.mode === allModeName) {
if (options.group) {
runGeneralGroup(routeTests, options.group, options.shardIndex, options.shardCount);
} else {
runGeneralSuites(routeTests);
}
}
if (options.mode === serializedModeName || options.mode === allModeName) {
runSerializedSuites(routeTests, options.shardIndex ?? 0, options.shardCount ?? 1);
}