mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
227 lines
9.8 KiB
JavaScript
227 lines
9.8 KiB
JavaScript
import assert from "node:assert/strict";
|
|
import { spawnSync } from "node:child_process";
|
|
import path from "node:path";
|
|
import { fileURLToPath } from "node:url";
|
|
import test from "node:test";
|
|
|
|
import {
|
|
defaultSuiteWeight,
|
|
loadShardDurations,
|
|
partitionGeneralServerSuites,
|
|
} from "../general-server-shard.mjs";
|
|
|
|
const repoRoot = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", "..");
|
|
const script = path.join(repoRoot, "scripts", "run-vitest-stable.mjs");
|
|
const durationsManifest = path.join(repoRoot, "scripts", "general-server-shard-durations.json");
|
|
const serializedDurationsManifest = path.join(
|
|
repoRoot,
|
|
"scripts",
|
|
"serialized-shard-durations.json",
|
|
);
|
|
|
|
function dryRun(args) {
|
|
const result = spawnSync(process.execPath, [script, ...args, "--dry-run"], {
|
|
cwd: repoRoot,
|
|
encoding: "utf8",
|
|
});
|
|
return result;
|
|
}
|
|
|
|
function dryRunJson(args) {
|
|
const result = dryRun(args);
|
|
assert.equal(result.status, 0, `expected success for ${args.join(" ")}: ${result.stderr}`);
|
|
return JSON.parse(result.stdout);
|
|
}
|
|
|
|
const SHARD_COUNT = 5;
|
|
const SERIALIZED_SHARD_COUNT = 5;
|
|
|
|
|
|
test("the serialized shards form a complete, non-overlapping partition", () => {
|
|
const shards = Array.from({ length: SERIALIZED_SHARD_COUNT }, (_, index) =>
|
|
dryRunJson(["--mode", "serialized", "--shard-index", String(index), "--shard-count", String(SERIALIZED_SHARD_COUNT)]),
|
|
);
|
|
|
|
const total = shards[0].serializedSuiteCount;
|
|
const selected = shards.flatMap((shard) => shard.selectedSerializedSuites);
|
|
assert.equal(selected.length, total, "every serialized suite must be selected exactly once");
|
|
assert.equal(new Set(selected).size, total, "serialized shards must not overlap");
|
|
});
|
|
|
|
test("the general-server shards form a complete, non-overlapping partition", () => {
|
|
const shards = Array.from({ length: SHARD_COUNT }, (_, index) =>
|
|
dryRunJson(["--mode", "general", "--group", "general-server", "--shard-index", String(index), "--shard-count", String(SHARD_COUNT)]),
|
|
);
|
|
|
|
const total = shards[0].generalServerSuiteCount;
|
|
assert.ok(total > 0, "expected a non-empty general-server suite set");
|
|
|
|
const seen = new Set();
|
|
let selectedTotal = 0;
|
|
for (const shard of shards) {
|
|
assert.equal(shard.generalServerSuiteCount, total, "suite count must be stable across shards");
|
|
for (const file of shard.selectedGeneralServerSuites) {
|
|
assert.ok(!seen.has(file), `suite assigned to more than one shard: ${file}`);
|
|
seen.add(file);
|
|
selectedTotal += 1;
|
|
}
|
|
}
|
|
|
|
// Every suite runs exactly once: union covers the whole set with no overlap.
|
|
assert.equal(selectedTotal, total, "every suite must be selected exactly once");
|
|
assert.equal(seen.size, total, "union of shards must cover the whole suite set");
|
|
});
|
|
|
|
test("a route/authz suite never leaks into the general-server shards", () => {
|
|
const shard = dryRunJson(["--mode", "general", "--group", "general-server", "--shard-index", "0", "--shard-count", SHARD_COUNT.toString()]);
|
|
for (const file of shard.selectedGeneralServerSuites) {
|
|
assert.ok(
|
|
!/[^/]*(?:route|routes|authz)[^/]*\.test\.ts$/.test(file),
|
|
`route/authz suite must stay in the serialized lane, not general-server: ${file}`,
|
|
);
|
|
}
|
|
});
|
|
|
|
test("shard flags are rejected for the workspaces-b group", () => {
|
|
const result = dryRun(["--mode", "general", "--group", "general-workspaces-b", "--shard-index", "0", "--shard-count", "3"]);
|
|
assert.notEqual(result.status, 0, "workspaces-b must not accept shard flags");
|
|
});
|
|
|
|
test("workspaces-a shards map to Vitest native --shard slices over a stable project list", () => {
|
|
const shards = [0, 1].map((index) =>
|
|
dryRunJson([
|
|
"--mode", "general", "--group", "general-workspaces-a",
|
|
"--shard-index", String(index), "--shard-count", "2",
|
|
]),
|
|
);
|
|
|
|
assert.deepEqual(
|
|
shards.map((shard) => shard.workspacesVitestShard),
|
|
["1/2", "2/2"],
|
|
"each matrix job must pass its own --shard slice to vitest",
|
|
);
|
|
// Vitest's --shard partitions each project's file list deterministically, so
|
|
// an identical project list across jobs is what guarantees complete,
|
|
// non-overlapping coverage of the lane.
|
|
assert.deepEqual(shards[0].workspaceProjects, shards[1].workspaceProjects);
|
|
assert.ok(shards[0].workspaceProjects.length > 0, "workspaces-a must run at least one project");
|
|
|
|
const unsharded = dryRunJson(["--mode", "general", "--group", "general-workspaces-a"]);
|
|
assert.deepEqual(
|
|
unsharded.workspaceProjects,
|
|
shards[0].workspaceProjects,
|
|
"sharding must not change which projects the lane covers",
|
|
);
|
|
assert.equal(unsharded.workspacesVitestShard, null);
|
|
});
|
|
|
|
test("duration-aware partition balances skewed weights better than round-robin", () => {
|
|
// Round-robin puts all three heavy suites on shard 0 (indexes 0, 3, 6).
|
|
const files = ["a", "b", "c", "d", "e", "f", "g", "h", "i"];
|
|
const durations = { a: 30000, d: 30000, g: 30000, b: 100, c: 100, e: 100, f: 100, h: 100, i: 100 };
|
|
|
|
const shards = partitionGeneralServerSuites(files, 3, durations);
|
|
const totals = shards.map((shard) => shard.totalWeight);
|
|
const maxTotal = Math.max(...totals);
|
|
const minTotal = Math.min(...totals);
|
|
assert.ok(
|
|
maxTotal - minTotal <= 200,
|
|
`expected near-even shard weights, got ${totals.join(", ")}`,
|
|
);
|
|
assert.equal(
|
|
shards.flatMap((shard) => shard.files).sort().join(","),
|
|
files.join(","),
|
|
"partition must cover every file exactly once",
|
|
);
|
|
});
|
|
|
|
test("the partition is deterministic for identical inputs", () => {
|
|
const files = Array.from({ length: 50 }, (_, index) => `suite-${index}.test.ts`);
|
|
const durations = Object.fromEntries(files.map((file, index) => [file, (index * 37) % 5000]));
|
|
|
|
const first = partitionGeneralServerSuites(files, 3, durations);
|
|
const second = partitionGeneralServerSuites(files, 3, durations);
|
|
assert.deepEqual(first, second, "same inputs must always produce the same partition");
|
|
});
|
|
|
|
test("suites missing from the manifest get the median weight", () => {
|
|
assert.equal(defaultSuiteWeight({ a: 100, b: 300, c: 900 }), 300);
|
|
assert.equal(defaultSuiteWeight({ a: 100, b: 300, c: 500, d: 900 }), 400);
|
|
assert.equal(defaultSuiteWeight({}), 1000, "empty manifest falls back to a fixed weight");
|
|
});
|
|
|
|
test("a missing or malformed manifest degrades to uniform weights", () => {
|
|
assert.deepEqual(loadShardDurations(path.join(repoRoot, "scripts", "no-such-manifest.json")), {});
|
|
|
|
const files = ["a", "b", "c", "d"];
|
|
const shards = partitionGeneralServerSuites(files, 2, {});
|
|
assert.equal(shards[0].files.length + shards[1].files.length, files.length);
|
|
assert.equal(Math.abs(shards[0].files.length - shards[1].files.length), 0);
|
|
});
|
|
|
|
test("the checked-in manifest loads and covers most of the current suite set", () => {
|
|
const durations = loadShardDurations(durationsManifest);
|
|
assert.ok(Object.keys(durations).length > 0, "manifest must parse to a non-empty duration map");
|
|
|
|
const shard = dryRunJson(["--mode", "general", "--group", "general-server", "--shard-index", "0", "--shard-count", "1"]);
|
|
const currentFiles = shard.selectedGeneralServerSuites;
|
|
const known = currentFiles.filter((file) => durations[file] !== undefined).length;
|
|
assert.ok(
|
|
known / currentFiles.length >= 0.5,
|
|
`manifest is stale: only ${known} of ${currentFiles.length} suites have recorded durations — regenerate it from a recent PR run (see the manifest's $comment)`,
|
|
);
|
|
});
|
|
|
|
test("the checked-in serialized manifest loads and covers most of the current suite set", () => {
|
|
const durations = loadShardDurations(serializedDurationsManifest);
|
|
assert.ok(Object.keys(durations).length > 0, "manifest must parse to a non-empty duration map");
|
|
|
|
const shard = dryRunJson(["--mode", "serialized", "--shard-index", "0", "--shard-count", "1"]);
|
|
const currentFiles = shard.selectedSerializedSuites;
|
|
const known = currentFiles.filter((file) => durations[file] !== undefined).length;
|
|
assert.ok(
|
|
known / currentFiles.length >= 0.5,
|
|
`manifest is stale: only ${known} of ${currentFiles.length} suites have recorded durations — regenerate it from a recent PR run (see the manifest's $comment)`,
|
|
);
|
|
});
|
|
|
|
test("the real serialized shard partition is duration-balanced", () => {
|
|
const durations = loadShardDurations(serializedDurationsManifest);
|
|
const fallback = defaultSuiteWeight(durations);
|
|
const shards = Array.from({ length: SERIALIZED_SHARD_COUNT }, (_, index) =>
|
|
dryRunJson(["--mode", "serialized", "--shard-index", String(index), "--shard-count", String(SERIALIZED_SHARD_COUNT)]),
|
|
);
|
|
|
|
const totals = shards.map((shard) =>
|
|
shard.selectedSerializedSuites.reduce((sum, file) => sum + (durations[file] ?? fallback), 0),
|
|
);
|
|
const maxTotal = Math.max(...totals);
|
|
const minTotal = Math.min(...totals);
|
|
// LPT keeps the spread within the heaviest single suite; use that as the bound.
|
|
const heaviest = Math.max(...Object.values(durations));
|
|
assert.ok(
|
|
maxTotal - minTotal <= heaviest,
|
|
`serialized shard weight spread ${maxTotal - minTotal}ms exceeds heaviest suite ${heaviest}ms: ${totals.join(", ")}`,
|
|
);
|
|
});
|
|
|
|
test("the real shard partition is duration-balanced", () => {
|
|
const durations = loadShardDurations(durationsManifest);
|
|
const fallback = defaultSuiteWeight(durations);
|
|
const shards = Array.from({ length: SHARD_COUNT }, (_, index) =>
|
|
dryRunJson(["--mode", "general", "--group", "general-server", "--shard-index", String(index), "--shard-count", String(SHARD_COUNT)]),
|
|
);
|
|
|
|
const totals = shards.map((shard) =>
|
|
shard.selectedGeneralServerSuites.reduce((sum, file) => sum + (durations[file] ?? fallback), 0),
|
|
);
|
|
const maxTotal = Math.max(...totals);
|
|
const minTotal = Math.min(...totals);
|
|
// LPT keeps the spread within the heaviest single suite; use that as the bound.
|
|
const heaviest = Math.max(...Object.values(durations));
|
|
assert.ok(
|
|
maxTotal - minTotal <= heaviest,
|
|
`shard weight spread ${maxTotal - minTotal}ms exceeds heaviest suite ${heaviest}ms: ${totals.join(", ")}`,
|
|
);
|
|
});
|