mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip uses browser tests to protect critical operator flows. > - The trusted pull request workflow runs the E2E catalog on three existing runners. > - Smoke Lab was one 168-second spec, so the shard scheduler could not divide it. > - The spec also repeated service-start calls, page loads, and full-page screenshots. > - This pull request removes that repeated work and divides the scenario catalog into two independent specs. > - The benefit is a shorter Smoke Lab run and a balanced E2E lane without more AWS capacity. ## Linked Issues or Issue Description Refs: #10629 **What existing behavior does this improve?** This improves the trusted pull request E2E lane and its Smoke Lab Playwright coverage. **Current behavior** Smoke Lab is one indivisible 168-second CI spec. It starts services for every scenario, loads the same evidence page twice, and captures a full-page success screenshot for all 56 lifecycle steps. **Proposed behavior** Start Smoke Lab services once per spec. Keep the per-scenario fixture reset. Capture one representative success screenshot per scenario and keep every failure screenshot. Run P1–P4 and P5–P7 as separate specs so the existing duration-aware scheduler can put them on different runners. **Reason and benefit** The optimized lifecycle reduced local Smoke Lab wall time from 57.68 seconds to 37.04 seconds. This is a 35.8% reduction. The two halves also let the existing three runners target about 125, 124, and 124 seconds of recorded spec work instead of about 168, 125, and 124 seconds. **Breaking changes** None. The same seven scenarios and eight lifecycle steps still run. The result API still records every step. Successful non-connect steps no longer attach redundant screenshots. ## What Changed - Reused one Smoke Lab service start within each spec while retaining isolated fixture installation for every scenario. - Removed the duplicate catalog evidence navigation. - Reduced success screenshots from 56 to 7 while retaining screenshots for every failed step. - Split the shared lifecycle runner into P1–P4 and P5–P7 specs. - Mark each successful split result as partial and keep dashboard health amber until one run covers the full catalog. - Updated the duration manifest and contributor docs for the split. ## Verification - `pnpm -r typecheck` passed on Node.js 24.20.0. - `pnpm build` passed on Node.js 24.20.0. - `node --test scripts/__tests__/e2e-shard.test.mjs` passed 9 tests. - `pnpm exec vitest run ui/src/pages/tools/smoke-lab-matrix.test.ts` passed 8 tests. - Both split specs passed together on Node.js 24.20.0 after the review fixes: 2 passed in 35.7 seconds; shell wall time was 36.86 seconds. - The pre-change Smoke Lab baseline passed with a 57.68-second shell wall time. The optimized unsplit A/B run passed with a 37.04-second shell wall time. - The full local E2E catalog passed 44 tests and skipped 2 tests. One existing `pipelines-tutorial-flow.spec.ts` assertion failed again when run alone. - The broad local unit run reproduced failures in untouched workspace-runtime suites. Typecheck, build, shard tests, and all changed browser coverage pass. CI remains the authoritative full-suite result. ## Risks - The split duration weights use the measured local reduction and the previous 168-second CI weight. They should be refreshed after two real pull request runs. - Service state is shared within each half. Fixture installation still runs before every scenario to reset connection, policy, and catalog state. - Each half records passed execution with partial coverage. Dashboard health recognizes the partial flag and stays amber because no single runner covers the full catalog. A failed half still records failed/red. - Fewer success screenshots reduce redundant artifacts. Every scenario keeps its connect screenshot, and every failure still captures evidence. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The exact deployment ID and context window are not exposed in this session. The model used agentic reasoning, code editing, shell execution, browser testing, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
22 lines
1.4 KiB
JSON
22 lines
1.4 KiB
JSON
{
|
|
"$comment": "Per-spec wall-clock durations (ms) for the Playwright e2e lane, used by scripts/e2e-shard.mjs to balance specs across the PR shard matrix. Averaged from two real PR runs of .github/workflows/pr.yml (actions runs 29775184389 and 29769465682, 2026-07-20). The split Smoke Lab weights scale its prior 168s CI duration by the measured local before/after ratio and scenario count; refresh them from the next two real PR runs. Specs missing here get the median weight, so the manifest only needs occasional refreshes: pull the 'Run e2e tests' logs from a recent PR run and sum the Playwright list-reporter durations per spec file.",
|
|
"unit": "ms",
|
|
"durations": {
|
|
"tests/e2e/app-not-connected.spec.ts": 26750,
|
|
"tests/e2e/application-delete-screenshot.spec.ts": 6750,
|
|
"tests/e2e/applications-crud.spec.ts": 16100,
|
|
"tests/e2e/apps-dark-mode-shots.spec.ts": 18000,
|
|
"tests/e2e/apps-prosumer-mcp-flow.spec.ts": 16800,
|
|
"tests/e2e/conference-room-typing-intro.spec.ts": 4900,
|
|
"tests/e2e/mcp-user-stories.spec.ts": 39000,
|
|
"tests/e2e/nux-phase4-screenshots.spec.ts": 19650,
|
|
"tests/e2e/onboarding.spec.ts": 3200,
|
|
"tests/e2e/pipelines-tutorial-flow.spec.ts": 30198,
|
|
"tests/e2e/planning-mode-visual-verification.spec.ts": 23250,
|
|
"tests/e2e/sidebar-takeover.spec.ts": 16650,
|
|
"tests/e2e/signoff-policy.spec.ts": 9696,
|
|
"tests/e2e/smoke-lab-p1-p4.spec.ts": 71000,
|
|
"tests/e2e/smoke-lab-p5-p7.spec.ts": 53000
|
|
}
|
|
}
|