Files
PaperClipAI/tests/runner-e2e/browser-bootstrap-diagnostics.spec.ts
DottaandPaperclip 9a76390cfa test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.

## Linked Issues or Issue Description

Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.

**What happened?**

Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.

**Expected behavior**

Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.

**Steps to reproduce**

Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.

**Paperclip version or commit**

Based on master at a68f3d8e3.

## What Changed

- Make synthetic exposure failures use the assigned port. Test both free
and occupied base pairs.
- Record document readiness, root children, service-worker control,
pending module paths, and recent module error or 304 statuses in private
runner evidence.
- Add browser tests for pending startup, failed module responses, and
diagnostic reset after navigation.
- Add a provider-free browser test for persisted task content and the
composer across three navigation/reload cycles.
- Guard trusted report-job lockfile resolution before frozen install and
AWS credential setup.
- Document the new evidence and test commands.

## Verification

- `pnpm test:e2e:runner:unit`: 443 tests passed.
- `pnpm test:e2e:runner:browser-support`: 10 tests passed.
- Exposure tests: 28 passed; 3 platform-specific skips.
- Harness typecheck, workspace typecheck, and full build passed.
- Task-reload browser test passed against a throwaway local instance. It
checks three navigation/reload cycles with a controlling service worker.
- All PR verification suites passed, including server, chat, workspace,
serialized server, Rust, runner, build, typecheck, and browser shards.
The existing sidebar navigation case showed an empty page on the first
CI attempt and passed on one retry.
- The monolithic local `pnpm test:run` was started, then stopped after
CI completed the equivalent suites. Additional local browser
reproduction attempts also hit embedded-Postgres startup failures; the
focused checks listed above completed successfully.
- Greptile: 5/5 on the current head, with no findings.

## Risks

This change affects tests and private test evidence. It does not change
production prompts or runtime behavior. Module paths omit queries; the
collector reads no response bodies or headers. The existing sanitizer
and publication allowlist still apply. A 304 response does not cause a
test failure.

The blank-page cause remains unconfirmed. The first CI attempt
reproduced it in an existing sidebar navigation case: the trace shows an
empty page after a service-worker-mediated 304 response for the large
editor module. The single retry passed. This PR adds coverage and
diagnostics; it does not claim to fix that intermittent symptom.

## Model Used

OpenAI Codex, GPT-6, with repository tools, code execution, and browser
tests. The exact deployed model variant and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:04:13 -05:00

45 lines
2.4 KiB
TypeScript

import { expect, test } from "@playwright/test";
import { observeBrowserBootstrap } from "./browser-bootstrap-diagnostics.js";
const html = '<div id="root"></div><script type="module" src="/entry.js?private=do-not-record"></script>';
test("records an empty root and pending module, then recognizes completed startup", async ({ page }) => {
let release!: () => void;
const gate = new Promise<void>(resolve => { release = resolve; });
await page.route("http://bootstrap.test/**", async route => {
if (new URL(route.request().url()).pathname === "/entry.js") {
await gate;
await route.fulfill({ contentType: "text/javascript", body: 'document.getElementById("root").innerHTML = "<main>Ready</main>";' });
} else await route.fulfill({ contentType: "text/html", body: html });
});
const diagnostics = observeBrowserBootstrap(page);
try {
await page.goto("http://bootstrap.test/", { waitUntil: "commit" });
await expect.poll(async () => (await diagnostics.snapshot()).pendingModules).toEqual(["/entry.js"]);
const before = await diagnostics.snapshot();
expect(before.state).toMatchObject({ rootPresent: true, rootChildCount: 0, serviceWorkerControlled: false });
expect(JSON.stringify(before)).not.toContain("do-not-record");
release();
await expect(page.getByRole("main")).toHaveText("Ready");
await expect.poll(async () => (await diagnostics.snapshot()).pendingModuleCount).toBe(0);
expect((await diagnostics.snapshot()).state?.rootChildCount).toBe(1);
} finally { release(); diagnostics.dispose(); }
});
for (const status of [304, 503]) {
test(`retains module status ${status} without credentials and clears it on navigation`, async ({ page }) => {
await page.route("http://bootstrap.test/**", async route => {
const pathname = new URL(route.request().url()).pathname;
await route.fulfill(pathname === "/entry.js" ? { status, body: "" }
: { contentType: "text/html", body: pathname === "/ready" ? '<div id="root">Ready</div>' : html });
});
const diagnostics = observeBrowserBootstrap(page);
try {
await page.goto("http://bootstrap.test/", { waitUntil: "load" });
expect((await diagnostics.snapshot()).moduleResponses).toEqual([{ path: "/entry.js", status }]);
await page.goto("http://bootstrap.test/ready");
expect((await diagnostics.snapshot()).moduleResponses).toEqual([]);
} finally { diagnostics.dispose(); }
});
}