mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
Verify document navigation in both task interfaces
Co-Authored-By: Paperclip <noreply@paperclip.ing>
This commit is contained in:
1 parent
103063f81d
commit
3a7349ddc6
2 files changed
+6
-2
No files matched your search
@@ -6,7 +6,7 @@ The correction keeps the shortened completion procedure while explicitly requiri
|
||||
|
||||
Only the manual instruction comparison uses the v3 final-answer observation. The existing `native-completion` suite retains its v2 verdict checks and does not acquire the browser-navigation requirement. A calibration demonstrates the same retained input can pass v2 and fail v3; original report files are never rewritten. The harness definition digest changes truthfully for future invocations. Its additional checks read only the actual run-attributed persisted provider final. A blocker label and action do not substitute for an explicit missing-access explanation. Document links must identify the exact task and one revisioned document at the instance origin; wrong origin, task, document, query, credentials or absent evidence fail. Exact call-ID joins, acceptance/termination, ordering, side-effect and budget checks remain enforced.
|
||||
|
||||
After saving the strict native snapshot, each completion cell clicks the rendered final-reply link in the browser. It must open the canonical saved document and show the original content marker. Retain the navigation receipt and screenshot independently; a valid-looking Markdown URL alone does not qualify navigation.
|
||||
After saving the strict native snapshot, each completion cell clicks the rendered final-reply link in the browser. It must open the canonical saved document and show the original content marker in exactly one visible classic document card or document-specific side-panel tab. The navigation check follows the actual configured UI instead of requiring a classic-only element. Retain the navigation receipt and screenshot independently; a valid-looking Markdown URL alone does not qualify navigation.
|
||||
|
||||
Calibrate with correct paraphrases and plausible wrong answers, including the action-only omission. Replay the retained observations only as labeled additional diagnostics; never overwrite or regrade the original results. Provider-free measurement v2 captures eighteen shared runnerd RPC projections plus six direct OpenCode HTTP projections across v4/v5 start, resume and continuation. The latter exercises the concrete OpenCode backend using a local fake server. Both boundaries measure Paperclip-supplied payloads, not provider stock prompts or model cognition. No synthetic receipt qualifies live model behavior.
|
||||
|
||||
|
||||
@@ -2717,7 +2717,11 @@ for (const execution of executions) {
|
||||
await expect(link).toBeVisible({ timeout: 30_000 });
|
||||
await link.click();
|
||||
await expect(page).toHaveURL(new URL(href, observation.documentLinkContext.appOrigin).href);
|
||||
const target = page.locator(`[id=${JSON.stringify(`document-${document.key}`)}]`);
|
||||
const target = page.locator([
|
||||
`[id=${JSON.stringify(`document-${document.key}`)}]:visible`,
|
||||
`[id=${JSON.stringify(`side-panel-content-document:${document.key}`)}]:visible`,
|
||||
].join(", "));
|
||||
await expect(target).toHaveCount(1);
|
||||
await expect(target).toBeVisible();
|
||||
await expect(target).toContainText(marker);
|
||||
opened = true;
|
||||
|
||||
Reference in new issue
Block a user