Files
PaperClipAI/tests/runner-e2e/question-documentation-scoring.ts
DottaandPaperclip e26d787928 Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must continue tasks using user answers without losing earlier
requirements or approval gates.
> - The wake prompt mixed human decisions with prior tool evidence and
repeated detailed question instructions.
> - Those instructions belong with the question tool, with a short
routing hint in the wake.
> - The Runner evals need to prove that answers, approvals, and
completed work survive later turns.
> - This PR shortens the prompts, separates authenticated answers, and
adds continuation tests with useful screenshots.

## Linked Issues or Issue Description

Refs #13517. This is a follow-up to the merged onboarding skill and
Runner E2E work. Related #13539 covers responses received while a run is
active; this PR preserves its cases and adds continuation coverage.
Existing continuation/recovery and question PRs were searched; none
covers this prompt/documentation and eval change.

**What existing behavior does this improve?**

The instructions sent when an agent continues a task, the native
human-input tool documentation, and the evidence captured by Runner
full-stack E2E.

**Current behavior**

The wake repeats a long question-tool guide. Human answers appear
alongside untrusted prior results. Screenshot capture can finish at DOM
load while the task still shows a spinner, even when backend behavior
checks pass.

**Proposed behavior**

Keep earlier requirements unless the user changes them. Treat
clarification as distinct from approval. Give authenticated human
responses a scoped field. Keep tool and agent results as evidence. Put
detailed question behavior in the tool descriptor and retain one routing
sentence in the native wake. Wait for the correct task and loaded
conversation before taking screenshots.

**Reason and benefit**

Reduce repeated prompt text and make authority boundaries clear. Test
that real question cards, later answers, approval gates, and completed
child tasks still work. Make screenshots useful for human review.

## What Changed

- Shorten shared continuation instructions for legacy and native
runners. Separate authenticated user responses from tool results and
agent summaries.
- Remove the detailed question guide from native wake prompts. Keep its
behavior in the canonical `request_human_input` descriptor and existing
payload schema. Regenerate semantic contracts and fixture hashes.
- Add five continuation cases across four local profiles. Add a
dedicated choice-then-text case for native Codex and native Claude. All
22 cells join the shared full E2E campaign.
- Cover revised scope, clarification without approval, hostile
instructions in a handoff file, and reuse of a completed child after
restart. Keep production instructions and fixed user facts.
- Capture continuation screenshots only when the intended task and
conversation have rendered. Add provider-free browser regressions for
loaders and wrong-task capture.
- Preserve current master’s extra tool and onboarding cases. The default
campaign now contains 166 cells; 35 manual everyday cells remain
separate.

## Verification

- `pnpm -r typecheck`: passed after replay on current master.
- `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed.
- `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
test:e2e:runner:browser-support`: 4 passed. These tests failed against
immediate screenshot capture and passed after the fix.
- Focused continuation and native-input tests: 36 passed locally. The
tool-authority suite could not initialize embedded PostgreSQL locally,
including one isolated retry; its 17 assertions did not run locally. The
full remote server shards passed on this PR commit.
- `pnpm build`: passed after replay on current master. `pnpm test:run`
was attempted locally but hit the same embedded PostgreSQL
initialization failure; the remaining local run was stopped after
complete remote CI passed. This is not claimed as a full local test
pass.
- [Full PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685):
passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All
server/chat/workspace/serialized shards, browser shards, Runner checks,
typecheck, build, canary and policy checks passed. The isolated native
Runner build and security checks also passed: 57 successful checks, with
two expected Storybook skips.
- Greptile reviewed the exact PR head at 5/5, with no findings or
unresolved review threads. The PR has no merge conflicts.
- [Live question-docs
report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/):
3/3 passed at source `83dd132f2` before replay on master. Native Codex
and Claude each asked a choice, waited, asked a text question, and saved
both answers. Claude also passed a completed-child restart case. All
three native turns are checked for absence of the old question block.
- [Earlier continuation
report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/):
all five continuation cases passed on native Claude. The report retains
campaign and revision provenance and separately shows two unresolved
onboarding behavior failures.
- [Before/after prompt
report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/):
full text, current recorded Claude inputs, and reproducible
reference-token counts. The controlled wake comparison removes 401
reference tokens; the net counted input reduction is 339 after charging
the larger tool description. These are text-size estimates, not measured
billing savings.

## Risks

- Prompt wording affects model behavior. Live results cover the stated
cases, not every provider or conversation. Legacy profiles are
registered but were not rerun for this change.
- The optional continuation field changes prompt data only; there is no
database migration or new production API.
- Authenticated answer projection excludes generated summaries and
agent-resolved interactions. It preserves the answer’s question or
approval scope.
- The screenshot guard can expose UI loading failures that earlier runs
hid. Backend grading alone no longer makes those captures valid.
- The two prior onboarding failures remain separate product issues: work
before acceptance and a missing saved plan. This PR does not claim the
entire onboarding suite passes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, code execution,
and browser verification. The exact deployed model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted tests above; the
full local database-startup limit is documented
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 09:32:54 -05:00

70 lines
4.0 KiB
TypeScript

import type { ContinuationCheckpoint } from "./continuation-scoring.js";
type Row = Record<string, any>;
const questions = (checkpoint?: ContinuationCheckpoint): Row[] =>
(checkpoint?.interactions ?? []).filter((value: any) => value.kind === "ask_user_questions") as Row[];
function form(row?: Row) {
const payload = row?.payload;
const canonical = payload?.questionSet?.questions;
if (canonical) return canonical;
return (payload?.questions ?? []).map((q: Row) => ({
...q, answerMode: q.selectionMode === "multi" ? "multi_select" : "single_select",
}));
}
/** Checks real recorded forms, responses, and native task inputs; no model judge. */
export function gradeQuestionDocumentation(checkpoints: ContinuationCheckpoint[], marker: string) {
const initial = checkpoints.find(c => c.phase === "initial");
const middle = checkpoints.find(c => c.phase === "answered");
const final = checkpoints.find(c => c.phase === "final");
const first = questions(initial)[0];
const next = questions(middle).filter(q => q.status === "pending");
const choice = form(first)[0];
const text = form(next[0])[0];
const options: Row[] = choice?.options ?? [];
const afternoon = options.find(o => /^afternoon\b/i.test(String(o.label).trim()));
const firstAnswer = questions(final).find(q => q.id === first?.id);
const secondAnswer = questions(final).find(q => q.id === next[0]?.id);
const inputs = (final?.runs ?? []).map((run: any) => run.runnerProfileJson?.nativeExecutionInput?.task?.prompt);
const checks = [
{
id: "meaningful-choice-form",
passed: questions(initial).length === 1 && first?.status === "pending" && form(first).length === 1
&& choice?.answerMode === "single_select" && options.length >= 2
&& new Set(options.map(o => o.id)).size === options.length
&& options.some(o => /^morning\b/i.test(String(o.label).trim())) && Boolean(afternoon),
detail: "The first card must show one single-choice question with distinct Morning and Afternoon options.",
},
{
id: "text-form-after-choice",
passed: questions(middle).length === 2 && next.length === 1 && form(next[0]).length === 1
&& next[0]?.id !== first?.id && text?.answerMode === "text"
&& !text?.options?.length && !text?.customAnswer
&& questions(middle).some(q => q.id === first?.id && q.status === "answered"),
detail: "Only after the choice is answered may the second, text-only question appear.",
},
{
id: "actual-question-answers",
passed: questions(final).length === 2 && firstAnswer?.status === "answered" && secondAnswer?.status === "answered"
&& firstAnswer?.result?.answers?.some((a: Row) => a.questionId === choice?.id && a.optionIds?.includes(afternoon?.id))
&& secondAnswer?.result?.answers?.some((a: Row) => a.questionId === text?.id && a.otherText?.includes(marker)),
detail: "Both real UI submissions must be persisted against their original question IDs, with no duplicate question cards.",
},
{
id: "question-guidance-not-in-wake",
passed: inputs.length >= 3 && inputs.every(prompt => typeof prompt === "string"
&& prompt.includes("Use Paperclip's request_human_input for durable task questions.")
&& !prompt.includes("## Questions that need a user response")
&& !prompt.includes("payload.questionSet")
&& !prompt.includes("Create a durable human question or approval card on the current Paperclip task bound to this run")),
detail: "Every recorded native turn must retain the short routing hint without the old question block or copied tool-format instructions.",
},
{
id: "both-answers-in-output",
passed: Boolean(final?.documents.some(d => d.key !== "plan" && d.body.includes(marker) && /\bafternoon\b/i.test(d.body))),
detail: "The saved note must use both the selected time and the literal reference supplied in the text answer.",
},
];
return checks.map(check => ({ ...check, passed: Boolean(check.passed) }));
}