mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must continue tasks using user answers without losing earlier requirements or approval gates. > - The wake prompt mixed human decisions with prior tool evidence and repeated detailed question instructions. > - Those instructions belong with the question tool, with a short routing hint in the wake. > - The Runner evals need to prove that answers, approvals, and completed work survive later turns. > - This PR shortens the prompts, separates authenticated answers, and adds continuation tests with useful screenshots. ## Linked Issues or Issue Description Refs #13517. This is a follow-up to the merged onboarding skill and Runner E2E work. Related #13539 covers responses received while a run is active; this PR preserves its cases and adds continuation coverage. Existing continuation/recovery and question PRs were searched; none covers this prompt/documentation and eval change. **What existing behavior does this improve?** The instructions sent when an agent continues a task, the native human-input tool documentation, and the evidence captured by Runner full-stack E2E. **Current behavior** The wake repeats a long question-tool guide. Human answers appear alongside untrusted prior results. Screenshot capture can finish at DOM load while the task still shows a spinner, even when backend behavior checks pass. **Proposed behavior** Keep earlier requirements unless the user changes them. Treat clarification as distinct from approval. Give authenticated human responses a scoped field. Keep tool and agent results as evidence. Put detailed question behavior in the tool descriptor and retain one routing sentence in the native wake. Wait for the correct task and loaded conversation before taking screenshots. **Reason and benefit** Reduce repeated prompt text and make authority boundaries clear. Test that real question cards, later answers, approval gates, and completed child tasks still work. Make screenshots useful for human review. ## What Changed - Shorten shared continuation instructions for legacy and native runners. Separate authenticated user responses from tool results and agent summaries. - Remove the detailed question guide from native wake prompts. Keep its behavior in the canonical `request_human_input` descriptor and existing payload schema. Regenerate semantic contracts and fixture hashes. - Add five continuation cases across four local profiles. Add a dedicated choice-then-text case for native Codex and native Claude. All 22 cells join the shared full E2E campaign. - Cover revised scope, clarification without approval, hostile instructions in a handoff file, and reuse of a completed child after restart. Keep production instructions and fixed user facts. - Capture continuation screenshots only when the intended task and conversation have rendered. Add provider-free browser regressions for loaders and wrong-task capture. - Preserve current master’s extra tool and onboarding cases. The default campaign now contains 166 cells; 35 manual everyday cells remain separate. ## Verification - `pnpm -r typecheck`: passed after replay on current master. - `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed. - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm test:e2e:runner:browser-support`: 4 passed. These tests failed against immediate screenshot capture and passed after the fix. - Focused continuation and native-input tests: 36 passed locally. The tool-authority suite could not initialize embedded PostgreSQL locally, including one isolated retry; its 17 assertions did not run locally. The full remote server shards passed on this PR commit. - `pnpm build`: passed after replay on current master. `pnpm test:run` was attempted locally but hit the same embedded PostgreSQL initialization failure; the remaining local run was stopped after complete remote CI passed. This is not claimed as a full local test pass. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685): passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All server/chat/workspace/serialized shards, browser shards, Runner checks, typecheck, build, canary and policy checks passed. The isolated native Runner build and security checks also passed: 57 successful checks, with two expected Storybook skips. - Greptile reviewed the exact PR head at 5/5, with no findings or unresolved review threads. The PR has no merge conflicts. - [Live question-docs report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/): 3/3 passed at source `83dd132f2` before replay on master. Native Codex and Claude each asked a choice, waited, asked a text question, and saved both answers. Claude also passed a completed-child restart case. All three native turns are checked for absence of the old question block. - [Earlier continuation report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/): all five continuation cases passed on native Claude. The report retains campaign and revision provenance and separately shows two unresolved onboarding behavior failures. - [Before/after prompt report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/): full text, current recorded Claude inputs, and reproducible reference-token counts. The controlled wake comparison removes 401 reference tokens; the net counted input reduction is 339 after charging the larger tool description. These are text-size estimates, not measured billing savings. ## Risks - Prompt wording affects model behavior. Live results cover the stated cases, not every provider or conversation. Legacy profiles are registered but were not rerun for this change. - The optional continuation field changes prompt data only; there is no database migration or new production API. - Authenticated answer projection excludes generated summaries and agent-resolved interactions. It preserves the answer’s question or approval scope. - The screenshot guard can expose UI loading failures that earlier runs hid. Backend grading alone no longer makes those captures valid. - The two prior onboarding failures remain separate product issues: work before acceptance and a missing saved plan. This PR does not claim the entire onboarding suite passes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, code execution, and browser verification. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted tests above; the full local database-startup limit is documented - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>
70 lines
4.0 KiB
TypeScript
70 lines
4.0 KiB
TypeScript
import type { ContinuationCheckpoint } from "./continuation-scoring.js";
|
|
|
|
type Row = Record<string, any>;
|
|
const questions = (checkpoint?: ContinuationCheckpoint): Row[] =>
|
|
(checkpoint?.interactions ?? []).filter((value: any) => value.kind === "ask_user_questions") as Row[];
|
|
function form(row?: Row) {
|
|
const payload = row?.payload;
|
|
const canonical = payload?.questionSet?.questions;
|
|
if (canonical) return canonical;
|
|
return (payload?.questions ?? []).map((q: Row) => ({
|
|
...q, answerMode: q.selectionMode === "multi" ? "multi_select" : "single_select",
|
|
}));
|
|
}
|
|
|
|
/** Checks real recorded forms, responses, and native task inputs; no model judge. */
|
|
export function gradeQuestionDocumentation(checkpoints: ContinuationCheckpoint[], marker: string) {
|
|
const initial = checkpoints.find(c => c.phase === "initial");
|
|
const middle = checkpoints.find(c => c.phase === "answered");
|
|
const final = checkpoints.find(c => c.phase === "final");
|
|
const first = questions(initial)[0];
|
|
const next = questions(middle).filter(q => q.status === "pending");
|
|
const choice = form(first)[0];
|
|
const text = form(next[0])[0];
|
|
const options: Row[] = choice?.options ?? [];
|
|
const afternoon = options.find(o => /^afternoon\b/i.test(String(o.label).trim()));
|
|
const firstAnswer = questions(final).find(q => q.id === first?.id);
|
|
const secondAnswer = questions(final).find(q => q.id === next[0]?.id);
|
|
const inputs = (final?.runs ?? []).map((run: any) => run.runnerProfileJson?.nativeExecutionInput?.task?.prompt);
|
|
const checks = [
|
|
{
|
|
id: "meaningful-choice-form",
|
|
passed: questions(initial).length === 1 && first?.status === "pending" && form(first).length === 1
|
|
&& choice?.answerMode === "single_select" && options.length >= 2
|
|
&& new Set(options.map(o => o.id)).size === options.length
|
|
&& options.some(o => /^morning\b/i.test(String(o.label).trim())) && Boolean(afternoon),
|
|
detail: "The first card must show one single-choice question with distinct Morning and Afternoon options.",
|
|
},
|
|
{
|
|
id: "text-form-after-choice",
|
|
passed: questions(middle).length === 2 && next.length === 1 && form(next[0]).length === 1
|
|
&& next[0]?.id !== first?.id && text?.answerMode === "text"
|
|
&& !text?.options?.length && !text?.customAnswer
|
|
&& questions(middle).some(q => q.id === first?.id && q.status === "answered"),
|
|
detail: "Only after the choice is answered may the second, text-only question appear.",
|
|
},
|
|
{
|
|
id: "actual-question-answers",
|
|
passed: questions(final).length === 2 && firstAnswer?.status === "answered" && secondAnswer?.status === "answered"
|
|
&& firstAnswer?.result?.answers?.some((a: Row) => a.questionId === choice?.id && a.optionIds?.includes(afternoon?.id))
|
|
&& secondAnswer?.result?.answers?.some((a: Row) => a.questionId === text?.id && a.otherText?.includes(marker)),
|
|
detail: "Both real UI submissions must be persisted against their original question IDs, with no duplicate question cards.",
|
|
},
|
|
{
|
|
id: "question-guidance-not-in-wake",
|
|
passed: inputs.length >= 3 && inputs.every(prompt => typeof prompt === "string"
|
|
&& prompt.includes("Use Paperclip's request_human_input for durable task questions.")
|
|
&& !prompt.includes("## Questions that need a user response")
|
|
&& !prompt.includes("payload.questionSet")
|
|
&& !prompt.includes("Create a durable human question or approval card on the current Paperclip task bound to this run")),
|
|
detail: "Every recorded native turn must retain the short routing hint without the old question block or copied tool-format instructions.",
|
|
},
|
|
{
|
|
id: "both-answers-in-output",
|
|
passed: Boolean(final?.documents.some(d => d.key !== "plan" && d.body.includes(marker) && /\bafternoon\b/i.test(d.body))),
|
|
detail: "The saved note must use both the selected time and the literal reference supplied in the text answer.",
|
|
},
|
|
];
|
|
return checks.map(check => ({ ...check, passed: Boolean(check.passed) }));
|
|
}
|