mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
186 lines
6.1 KiB
TypeScript
186 lines
6.1 KiB
TypeScript
import { createRequire } from "node:module";
|
|
import { readFile } from "node:fs/promises";
|
|
import type { Matcher } from "./types.js";
|
|
|
|
interface JsonSchemaValidator {
|
|
(value: unknown): boolean;
|
|
errors?: unknown;
|
|
}
|
|
|
|
interface Ajv2020Instance {
|
|
compile(schema: Record<string, unknown>): JsonSchemaValidator;
|
|
}
|
|
|
|
const runnerRequire = createRequire(
|
|
new URL("../../packages/paperclip-runner/package.json", import.meta.url),
|
|
);
|
|
const Ajv2020 = runnerRequire("ajv/dist/2020.js").default as new (options: {
|
|
allErrors: boolean;
|
|
strict: boolean;
|
|
}) => Ajv2020Instance;
|
|
const jsonSchemaCompiler = new Ajv2020({ allErrors: true, strict: false });
|
|
|
|
export interface MatcherObservation {
|
|
message?: string;
|
|
issueStatus?: string;
|
|
runStatus?: string;
|
|
runtimeMode?: string;
|
|
environment?: string;
|
|
files?: Record<string, string>;
|
|
artifacts?: Array<{ name: string; mimeType?: string }>;
|
|
json?: unknown;
|
|
}
|
|
|
|
export interface MatcherResult {
|
|
matcher: Matcher;
|
|
passed: boolean;
|
|
detail: string;
|
|
}
|
|
|
|
function normalizeMessage(value: string | undefined) {
|
|
return (
|
|
(value ?? "")
|
|
.replace(/\r\n/g, "\n")
|
|
// Some provider renderers escape underscores in plain-text identifiers
|
|
// before persisting Markdown. Treat that presentation-only escape as the
|
|
// same visible marker for exact/contains/ordered message assertions.
|
|
.replace(/\\_/g, "_")
|
|
.replace(/[ \t]+/g, " ")
|
|
.trim()
|
|
);
|
|
}
|
|
|
|
function readJsonPath(value: unknown, path: string): unknown {
|
|
return path
|
|
.split(".")
|
|
.filter(Boolean)
|
|
.reduce<unknown>((current, segment) => {
|
|
if (!current || typeof current !== "object") return undefined;
|
|
return (current as Record<string, unknown>)[segment];
|
|
}, value);
|
|
}
|
|
|
|
function countOccurrences(value: string, expected: string): number {
|
|
if (!expected) return 0;
|
|
let count = 0;
|
|
let cursor = 0;
|
|
while (cursor <= value.length - expected.length) {
|
|
const index = value.indexOf(expected, cursor);
|
|
if (index < 0) break;
|
|
count += 1;
|
|
cursor = index + expected.length;
|
|
}
|
|
return count;
|
|
}
|
|
|
|
export async function evaluateMatcher(
|
|
matcher: Matcher,
|
|
observation: MatcherObservation,
|
|
): Promise<MatcherResult> {
|
|
const message = normalizeMessage(observation.message);
|
|
let passed = false;
|
|
let actual: unknown;
|
|
if (matcher.kind === "message_exact") {
|
|
actual = message;
|
|
passed = message === normalizeMessage(matcher.expected);
|
|
} else if (matcher.kind === "message_contains") {
|
|
actual = message;
|
|
passed = message.includes(normalizeMessage(matcher.expected));
|
|
} else if (matcher.kind === "message_occurrences") {
|
|
actual = countOccurrences(message, normalizeMessage(matcher.expected));
|
|
passed = actual === matcher.count;
|
|
} else if (matcher.kind === "message_regex") {
|
|
actual = message;
|
|
passed = new RegExp(matcher.pattern, matcher.flags).test(message);
|
|
} else if (matcher.kind === "message_ordered") {
|
|
actual = message;
|
|
let cursor = 0;
|
|
passed = matcher.expected.every((expected) => {
|
|
const normalizedExpected = normalizeMessage(expected);
|
|
const index = message.indexOf(normalizedExpected, cursor);
|
|
if (index < 0) return false;
|
|
cursor = index + normalizedExpected.length;
|
|
return true;
|
|
});
|
|
} else if (matcher.kind === "issue_status") {
|
|
actual = observation.issueStatus;
|
|
passed = actual === matcher.expected;
|
|
} else if (matcher.kind === "run_status") {
|
|
actual = observation.runStatus;
|
|
passed = actual === matcher.expected;
|
|
} else if (matcher.kind === "runtime_mode") {
|
|
actual = observation.runtimeMode;
|
|
passed = actual === matcher.expected;
|
|
} else if (matcher.kind === "environment") {
|
|
actual = observation.environment;
|
|
passed = actual === matcher.expected;
|
|
} else if (
|
|
matcher.kind === "file_exists" ||
|
|
matcher.kind === "file_exact" ||
|
|
matcher.kind === "file_contains"
|
|
) {
|
|
try {
|
|
actual =
|
|
observation.files?.[matcher.path] ??
|
|
(await readFile(matcher.path, "utf8"));
|
|
passed =
|
|
matcher.kind === "file_exists" ||
|
|
(matcher.kind === "file_exact"
|
|
? String(actual) === matcher.expected
|
|
: String(actual).includes(matcher.expected));
|
|
} catch {
|
|
actual = undefined;
|
|
passed = false;
|
|
}
|
|
} else if (matcher.kind === "artifact_exists") {
|
|
actual = observation.artifacts ?? [];
|
|
passed = (observation.artifacts ?? []).some(
|
|
(artifact) =>
|
|
artifact.name === matcher.name &&
|
|
(!matcher.mimeType || artifact.mimeType === matcher.mimeType),
|
|
);
|
|
} else if (matcher.kind === "json_path") {
|
|
actual = readJsonPath(observation.json, matcher.path);
|
|
passed = JSON.stringify(actual) === JSON.stringify(matcher.expected);
|
|
} else {
|
|
const validate = jsonSchemaCompiler.compile(matcher.schema);
|
|
passed = validate(observation.json);
|
|
actual = passed
|
|
? observation.json
|
|
: { value: observation.json, errors: validate.errors ?? [] };
|
|
}
|
|
return {
|
|
matcher,
|
|
passed,
|
|
detail: passed
|
|
? "matched"
|
|
: `expected ${JSON.stringify(matcher)}; observed ${JSON.stringify(actual)}`,
|
|
};
|
|
}
|
|
|
|
export async function evaluateMatchers(
|
|
matchers: readonly Matcher[],
|
|
observation: MatcherObservation,
|
|
): Promise<MatcherResult[]> {
|
|
return Promise.all(
|
|
matchers.map((matcher) => evaluateMatcher(matcher, observation)),
|
|
);
|
|
}
|
|
|
|
|
|
export function persistedFinalRunMessage(
|
|
comments: Array<{ id: string; createdByRunId?: string | null; body?: string | null }>,
|
|
run: { id: string; resultJson?: Record<string, unknown> | null },
|
|
): string {
|
|
const runComments = comments.filter(comment => comment.createdByRunId === run.id);
|
|
const decision = run.resultJson?.presentationDecision;
|
|
const selectedId = decision && typeof decision === "object" && !Array.isArray(decision)
|
|
? (decision as Record<string, unknown>).commentId : null;
|
|
// Read the persisted comment selected for the user, not a finish summary or
|
|
// attachment preparation message. Missing selected evidence must still fail.
|
|
if (typeof selectedId === "string") {
|
|
return runComments.find(comment => comment.id === selectedId)?.body ?? "";
|
|
}
|
|
return runComments.map(comment => comment.body ?? "").join("\n");
|
|
}
|