Files
DottaandPaperclip 0f5fafe16b fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - The Runner connects provider sessions to task state, replies, and
delegated work.
> - Full-stack tests found lost final replies, rejected helper calls
that stopped the parent, and unnecessary process restarts.
> - A completed child could also receive a new assignment wake that the
scheduler then cancelled.
> - This change fixes those boundaries and gives agents clearer teammate
instructions.
> - The tests retain strict completion and process-continuity
requirements.

## Linked Issues or Issue Description

**What happened?**

A generated attachment comment could suppress an agent's final reply. A
known Codex helper could stop its parent when it requested a Paperclip
tool. Native Daytona processes restarted between turns because Paperclip
minted an unused GitHub broker token. Reassigning a completed child
queued a run that immediately cancelled. Revision instructions also
allowed agents to do work assigned to a named teammate themselves.

**Expected behavior**

Keep the final reply. Reject helper tool requests without borrowing
parent authority or stopping the parent. Keep an unconfigured sandbox
process alive between turns. Treat assignment-only changes to completed
tasks as metadata changes. Preserve explicit teammate assignments during
revisions.

**Steps to reproduce**

Run the retained Runner E2E cases for file handoff, teammate reuse,
Daytona warm continuity, and Legacy Claude interview/plan acceptance.
The focused regression tests reproduce the reply, helper,
process-lifetime, and assignment-wake defects without provider calls.

**Paperclip version or commit**

The live lifetime and completion campaign used `db3857807`. This PR
replays the changes on master `c65fc9e3c`. See Verification for the
limits of that evidence.

**Deployment mode**

Isolated local instances and native Runner sessions in Daytona
sandboxes.

Related work: #13546 handles a different queued-run issue after an
issue-lock compare-and-set failure. This PR prevents the unnecessary
assignment wake earlier. #13410 covers retained user services; this PR
covers the provider process. No duplicate fix was found.

## What Changed

- Exclude generated deliverable-binding comments from final-reply
deduplication. Preserve the attachment and explicit user-facing replies.
- Reject Paperclip tool and input requests from known Codex helper
threads without terminating the parent. Keep unknown-thread rejection
intact.
- Explain how to hire or reuse a persistent teammate and preserve named
delegation on revisions. Update generated protocol fixtures.
- Use stable, token-free GitHub wrappers for unconfigured native
sandboxes. Preserve credential isolation, configured-account rotation,
and cleanup after partial staging failures.
- Do not queue assignment-only wakes for done or cancelled tasks. Keep
explicit reopening behavior.
- Make warm-continuity fixtures create real review cards. Read the
persisted final response selected by production presentation logic.
Missing selected evidence still fails.

## Verification

- Before rebase: 560 focused route, native-executor, and launcher tests
passed. The new regressions were reproduced before their fixes.
- Live E2E: Legacy Claude interview/plan acceptance passed 3/3
repetitions. Daytona warm continuity passed 2/3 full repetitions. Each
successful run retained one process and provider session for all three
turns.
- The remaining Daytona repetition stopped after a same-URL browser
reload left the page blank. Both completed turns retained the same
process. Its failed verdict remains unchanged; this PR does not claim
the blank-page cause is fixed.
- Reports:
https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html
and
https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html
- Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414
Runner E2E harness unit tests and its typecheck passed. Codex protocol
tests: 88 passed, 2 ignored.
- Latest-head CI: 55 successful checks and 2 intentional skips.
Greptile: 5/5 with no review threads. The unchanged workspace exposure
tests hit a fixed-port collision on the first CI attempt; their local
suite passed (25 tests, 3 platform skips), and the CI shard passed on
one retry.
- The duplicate local `pnpm test:run` was stopped after the full hosted
general and serialized test shards passed. It did not finish locally and
is not counted as a local full-suite pass.

## Risks

Configured GitHub accounts retain run-scoped credential rotation and can
still restart warm processes. That limitation requires a separate
design. Known provider helpers cannot use Paperclip coordination tools
directly; they must return findings to the parent. The delegation prompt
is an instruction, not an enforced guarantee; Codex Mini hiring/reuse
failures remain open. No schema or workflow changes are included.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
parallel coding agents. The exact deployment suffix and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub issues and shared reports)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full hosted test
shards also pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 08:56:35 -05:00

186 lines
6.1 KiB
TypeScript

import { createRequire } from "node:module";
import { readFile } from "node:fs/promises";
import type { Matcher } from "./types.js";
interface JsonSchemaValidator {
(value: unknown): boolean;
errors?: unknown;
}
interface Ajv2020Instance {
compile(schema: Record<string, unknown>): JsonSchemaValidator;
}
const runnerRequire = createRequire(
new URL("../../packages/paperclip-runner/package.json", import.meta.url),
);
const Ajv2020 = runnerRequire("ajv/dist/2020.js").default as new (options: {
allErrors: boolean;
strict: boolean;
}) => Ajv2020Instance;
const jsonSchemaCompiler = new Ajv2020({ allErrors: true, strict: false });
export interface MatcherObservation {
message?: string;
issueStatus?: string;
runStatus?: string;
runtimeMode?: string;
environment?: string;
files?: Record<string, string>;
artifacts?: Array<{ name: string; mimeType?: string }>;
json?: unknown;
}
export interface MatcherResult {
matcher: Matcher;
passed: boolean;
detail: string;
}
function normalizeMessage(value: string | undefined) {
return (
(value ?? "")
.replace(/\r\n/g, "\n")
// Some provider renderers escape underscores in plain-text identifiers
// before persisting Markdown. Treat that presentation-only escape as the
// same visible marker for exact/contains/ordered message assertions.
.replace(/\\_/g, "_")
.replace(/[ \t]+/g, " ")
.trim()
);
}
function readJsonPath(value: unknown, path: string): unknown {
return path
.split(".")
.filter(Boolean)
.reduce<unknown>((current, segment) => {
if (!current || typeof current !== "object") return undefined;
return (current as Record<string, unknown>)[segment];
}, value);
}
function countOccurrences(value: string, expected: string): number {
if (!expected) return 0;
let count = 0;
let cursor = 0;
while (cursor <= value.length - expected.length) {
const index = value.indexOf(expected, cursor);
if (index < 0) break;
count += 1;
cursor = index + expected.length;
}
return count;
}
export async function evaluateMatcher(
matcher: Matcher,
observation: MatcherObservation,
): Promise<MatcherResult> {
const message = normalizeMessage(observation.message);
let passed = false;
let actual: unknown;
if (matcher.kind === "message_exact") {
actual = message;
passed = message === normalizeMessage(matcher.expected);
} else if (matcher.kind === "message_contains") {
actual = message;
passed = message.includes(normalizeMessage(matcher.expected));
} else if (matcher.kind === "message_occurrences") {
actual = countOccurrences(message, normalizeMessage(matcher.expected));
passed = actual === matcher.count;
} else if (matcher.kind === "message_regex") {
actual = message;
passed = new RegExp(matcher.pattern, matcher.flags).test(message);
} else if (matcher.kind === "message_ordered") {
actual = message;
let cursor = 0;
passed = matcher.expected.every((expected) => {
const normalizedExpected = normalizeMessage(expected);
const index = message.indexOf(normalizedExpected, cursor);
if (index < 0) return false;
cursor = index + normalizedExpected.length;
return true;
});
} else if (matcher.kind === "issue_status") {
actual = observation.issueStatus;
passed = actual === matcher.expected;
} else if (matcher.kind === "run_status") {
actual = observation.runStatus;
passed = actual === matcher.expected;
} else if (matcher.kind === "runtime_mode") {
actual = observation.runtimeMode;
passed = actual === matcher.expected;
} else if (matcher.kind === "environment") {
actual = observation.environment;
passed = actual === matcher.expected;
} else if (
matcher.kind === "file_exists" ||
matcher.kind === "file_exact" ||
matcher.kind === "file_contains"
) {
try {
actual =
observation.files?.[matcher.path] ??
(await readFile(matcher.path, "utf8"));
passed =
matcher.kind === "file_exists" ||
(matcher.kind === "file_exact"
? String(actual) === matcher.expected
: String(actual).includes(matcher.expected));
} catch {
actual = undefined;
passed = false;
}
} else if (matcher.kind === "artifact_exists") {
actual = observation.artifacts ?? [];
passed = (observation.artifacts ?? []).some(
(artifact) =>
artifact.name === matcher.name &&
(!matcher.mimeType || artifact.mimeType === matcher.mimeType),
);
} else if (matcher.kind === "json_path") {
actual = readJsonPath(observation.json, matcher.path);
passed = JSON.stringify(actual) === JSON.stringify(matcher.expected);
} else {
const validate = jsonSchemaCompiler.compile(matcher.schema);
passed = validate(observation.json);
actual = passed
? observation.json
: { value: observation.json, errors: validate.errors ?? [] };
}
return {
matcher,
passed,
detail: passed
? "matched"
: `expected ${JSON.stringify(matcher)}; observed ${JSON.stringify(actual)}`,
};
}
export async function evaluateMatchers(
matchers: readonly Matcher[],
observation: MatcherObservation,
): Promise<MatcherResult[]> {
return Promise.all(
matchers.map((matcher) => evaluateMatcher(matcher, observation)),
);
}
export function persistedFinalRunMessage(
comments: Array<{ id: string; createdByRunId?: string | null; body?: string | null }>,
run: { id: string; resultJson?: Record<string, unknown> | null },
): string {
const runComments = comments.filter(comment => comment.createdByRunId === run.id);
const decision = run.resultJson?.presentationDecision;
const selectedId = decision && typeof decision === "object" && !Array.isArray(decision)
? (decision as Record<string, unknown>).commentId : null;
// Read the persisted comment selected for the user, not a finish summary or
// attachment preparation message. Missing selected evidence must still fail.
if (typeof selectedId === "string") {
return runComments.find(comment => comment.id === selectedId)?.body ?? "";
}
return runComments.map(comment => comment.body ?? "").join("\n");
}