fix(hermes): settle native usage before governed confirmations

Extend the managed human-input boundary to confirmations and checkbox
confirmations. Preserve native receipts before controller parking and
cover immediate shutdown for all three canonical input kinds.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip committed 2026-10-08 12:28:13 -05:00
1 parent f29f7d89b9
commit efd5f5d74b
17 files changed
+106 -56

No files matched your search

@@ -1068,3 +1068,44 @@ Codex regressions pass 318 checks. The final result-copy and ordering tests pass
51 checks. Runner TypeScript and generated contracts pass. The pinned runtime
and Rust inputs are unchanged from `1f3af78114`. This remains deterministic
proof. Clean-head live billing, fresh cloud CI and review remain required.
### 2026-10-08 governed confirmation correction
The clean `d54a98895c` Mac question workflow passes in 71.476 seconds. A second
workflow preserves the exact pending interaction across a controller restart
and a fresh browser document, then resumes Hermes to completion in 107.945
seconds. Both have zero automatic retries and passing cleanup. Each has two
settled, reported OpenRouter receipts with healthy configured budgets. Their
reported totals are $0.002000649 and $0.001891331. Inspected screenshots show
the pending form, Cobalt answer, thought card, one final response, Done state
and available composer. This proves assigned Paperclip human input, not live
native `clarify_callback`. That head has 54 passing cloud checks, two skips,
and fresh Greptile 5/5 with no unresolved threads. Its credential-free Linux
job verifies the recorded closure and passes four native fixtures.
A separate clean-head planning workflow exposes the same early-stop problem
for `request_confirmation`. The plan is accepted, but its first run remains
pending and unpriced; the company pauses and the wake fails. The owned test
launcher was cancelled after preserving that public state. Its original
243.189-second infrastructure failure and `not_started` cleanup verdict remain
unchanged. Independent checks confirm no owned process group, listener or
temporary instance remains. Unknown spend stays fully reserved.
The bridge and both event boundaries now accept the three closed canonical
human-input kinds: questions, confirmations and checkbox confirmations. Each
must be an applied pending result from assigned `request_human_input`, with
`wake_assignee` and the original identity. Other kinds, tools, resolved results
and prose cannot stop work. This delays the fact only; the controller retains
wait, approval and accounting authority. The internal launch policy is named
for human input. The native regression covers immediate shutdown for each
kind. The bridge-only closure update retains verified dependencies and the
interpreter; the shared account cache is unchanged. Fresh live confirmation
proof, cloud CI and review remain required before claiming this correction
qualified. Actual Daytona and the broader release gates remain pending.
The correction passes 57 focused ordering/cancellation checks, 315 shared
Runner regressions and seven distribution fixtures. Runner TypeScript and
generated protocol/sidecar contracts pass. All six credential-free native
fixtures pass in 87.507 seconds, including 38 pinned Python checks and immediate
shutdown with a complete usage delta for each of the three human-input kinds.
These fixtures do not establish live approval, billing or Daytona proof.
+3 -2
View File
@@ -135,11 +135,12 @@ extend the task's execution deadline.
authority. Missing, invalid or timed-out reads do not invent usage or cost.
This optional execution-result field carries an existing v1 PRP event; it
does not change persisted execution inputs or the Rust wire contract.
- An applied pending `request_human_input` response stops Hermes in its native
- An applied pending `request_human_input` response for a question, confirmation
or checkbox confirmation stops Hermes in its native
tool completion callback, before another model request. The bridge publishes
that tool's completed result after native finalization. Both Runner event
pumps read the prompt usage receipt before forwarding this completion to the
controller. The tool bridge also defers its own completed question fact until
controller. The tool bridge also defers its own completed human-input fact until
the native terminal notification, which follows that receipt. The managed
Hermes profile selects this launch policy; other profiles retain their
existing order. Other tool activity still streams immediately. Tool identities
@@ -23,8 +23,8 @@ const completion = {
evidence: [], verification: [{ commandOrCheck: 'native transport fixture', status: 'passed' }], attentionRequests: [], artifacts: [],
};
for (const question of [false, true]) test(question
? 'pinned Hermes retains prompt usage before a committed question triggers immediate Rust-sidecar shutdown'
for (const interactionKind of [null, "ask_user_questions", "request_confirmation", "request_checkbox_confirmation"]) test(interactionKind
? `pinned Hermes retains prompt usage before committed ${interactionKind} triggers immediate Rust-sidecar shutdown`
: 'pinned Hermes streams images and semantic completion through Rust PRP and the production sidecar', {
skip: process.env.PAPERCLIP_HERMES_QUALIFY !== '1', timeout: 180_000,
}, async t => {
@@ -52,10 +52,10 @@ for (const question of [false, true]) test(question
if (!body.messages.some(message => message.role === 'tool')) {
send({ role: 'assistant', reasoning_content: 'Checking the native runner path.' });
await new Promise(resolve => setTimeout(resolve, 80));
const finish = body.tools.find(tool => tool.function.name.endsWith(question ? 'request_human_input' : 'paperclip_finish'));
const finish = body.tools.find(tool => tool.function.name.endsWith(interactionKind ? 'request_human_input' : 'paperclip_finish'));
assert.ok(finish, 'Paperclip completion authority was not exposed to the native model');
send({ tool_calls: [{ index: 0, id: 'native-prp-finish', type: 'function', function: { name: finish.function.name,
arguments: JSON.stringify(question ? { marker: 'governed-question' } : completion) } }] });
arguments: JSON.stringify(interactionKind ? { marker: 'governed-question' } : completion) } }] });
send({}, 'tool_calls');
} else {
send({ role: 'assistant', content: 'Native Rust ' });
@@ -109,12 +109,12 @@ for (const question of [false, true]) test(question
prpIdentity: { runnerInstanceId: 'hermes-prp-runner', environmentLeaseId: 'hermes-prp-lease', runId: identity.runId, normalizedSessionId: identity.sessionId, turnId: 'hermes-prp-turn', itemId: 'hermes-prp-item' },
});
const backend = createRunnerdNativeSessionBackend(input, { runnerInstanceId: 'hermes-prp-runner', transportFactory: () => bundle.transport,
...(question ? {
dynamicTools: [{ name: 'request_human_input', description: 'Simulated authenticated question operation',
...(interactionKind ? {
dynamicTools: [{ name: 'request_human_input', description: 'Simulated authenticated human input operation',
inputSchema: { type: 'object', properties: { marker: { type: 'string' } }, required: ['marker'] } }],
dynamicToolHandler: async () => ({ disposition: 'applied', interaction: {
id: 'fixture-question', companyId: identity.companyId, issueId: identity.issueId, sourceRunId: identity.runId,
kind: 'ask_user_questions', status: 'pending', continuationPolicy: 'wake_assignee',
kind: interactionKind, status: 'pending', continuationPolicy: 'wake_assignee',
} }),
} : {}),
});
@@ -124,16 +124,16 @@ for (const question of [false, true]) test(question
let cancellationOutcome;
for await (const event of session.events()) {
events.push(event);
if (question && event.eventType === 'item.completed' && event.payload.kind === 'dynamicToolCall') {
cancellationOutcome = session.cancel({ reason: 'committed question wait', signal: new AbortController().signal }).cleanup
if (interactionKind && event.eventType === 'item.completed' && event.payload.kind === 'dynamicToolCall') {
cancellationOutcome = session.cancel({ reason: 'committed human input wait', signal: new AbortController().signal }).cleanup
.then(() => ({ stopped: true }), error => ({ error }));
break;
}
if (['run.terminal', 'turn.completed', 'turn.failed', 'turn.cancelled', 'turn.interrupted'].includes(event.eventType)) break;
}
if (question) {
assert.ok(cancellationOutcome, 'The committed question completion did not reach the native boundary');
await session.close({ reason: 'immediate question shutdown' });
if (interactionKind) {
assert.ok(cancellationOutcome, 'The committed human input completion did not reach the native boundary');
await session.close({ reason: 'immediate human input shutdown' });
const stop = await cancellationOutcome;
if (stop.error) assert.equal(stop.error.code, 'already_terminal', 'Provider cancellation failed before terminal settlement');
const usage = await session.accountingUsageEvent?.();
@@ -141,8 +141,8 @@ for (const question of [false, true]) test(question
input: usage.payload.usage.runDelta.inputTokens, output: usage.payload.usage.runDelta.outputTokens,
complete: usage.payload.usage.runDeltaComplete,
}, { input: 10, output: 5, complete: true }, 'Immediate shutdown lost the native prompt receipt');
assert.ok(events.some(event => event.payload.kind === 'usage'), 'The question completion reached PRP before its usage receipt');
assert.equal(requests.length, 1, 'Hermes started another model request after the committed question');
assert.ok(events.some(event => event.payload.kind === 'usage'), 'The human input completion reached PRP before its usage receipt');
assert.equal(requests.length, 1, 'Hermes started another model request after the committed human input');
return;
}
assert.ok(events.some(event => event.eventType === 'turn.completed'), JSON.stringify(events));
@@ -181,7 +181,7 @@ function createTransportBackedNativeSessionBackend(
transportFactory: options.transportFactory,
dynamicTools: options.dynamicTools,
dynamicToolHandler: options.dynamicToolHandler,
deferCommittedQuestionResults: input.provider.kind === "acpx" && input.provider.agent === "hermes",
deferCommittedHumanInputResults: input.provider.kind === "acpx" && input.provider.agent === "hermes",
completionFeedback: options.completionFeedback,
environment: options.environment,
workingDirectoryAuthority: options.workingDirectoryAuthority,
@@ -15,7 +15,7 @@ import type {
AcpPermissionDecision,
} from "acpx/runtime";
import { acpxProfileClientCapabilities, bindAcpxExtensionTurn, validateAcpxRichEvent, createAcpxProfileExtensionAdapter, isHermesCommittedQuestionCompletion, type AcpxExtensionInput } from "../drivers/acpx/profile-extensions.js";
import { acpxProfileClientCapabilities, bindAcpxExtensionTurn, validateAcpxRichEvent, createAcpxProfileExtensionAdapter, isHermesCommittedHumanInputCompletion, type AcpxExtensionInput } from "../drivers/acpx/profile-extensions.js";
import type { PaperclipQuestionSet } from "../contracts/question-set.js";
import { readProviderUsageBilling, type ProviderUsageBilling } from "../contracts/usage-billing.js";
import { createAcpxToolEventNormalizer, createGrokMessageNormalizer } from "../provider-events.js";
@@ -635,7 +635,7 @@ async function pumpTurn(
toolEvidence?.tool(event);
const normalized = normalizeMessage(normalizeToolEvent(boundRuntimeEventForNormalization(event)));
if (openParams?.agent === "hermes" && event.type === "tool_call"
&& isHermesCommittedQuestionCompletion({ ...normalized, rawOutput: event.rawOutput }, runId ?? "")) {
&& isHermesCommittedHumanInputCompletion({ ...normalized, rawOutput: event.rawOutput }, runId ?? "")) {
if (committedQuestions.length >= MAX_PENDING_INPUTS) throw new Error("Hermes committed question limit exceeded");
committedQuestions.push(normalized);
continue;
@@ -1,7 +1,7 @@
import { isProviderMode } from "../../contracts/provider-mode.js";
import { acpxProfileActivity, type AcpxActivityAdapter, type AcpxToolEvidence } from "./profile-activity.js";
import { requireAcpxResponseDelivery } from "./response-delivery.js";
import { acpxProfileClientCapabilities, bindAcpxExtensionTurn, validateAcpxRichEvent, createAcpxProfileExtensionAdapter, isHermesCommittedQuestionCompletion, type AcpxExtensionInput } from "./profile-extensions.js";
import { acpxProfileClientCapabilities, bindAcpxExtensionTurn, validateAcpxRichEvent, createAcpxProfileExtensionAdapter, isHermesCommittedHumanInputCompletion, type AcpxExtensionInput } from "./profile-extensions.js";
import { createHash, randomBytes } from "node:crypto";
import type {
@@ -1467,7 +1467,7 @@ class CodexAcpxSession implements HarnessSession {
const projected = activity.toolExecutionId && event.type === "tool_call" && typeof event.toolCallId === "string"
? { ...event, toolCallId: activity.toolExecutionId(event.toolCallId) } : event;
const normalized = normalizeMessage(normalizeToolEvent(projected));
if (this.#agent === "hermes" && isHermesCommittedQuestionCompletion(normalized, this.#input.runId)) {
if (this.#agent === "hermes" && isHermesCommittedHumanInputCompletion(normalized, this.#input.runId)) {
if (committedQuestions.length >= 16) throw new Error("Hermes committed question limit exceeded");
committedQuestions.push(normalized);
continue;
@@ -1,5 +1,5 @@
/** Reviewed native closures. Setup cannot adopt a digest from downloaded files. */
export const HERMES_CLOSURES: Readonly<Record<string, string>> = Object.freeze({
"darwin-arm64": "f443a6867c914f49d7308c7421dbce9ba4a0d024b0c718a41a78845c784a7bd0",
"linux-x64": "c9879c3b69357d17d375d8fd4897f18496cebb21256c1e8bb9220d779d995921",
"darwin-arm64": "2230a296b80cb79079f0e6223449e18affb0e9543a476fb5901fb0524f18a0f5",
"linux-x64": "5c49a3018d6f0875bdbc318f934d05814db9c40add9d91b5e834566b3b8b32b4",
});
@@ -1,14 +1,14 @@
import { describe, expect, it } from "vitest";
import { createAcpxProfileExtensionAdapter, isHermesCommittedQuestionCompletion, validateAcpxRichEvent } from "./profile-extensions.js";
import { createAcpxProfileExtensionAdapter, isHermesCommittedHumanInputCompletion, validateAcpxRichEvent } from "./profile-extensions.js";
describe("Hermes native extensions", () => {
const committed = { disposition: "applied", interaction: { id: "question", companyId: "company", issueId: "issue",
sourceRunId: "run", kind: "ask_user_questions", status: "pending", continuationPolicy: "wake_assignee" } };
const completion = { type: "tool_call", tag: "tool_call_update", status: "completed", title: "mcp__paperclip__request_human_input" };
it("delays only the current run's structured committed question completion", () => {
it.each(["ask_user_questions", "request_confirmation", "request_checkbox_confirmation"])("delays only the current run's structured committed human input completion (%s)", kind => {
const committed = { disposition: "applied", interaction: { id: "input", companyId: "company", issueId: "issue",
sourceRunId: "run", kind, status: "pending", continuationPolicy: "wake_assignee" } };
for (const rawOutput of [committed, JSON.stringify(committed), { result: committed },
{ result: JSON.stringify(committed) }, { result: "Saved", structuredContent: committed }]) {
expect(isHermesCommittedQuestionCompletion({ ...completion, rawOutput }, "run")).toBe(true);
expect(isHermesCommittedHumanInputCompletion({ ...completion, rawOutput }, "run")).toBe(true);
}
for (const event of [{ ...completion, title: "terminal", rawOutput: committed },
{ ...completion, status: "failed", rawOutput: committed }, { ...completion, tag: "tool_call", rawOutput: committed },
@@ -17,10 +17,12 @@ describe("Hermes native extensions", () => {
{ ...completion, rawOutput: { ...committed, disposition: "rejected" } },
...["status", "kind", "continuationPolicy", "sourceRunId", "id", "companyId", "issueId"].map(key => ({
...completion, rawOutput: { ...committed, interaction: { ...committed.interaction, [key]: "" } },
}))]) expect(isHermesCommittedQuestionCompletion(event, "run")).toBe(false);
expect(isHermesCommittedQuestionCompletion({ ...completion, rawOutput: committed }, "other-run")).toBe(false);
expect(isHermesCommittedQuestionCompletion({ ...completion, title: completion.title + ": Choose the color", rawOutput: committed }, "run")).toBe(true);
expect(isHermesCommittedQuestionCompletion({ ...completion, title: completion.title + "_other", rawOutput: committed }, "run")).toBe(false);
}))]) expect(isHermesCommittedHumanInputCompletion(event, "run")).toBe(false);
expect(isHermesCommittedHumanInputCompletion({ ...completion, rawOutput: committed }, "other-run")).toBe(false);
expect(isHermesCommittedHumanInputCompletion({ ...completion, title: completion.title + ": Choose the color", rawOutput: committed }, "run")).toBe(true);
expect(isHermesCommittedHumanInputCompletion({ ...completion, title: completion.title + "_other", rawOutput: committed }, "run")).toBe(false);
expect(isHermesCommittedHumanInputCompletion({ ...completion, rawOutput: { ...committed,
interaction: { ...committed.interaction, kind: "future_interaction" } } }, "run")).toBe(false);
});
const adapter = () => createAcpxProfileExtensionAdapter("hermes", { sessionId: "session", turnId: "turn", workspacePath: "/workspace" })!;
it("keeps native child identities stable across start, progress and completion", async () => {
@@ -115,7 +115,7 @@ export function acpxProfileClientCapabilities(agent: QualifiedAcpxAgent): Record
/** Delay this tool's display completion until its native prompt receipt is read.
* This grants no wait or accounting authority; the controller validates both.
*/
export function isHermesCommittedQuestionCompletion(event: AcpRuntimeEventShape, runId: string): boolean {
export function isHermesCommittedHumanInputCompletion(event: AcpRuntimeEventShape, runId: string): boolean {
if (event.type !== "tool_call" || event.tag !== "tool_call_update" || event.status !== "completed"
|| !(event.title === "mcp__paperclip__request_human_input"
|| event.title?.startsWith("mcp__paperclip__request_human_input: "))) return false;
@@ -134,7 +134,9 @@ export function isHermesCommittedQuestionCompletion(event: AcpRuntimeEventShape,
}
if (!result || "error" in result || result.disposition !== "applied") return false;
const interaction = object(result.interaction);
return interaction !== null && interaction.kind === "ask_user_questions" && interaction.status === "pending"
return interaction !== null && typeof interaction.kind === "string"
&& ["ask_user_questions", "request_confirmation", "request_checkbox_confirmation"].includes(interaction.kind)
&& interaction.status === "pending"
&& interaction.continuationPolicy === "wake_assignee" && interaction.sourceRunId === runId
&& ["id", "companyId", "issueId"].every(key => typeof interaction[key] === "string" && Boolean(interaction[key]));
}
@@ -948,7 +948,7 @@ export class CodexAppServerDriver implements HarnessDriver {
skillInputs: this.#options.skillInputs,
reasoningEffort: this.#options.reasoningEffort,
dynamicToolHandler: this.#options.dynamicToolHandler,
deferCommittedQuestionResults: this.#options.deferCommittedQuestionResults,
deferCommittedHumanInputResults: this.#options.deferCommittedHumanInputResults,
completionFeedback: this.#options.completionFeedback,
});
}
@@ -44,11 +44,11 @@ import {
} from "./codex-app-server-driver.test-support.js";
describe("Codex app-server Codex driver", () => {
it.each([true, false])("orders committed bridge question results after native usage only for the managed policy (%s)", async defer => {
it.each([true, false].flatMap(defer => ["ask_user_questions", "request_confirmation", "request_checkbox_confirmation"].map(kind => ({ defer, kind }))))("orders committed bridge human input results after native usage only for the managed policy ($kind, $defer)", async ({ defer, kind }) => {
const transport = new FakeCodexTransport();
const committed = { disposition: "applied", interaction: { id: "question", companyId: "company", issueId: "issue",
sourceRunId: "run-question", kind: "ask_user_questions", status: "pending", continuationPolicy: "wake_assignee" } };
const session = await makeDriver([transport], { deferCommittedQuestionResults: defer,
sourceRunId: "run-question", kind, status: "pending", continuationPolicy: "wake_assignee" } };
const session = await makeDriver([transport], { deferCommittedHumanInputResults: defer,
dynamicTools: [{ name: "request_human_input", inputSchema: { type: "object" } }],
dynamicToolHandler: async () => committed,
}).openSession({ runId: "run-question", normalizedSessionId: "normalized-question", workingDirectory: WORKSPACE });
@@ -48,8 +48,8 @@ export interface CodexAppServerDriverOptions {
}) => CodexAppServerTransport;
/** Additional control-plane tools exposed to the provider for this run. */
dynamicTools?: readonly Readonly<Record<string, unknown>>[];
/** Native Hermes stops at a committed question and sends final usage first. */
deferCommittedQuestionResults?: boolean;
/** Native Hermes stops at committed human input and sends final usage first. */
deferCommittedHumanInputResults?: boolean;
/** Executes an admitted additional tool call. Completion tools remain driver-owned. */
dynamicToolHandler?: (call: {
tool: string;
@@ -123,7 +123,7 @@ export class CodexHarnessSession
}
this.runId = input.runId;
this.lastAccountingUsageEvent = null;
this.committedQuestionResults.clear();
this.committedHumanInputResults.clear();
this.result = null;
this.resultFingerprint = null;
this.resultCallId = null;
@@ -1,4 +1,4 @@
import { isAcpxCanonicalInputMethod, isHermesCommittedQuestionCompletion } from "../acpx/profile-extensions.js";
import { isAcpxCanonicalInputMethod, isHermesCommittedHumanInputCompletion } from "../acpx/profile-extensions.js";
import { isSemanticToolOutcomeUnknownError } from "../../contracts/native-session-backend.js";
import type { HarnessRuntimeRequest, PaperclipQuestionSet } from "../../contracts/harness-driver.js";
import {
@@ -142,11 +142,11 @@ async function handleServerRequestBody(
result,
},
};
if (state.deferCommittedQuestionResults && tool === "request_human_input"
&& isHermesCommittedQuestionCompletion({ type: "tool_call", tag: "tool_call_update", status: "completed",
if (state.deferCommittedHumanInputResults && tool === "request_human_input"
&& isHermesCommittedHumanInputCompletion({ type: "tool_call", tag: "tool_call_update", status: "completed",
title: "mcp__paperclip__request_human_input", rawOutput: result }, state.runId)) {
if (state.committedQuestionResults.size >= 16) throw new Error("Hermes committed question result limit exceeded");
state.committedQuestionResults.set(callId, { payload: structuredClone(completed), turnId, itemId: callId });
if (state.committedHumanInputResults.size >= 16) throw new Error("Hermes committed human input result limit exceeded");
state.committedHumanInputResults.set(callId, { payload: structuredClone(completed), turnId, itemId: callId });
} else {
state.emit("item.completed", completed, { turnId, itemId: callId });
}
@@ -105,8 +105,8 @@ export class CodexSessionState {
readonly dynamicTools: readonly Readonly<Record<string, unknown>>[];
readonly completionFeedback: CodexAppServerDriverOptions["completionFeedback"];
readonly dynamicToolHandler: CodexAppServerDriverOptions["dynamicToolHandler"];
readonly deferCommittedQuestionResults: boolean;
readonly committedQuestionResults = new Map<string, { payload: Record<string, unknown>; turnId: string; itemId: string }>();
readonly deferCommittedHumanInputResults: boolean;
readonly committedHumanInputResults = new Map<string, { payload: Record<string, unknown>; turnId: string; itemId: string }>();
readonly eventQueue = new AsyncQueue<PrpEvent>();
sourceSequence: number;
activeTurnId: string | null;
@@ -182,7 +182,7 @@ export class CodexSessionState {
dynamicTools: readonly Readonly<Record<string, unknown>>[];
completionFeedback?: CodexAppServerDriverOptions["completionFeedback"];
dynamicToolHandler?: CodexAppServerDriverOptions["dynamicToolHandler"];
deferCommittedQuestionResults?: boolean;
deferCommittedHumanInputResults?: boolean;
}) {
this.codexUsageBaseline = input.codexUsageBaseline ?? null;
if (this.codexUsageBaseline) this.usageSnapshot = codexRunUsage(this.codexUsageBaseline);
@@ -207,7 +207,7 @@ export class CodexSessionState {
this.reasoningEffort = input.reasoningEffort;
this.dynamicTools = input.dynamicTools;
this.dynamicToolHandler = input.dynamicToolHandler;
this.deferCommittedQuestionResults = input.deferCommittedQuestionResults === true;
this.deferCommittedHumanInputResults = input.deferCommittedHumanInputResults === true;
this.completionFeedback = input.completionFeedback;
this.currentGoal = input.goal === undefined ? null : structuredClone(input.goal);
for (const entry of input.lineage ?? [input.opened.lineage]) {
@@ -443,9 +443,9 @@ export class CodexSessionState {
// Native terminal notification follows the final prompt receipt. Release
// the bridge's own completed tool facts here, before publishing terminal,
// so controller parking cannot interrupt delivery of that receipt.
for (const [id, completed] of this.committedQuestionResults) {
for (const [id, completed] of this.committedHumanInputResults) {
if (completed.turnId !== refs.turnId) continue;
this.committedQuestionResults.delete(id);
this.committedHumanInputResults.delete(id);
this.emit("item.completed", completed.payload, { turnId: completed.turnId, itemId: completed.itemId });
}
}
@@ -58,7 +58,7 @@ def committed_human_input_wait(name, result):
return False
interaction = result.get("interaction")
return (result.get("disposition") == "applied" and isinstance(interaction, dict)
and interaction.get("kind") == "ask_user_questions"
and interaction.get("kind") in ("ask_user_questions", "request_confirmation", "request_checkbox_confirmation")
and interaction.get("status") == "pending"
and interaction.get("continuationPolicy") == "wake_assignee"
and all(isinstance(interaction.get(key), str) and bool(interaction[key])
@@ -268,8 +268,12 @@ class Controls(unittest.IsolatedAsyncioTestCase):
"status": "pending", "continuationPolicy": "wake_assignee"}}
name = "mcp__paperclip__request_human_input"
updates = []
for carrier in [result, {"result": result}, {"result": json.dumps(result)},
{"result": "Question saved", "structuredContent": result}]:
carriers = []
for kind in ("ask_user_questions", "request_confirmation", "request_checkbox_confirmation"):
committed = {**result, "interaction": {**result["interaction"], "kind": kind}}
carriers.extend([committed, {"result": committed}, {"result": json.dumps(committed)},
{"result": "Human input saved", "structuredContent": committed}])
for carrier in carriers:
with self.subTest(carrier=carrier):
self.state.cancel_event.clear()
self.bridge._committed_wait_updates.clear()
@@ -313,7 +317,7 @@ class Controls(unittest.IsolatedAsyncioTestCase):
("mcp__paperclip__request_human_input", {"result": "Saved a pending question"}),
("mcp__paperclip__request_human_input", {"result": original, "error": "denied"}),
("mcp__paperclip__request_human_input", {**original, "disposition": "rejected"})]
for key, value in [("status", "resolved"), ("kind", "request_confirmation"),
for key, value in [("status", "resolved"), ("kind", "future_interaction"), ("kind", []),
("continuationPolicy", "none"), ("sourceRunId", None), ("id", "")]:
cases.append(("mcp__paperclip__request_human_input",
{**original, "interaction": {**original["interaction"], key: value}}))