Files
PaperClipAI/server/src/services/native-runtime/native-execution-input.ts
T
DottaandPaperclip 2ec82c5774 fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use tools through company-scoped connections and provider
sessions.
> - Resolving a tool connection currently forces a fresh session even
when the provider can load new tools into the existing conversation.
> - A fresh provider conversation can receive too little history to
continue the task.
> - This pull request adds explicit tool-refresh capabilities and uses
them in both runner paths.
> - Fresh attempts receive bounded task history with source IDs and
retrieval instructions.
> - The benefit is that agents can continue the same task after a
connection changes.

## Linked Issues or Issue Description

**What happened?**

A resolved tool connection forced a fresh provider conversation. The new
conversation could lose the original goal and prior answers. Claude also
rejected resume when only the MCP server set changed.

**Expected behavior**

Resume the provider conversation when its harness can refresh tools.
When a fresh session is required, supply enough bounded history to
continue the task. Preserve company, agent, task, workspace, model,
instruction, and skill checks.

**Steps to reproduce**

1. Start a conversation and agree on a task and its constraints.
2. Request and connect a tool needed for the task.
3. Continue the conversation after the connection resolves.
4. Check that the agent remembers the task and can use the new tool.

**Paperclip version or commit**

The bug was reproduced on master at `c46e41e81`. This branch is rebased
on current master.

**Deployment mode**

Self-hosted server. Both legacy adapters and the native runner are
affected.

Related public work: Refs #13282 for task-backed conversations. Refs
#13057 for the broader session-compaction proposal. Refs #14659 for
another report about local CLI session continuity. This change fixes
tool-connection continuation. Provider authentication repairs keep their
existing recovery behavior.

## What Changed

- Expose tool-refresh support in native harness descriptors and legacy
adapter metadata.
- Request tool refresh after connection resolution. Keep provider
authentication repair as a fresh-session wake.
- Reload current tools and credentials while retaining supported Claude,
Codex, Grok, and other provider conversations.
- Allow MCP-only changes during qualified native recovery. Keep all
other compatibility checks.
- Refresh managed-provider and ACPX tool bindings when attaching a new
run.
- Add a fresh-session handoff for both runner paths. Bound database
reads, excerpts, and the final packet to 24,000 bytes.
- Include the original request, recent messages, decisions, plans, prior
answers, and source IDs. Mark omitted content. Apply reset boundaries,
wake cutoffs, quarantine, and secret redaction.
- Add regression tests and document the capabilities and handoff
behavior.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed.
- Final review fixes passed 314 server tests, 333 adapter utility tests,
139 native-session runtime tests, and 12 managed-provider Rust tests.
They verify historical quarantine, raised budgets across attachment, no
history reads on successful resume, and handoff delivery on fresh retry.
- Broader branch verification also passed 1,401 adapter utility tests,
1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core
tests.
- Live Claude CLI and Grok ACP probes preserved the provider session ID,
recalled a prior task constraint, and called a newly added read-only MCP
tool.
- GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55
successful checks and 4 skipped checks. This includes all test shards,
all eight browser shards, runner checks, and the Grok clean public npm
install canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37057514976).
- A full local test attempt encountered a separate Git snapshot timeout.
All affected local suites passed after the final edits, and the full CI
test gates passed.
- Review the capability matrix in `packages/paperclip-runner/README.md`.
Repeat the four reproduction steps with a supported provider and with an
unsupported harness.

## Risks

- Provider tool refresh can fail. Existing recovery falls back to a
fresh conversation where policy permits it.
- A new transport can replace an old process while preserving the
provider conversation. Tests cover current credentials and unchanged
identity.
- Long history can omit older context. Explicit markers and source IDs
let the agent retrieve needed context within task scope.
- Unknown and unqualified harnesses use the fresh-session path. No
database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live
provider testing. The exact model ID and context-window size are not
exposed in this session. Claude and Grok also ran as test subjects.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:11:44 -05:00

324 lines
14 KiB
TypeScript

import type { PaperclipTurnContext } from "@paperclipai/adapter-utils/server-utils";
import { createHash } from "node:crypto";
import { buildNativeContinuationPrompt } from "./native-continuation.js";
import type {
NativeAcpxAgent,
NativeAcpxPermissionMode,
NativeCodexApprovalPolicy,
NativeExecutionInputV5,
NativeCompletionSource,
NativeInteractionResponseEnvelope,
NativeOpenCodePermissionMode,
NativePlanningContext,
NativeRuntimeContextSnapshot,
StrictCompletionContractInput,
} from "../../vendor/paperclip-runner/index.js";
import {
parseNativeExecutionInput,
resolveQualifiedAcpxProfile,
} from "../../vendor/paperclip-runner/index.js";
import {
isPaperclipExternalChatContractTurn,
isPaperclipExternalChatQuestionResponseTurn,
selectPaperclipPromptSections,
} from "@paperclipai/adapter-utils/server-utils";
const NATIVE_GITHUB_ATTACHMENT_RECOVERY_GUIDANCE = [
"## GitHub attachment recovery navigation",
"Paperclip owns recovery navigation for unavailable GitHub attachments. It may append an authenticated task link after an accepted response, only when the current source remains authorized and a safe configured Board URL is available. The model does not select or authorize that link.",
"A task URL missing from your prompt or tool results is not evidence that no task link can be provided; do not claim that a link is unavailable merely because you cannot see its URL. Do not invent a URL or promise that a link will appear. Briefly explain the unavailable input and ask the user to attach it directly to this Paperclip task or paste the needed text. Never infer the file's contents or substitute an older file.",
].join("\n");
/** Closed constructor: callers cannot spread legacy context or environment data. */
export function buildNativeExecutionInput(input: {
companyId: string;
runId: string;
issue: {
id: string;
identifier: string | null;
title: string;
description: string | null;
workMode: string;
};
taskPrompt: string;
initialCommunicationGuidance?: string | null;
/** Bounded, redacted background restored only after a fresh provider bootstrap. */
freshSessionHandoff?: string | null;
/**
* The already-sanitized Paperclip wake envelope for this run. Native drivers
* receive a closed execution input rather than the legacy adapter context,
* so the constructor must deliberately project the same bounded wake delta
* that legacy adapters place in their provider prompt.
*/
wakePayload?: unknown;
/** Additive source ownership emitted by the server task/wake builders. */
turnContext?: unknown;
resumedSession?: boolean;
previousTurn?: { runId: string; task: { title: string; description: string | null } } | null;
conversationMode?: boolean;
agentId: string;
workspace: {
id: string;
cwd: string;
repoUrl: string | null;
repoRef: string | null;
branchName: string | null;
};
normalizedSessionId: string | null;
provider?: "codex" | "opencode" | "claude_managed" | "aws_agentcore" | "acpx";
acpxAgent?: NativeAcpxAgent;
codexApprovalPolicy?: NativeCodexApprovalPolicy;
codexReasoningEffort?: string;
opencodePermissionMode?: NativeOpenCodePermissionMode;
acpxPermissionMode?: NativeAcpxPermissionMode;
model?: string | null;
managedProfile?: Extract<
NativeExecutionInputV5["provider"],
{ kind: "claude_managed" }
>["managedProfile"];
maxSessionListCostUsd?: number;
agentCoreProfile?: Extract<
NativeExecutionInputV5["provider"],
{ kind: "aws_agentcore" }
>["agentCoreProfile"];
maxEstimatedSessionCostUsd?: number;
invocationLimits?: Extract<
NativeExecutionInputV5["provider"],
{ kind: "aws_agentcore" }
>["invocationLimits"];
lifecyclePolicy?: NativeExecutionInputV5["session"]["lifecyclePolicy"];
executionMode?: "default" | "plan";
planningContext?: NativePlanningContext | null;
interactionResponses?: NativeInteractionResponseEnvelope[];
completionContract: {
id: string;
sha256: string;
schemaVersion: string;
contract: StrictCompletionContractInput;
sources?: Array<{ id: string; source: NativeCompletionSource }>;
};
runtimeContext: NativeRuntimeContextSnapshot;
}): NativeExecutionInputV5 {
if (input.issue.workMode !== "standard" && input.issue.workMode !== "planning" && input.issue.workMode !== "ask") {
throw new Error("native_execution_input_invalid: issue work mode must be standard, planning, or ask");
}
const executionMode = input.executionMode
?? (input.issue.workMode === "planning" ? "plan" : "default");
const acpxProfile = input.provider === "acpx"
? resolveQualifiedAcpxProfile(
input.acpxAgent ?? "codex",
input.model ?? "",
)
: null;
// Answers are materialized from the authoritative interaction only for this
// invocation; do not persist a duplicate answer in the durable wake snapshot.
const wake =
input.wakePayload &&
typeof input.wakePayload === "object" &&
!Array.isArray(input.wakePayload)
? (input.wakePayload as Record<string, unknown>)
: null;
const question = wake?.externalChatQuestionResponse
? input.interactionResponses?.find(
(response) =>
response.interactionId === wake.interactionId &&
response.kind === "ask_user_questions" &&
response.response.status === "answered",
)
: null;
const answerResult = question?.response.result as
Record<string, unknown> | undefined;
// The server supplies only the revalidated answer chain, in source order.
// Keep prior choices available even when this continuation starts a fresh
// provider session; never recover them from model prose or a transcript.
const answerChain = question
? input.interactionResponses?.filter((response) =>
response.kind === "ask_user_questions" &&
response.response.status === "answered" &&
typeof (response.response.result as Record<string, unknown> | undefined)
?.summaryMarkdown === "string")
: null;
const answerSummary =
answerChain && answerChain.length > 1 &&
answerChain.at(-1)?.interactionId === question?.interactionId
? answerChain.map((response, index) => {
const label = index === answerChain.length - 1
? "Latest answered question" : `Earlier answer ${index + 1}`;
return `${label}:\n${(response.response.result as Record<string, unknown>).summaryMarkdown}`;
}).join("\n\n")
: answerResult?.summaryMarkdown;
const wakePayload =
question && typeof answerSummary === "string"
? {
...wake,
questionResponse: {
interactionId: question.interactionId,
summaryMarkdown: answerSummary,
},
}
: input.wakePayload;
// Build the full bootstrap through the same owner as legacy adapters.
// Verified native resume selection stays at the existing session boundary.
const { taskContextNote, wakePrompt } = selectPaperclipPromptSections({
paperclipTaskMarkdownAssignment: input.taskPrompt,
paperclipWake: wakePayload,
conversationMode: input.conversationMode,
}, {
resumedSession: false,
includeCommunicationGuidance: false,
nativeWakeReaderAvailable: true,
});
const externalChatTurn =
isPaperclipExternalChatContractTurn(wakePayload) ||
isPaperclipExternalChatQuestionResponseTurn(wakePayload);
const taskPrompt = [
wakePrompt,
// Durable task questions must survive the current provider turn.
// Keep routing visible before deferred tool discovery; usage belongs in the tool schema.
"Use Paperclip's request_human_input for durable task questions.",
externalChatTurn && wake?.externalChatProvider === "github"
? NATIVE_GITHUB_ATTACHMENT_RECOVERY_GUIDANCE
: "",
taskContextNote,
]
.filter((section) => section.length > 0)
.join("\n\n");
const completionSources = !externalChatTurn && !input.conversationMode && input.taskPrompt.trim()
? verifiedCompletionSources(input.turnContext, input.completionContract.sources ?? [])
: [];
return parseNativeExecutionInput({
schema: "paperclip.native-execution-input.v5",
...((input.initialCommunicationGuidance || input.freshSessionHandoff) ? {
initialCommunicationGuidance: [input.initialCommunicationGuidance, input.freshSessionHandoff].filter(Boolean).join("\n\n"),
} : {}),
...(input.resumedSession && input.previousTurn && !input.conversationMode ? {
continuationPrompt: buildNativeContinuationPrompt({
wakePayload: input.wakePayload,
previousRunId: input.previousTurn.runId,
previousIssue: input.previousTurn.task,
allowExternalChat: true,
issue: input.issue,
}),
} : {}),
executionMode,
planningContext: input.planningContext ?? null,
binding: {
companyId: input.companyId,
runId: input.runId,
issueId: input.issue.id,
agentId: input.agentId,
executionWorkspaceId: input.workspace.id,
},
task: {
identifier: input.issue.identifier ?? input.issue.id,
// The issue title is durable background context and may itself contain an
// exact-output instruction from the thread's first message. Repeating it
// as the native turn title can override a newer provider message in small
// models. Keep the canonical title and description in task.prompt as
// explicitly labeled background, but give authenticated external-chat
// turns neutral structured fields.
title: externalChatTurn ? "External chat follow-up" : input.issue.title,
description: externalChatTurn ? null : input.issue.description,
prompt: taskPrompt,
workMode: input.issue.workMode,
},
workspace: {
cwd: input.workspace.cwd,
repoUrl: input.workspace.repoUrl,
repoRef: input.workspace.repoRef,
branchName: input.workspace.branchName,
},
session: {
normalizedSessionId: input.normalizedSessionId,
driverKind: input.provider === "opencode"
? "opencode_server"
: input.provider === "claude_managed"
? "claude_managed_agents_api"
: input.provider === "aws_agentcore"
? "aws_agentcore_harness_api"
: input.provider === "acpx"
? "acpx_runtime"
: "codex_app_server",
protocolVersion: 1,
lifecyclePolicy: input.lifecyclePolicy ?? { mode: "per_turn", idleTimeoutMs: null },
},
provider: input.provider === "claude_managed"
? {
kind: "claude_managed",
model: input.model,
managedProfile: input.managedProfile,
maxSessionListCostUsd: input.maxSessionListCostUsd,
}
: input.provider === "aws_agentcore"
? {
kind: "aws_agentcore",
model: input.model,
agentCoreProfile: input.agentCoreProfile,
maxEstimatedSessionCostUsd: input.maxEstimatedSessionCostUsd,
invocationLimits: input.invocationLimits,
}
: input.provider === "acpx"
? {
kind: "acpx",
agent: acpxProfile!.agent,
model: input.model,
permissionMode: input.acpxPermissionMode ?? "approve-all",
profile: {
driverKind: acpxProfile!.driverKind,
protocolVersion: acpxProfile!.protocolVersion,
acpxVersion: acpxProfile!.acpxVersion,
agent: acpxProfile!.agent,
agentProfileVersion: acpxProfile!.agentProfileVersion,
agentServerPackage: acpxProfile!.agentServerPackage,
agentServerVersion: acpxProfile!.agentServerVersion,
agentRuntimePackage: acpxProfile!.agentRuntimePackage,
agentRuntimeVersion: acpxProfile!.agentRuntimeVersion,
commandDigest: acpxProfile!.commandDigest,
},
}
: input.provider === "opencode"
? {
kind: "opencode",
model: input.model,
permissionMode: input.opencodePermissionMode ?? "allow",
}
: {
kind: "codex",
model: input.model ?? null,
approvalPolicy: input.codexApprovalPolicy ?? "never",
...(input.codexReasoningEffort ? { reasoningEffort: input.codexReasoningEffort } : {}),
},
completionContract: {
id: input.completionContract.id,
sha256: input.completionContract.sha256,
schemaVersion: input.completionContract.schemaVersion,
contract: input.completionContract.contract,
},
...(completionSources.length ? { completionSources: {
promptSha256: createHash("sha256").update(taskPrompt).digest("hex"),
contractRevision: input.completionContract.contract.revision,
criteria: completionSources,
} } : {}),
interactionResponses: input.interactionResponses ?? [],
credentialBindings: [],
runtimeContext: input.runtimeContext,
}) as NativeExecutionInputV5;
}
/** Verify source identities and revisions, never guess provenance from requirement text. */
function verifiedCompletionSources(
value: unknown,
sources: Array<{ id: string; source: NativeCompletionSource }>,
): Array<{ id: string; source: NativeCompletionSource }> {
if (!value || typeof value !== "object" || Array.isArray(value)) return [];
const context = value as Partial<PaperclipTurnContext>;
if (context.version !== 1) return [];
return sources.filter(({ source }) => {
if (source.kind === "description") {
return context.assignment?.owner === "task_markdown" && context.assignment.description?.id === source.id && context.assignment.description.revision === source.revision;
}
return context.events?.owner === "wake_prompt" && Array.isArray(context.events.comments) && context.events.comments.some((comment) => comment.id === source.id && comment.revision === source.revision);
});
}