Files
PaperClipAI/server/src/__tests__/legacy-runtime-tool-refresh.test.ts
T
DottaandPaperclip 2ec82c5774 fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use tools through company-scoped connections and provider
sessions.
> - Resolving a tool connection currently forces a fresh session even
when the provider can load new tools into the existing conversation.
> - A fresh provider conversation can receive too little history to
continue the task.
> - This pull request adds explicit tool-refresh capabilities and uses
them in both runner paths.
> - Fresh attempts receive bounded task history with source IDs and
retrieval instructions.
> - The benefit is that agents can continue the same task after a
connection changes.

## Linked Issues or Issue Description

**What happened?**

A resolved tool connection forced a fresh provider conversation. The new
conversation could lose the original goal and prior answers. Claude also
rejected resume when only the MCP server set changed.

**Expected behavior**

Resume the provider conversation when its harness can refresh tools.
When a fresh session is required, supply enough bounded history to
continue the task. Preserve company, agent, task, workspace, model,
instruction, and skill checks.

**Steps to reproduce**

1. Start a conversation and agree on a task and its constraints.
2. Request and connect a tool needed for the task.
3. Continue the conversation after the connection resolves.
4. Check that the agent remembers the task and can use the new tool.

**Paperclip version or commit**

The bug was reproduced on master at `c46e41e81`. This branch is rebased
on current master.

**Deployment mode**

Self-hosted server. Both legacy adapters and the native runner are
affected.

Related public work: Refs #13282 for task-backed conversations. Refs
#13057 for the broader session-compaction proposal. Refs #14659 for
another report about local CLI session continuity. This change fixes
tool-connection continuation. Provider authentication repairs keep their
existing recovery behavior.

## What Changed

- Expose tool-refresh support in native harness descriptors and legacy
adapter metadata.
- Request tool refresh after connection resolution. Keep provider
authentication repair as a fresh-session wake.
- Reload current tools and credentials while retaining supported Claude,
Codex, Grok, and other provider conversations.
- Allow MCP-only changes during qualified native recovery. Keep all
other compatibility checks.
- Refresh managed-provider and ACPX tool bindings when attaching a new
run.
- Add a fresh-session handoff for both runner paths. Bound database
reads, excerpts, and the final packet to 24,000 bytes.
- Include the original request, recent messages, decisions, plans, prior
answers, and source IDs. Mark omitted content. Apply reset boundaries,
wake cutoffs, quarantine, and secret redaction.
- Add regression tests and document the capabilities and handoff
behavior.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed.
- Final review fixes passed 314 server tests, 333 adapter utility tests,
139 native-session runtime tests, and 12 managed-provider Rust tests.
They verify historical quarantine, raised budgets across attachment, no
history reads on successful resume, and handoff delivery on fresh retry.
- Broader branch verification also passed 1,401 adapter utility tests,
1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core
tests.
- Live Claude CLI and Grok ACP probes preserved the provider session ID,
recalled a prior task constraint, and called a newly added read-only MCP
tool.
- GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55
successful checks and 4 skipped checks. This includes all test shards,
all eight browser shards, runner checks, and the Grok clean public npm
install canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37057514976).
- A full local test attempt encountered a separate Git snapshot timeout.
All affected local suites passed after the final edits, and the full CI
test gates passed.
- Review the capability matrix in `packages/paperclip-runner/README.md`.
Repeat the four reproduction steps with a supported provider and with an
unsupported harness.

## Risks

- Provider tool refresh can fail. Existing recovery falls back to a
fresh conversation where policy permits it.
- A new transport can replace an old process while preserving the
provider conversation. Tests cover current credentials and unchanged
identity.
- Long history can omit older context. Explicit markers and source IDs
let the agent retrieve needed context within task scope.
- Unknown and unqualified harnesses use the fresh-session path. No
database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live
provider testing. The exact model ID and context-window size are not
exposed in this session. Claude and Grok also ran as test subjects.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:11:44 -05:00

74 lines
5.0 KiB
TypeScript

import fs from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { afterEach, describe, expect, it, vi } from "vitest";
import type { AdapterExecutionContext } from "@paperclipai/adapter-utils";
// Pi resolves its session directory at module load. Keep that test directory
// under the temporary root instead of the operator's real home.
vi.mock("node:os", async (importOriginal) => {
const actual = await importOriginal<typeof import("node:os")>();
return { ...actual, default: { ...actual.default, homedir: actual.tmpdir }, homedir: actual.tmpdir };
});
vi.mock("@paperclipai/adapter-utils/execution-target", async (importOriginal) => {
const actual = await importOriginal<typeof import("@paperclipai/adapter-utils/execution-target")>();
return { ...actual, runAdapterExecutionTargetProcess: vi.fn() };
});
import { runAdapterExecutionTargetProcess } from "@paperclipai/adapter-utils/execution-target";
import { execute as gemini } from "@paperclipai/adapter-gemini-local/server";
import { execute as kimi } from "@paperclipai/adapter-kimi-local/server";
import { execute as cursor } from "@paperclipai/adapter-cursor-local/server";
import { execute as opencode } from "@paperclipai/adapter-opencode-local/server";
import { execute as pi } from "@paperclipai/adapter-pi-local/server";
const roots: string[] = [];
afterEach(async () => { vi.clearAllMocks(); await Promise.all(roots.splice(0).map(root => fs.rm(root, { recursive: true, force: true }))); });
describe("legacy environment tool access on resumed conversations", () => {
it.each([
["gemini_local", gemini, "--resume"], ["kimi_local", kimi, "-r"],
["cursor", cursor, "--resume"], ["opencode_local", opencode, "--session"], ["pi_local", pi, "--session"],
] as const)("refreshes %s without dropping its conversation", async (type, execute, resumeFlag) => {
const root = await fs.mkdtemp(path.join(os.tmpdir(), "paperclip-cli-tool-refresh-")); roots.push(root);
const command = path.join(root, "agent");
await fs.writeFile(command, "#!/bin/sh\nprintf 'provider model\\nopenai gpt-5\\n'\n", { mode: 0o755 });
const sessionId = type === "pi_local" ? path.join(root, "session.jsonl") : "existing-session";
if (type === "pi_local") await fs.writeFile(sessionId, JSON.stringify({ type: "session", cwd: root }) + "\n");
const stdout = type === "kimi_local"
? JSON.stringify({ role: "meta", type: "session.resume_hint", session_id: sessionId })
: type === "opencode_local"
? JSON.stringify({ type: "text", sessionID: sessionId, part: { text: "done" } })
: type === "pi_local"
? JSON.stringify({ type: "agent_end", messages: [{ role: "assistant", content: "done" }] })
: JSON.stringify({ type: "result", subtype: "success", session_id: sessionId, result: "done" });
const invocations: Array<{ args: string[]; env: Record<string, string> }> = [];
vi.mocked(runAdapterExecutionTargetProcess).mockImplementation(async (_runId, _target, _command, args, options) => {
invocations.push({ args, env: options.env ?? {} });
return { exitCode: 0, signal: null, timedOut: false, stdout, stderr: "", pid: null, startedAt: null };
});
const getFreshSessionHandoff = vi.fn(async () => "FRESH_HANDOFF_ONLY");
const ctx: AdapterExecutionContext = {
getFreshSessionHandoff,
runId: "refresh", agent: { id: "agent", companyId: "company", name: "Agent", adapterType: type, adapterConfig: {} },
runtime: { sessionId, sessionParams: { sessionId, cwd: root }, sessionDisplayId: sessionId, taskKey: null },
config: { command, cwd: root, model: "openai/gpt-5", engine: "cli", env: { HOME: root, XDG_CONFIG_HOME: root, OPENCODE_ALLOW_ALL_MODELS: "1" } },
context: { refreshTools: true, paperclipFreshSessionHandoffMarkdown: "FRESH_HANDOFF_ONLY" }, onLog: async () => {},
runtimeTools: { version: 1, guidance: "Tools available", mcpEndpoint: "https://example.test/current-tools/mcp",
rest: { connectionsSearch: "https://example.test/search", connectionRequest: "https://example.test/request" },
bearerToken: "current-tool-token", expiresAt: "2027-01-01T00:00:00Z", tools: ["connections_search", "connection_request"] },
};
const result = await execute(ctx);
expect(result.exitCode).toBe(0);
const invocation = invocations.at(-1)!;
expect(invocation.args[invocation.args.indexOf(resumeFlag) + 1]).toBe(sessionId);
expect(invocation.env.PAPERCLIP_RUNTIME_TOOLS_MCP_URL).toBe(ctx.runtimeTools!.mcpEndpoint);
expect(invocation.env.PAPERCLIP_RUNTIME_TOOLS_TOKEN).toBe("current-tool-token");
expect(invocation.args.join(" ")).not.toContain("FRESH_HANDOFF_ONLY");
expect(result.sessionParams?.sessionId).toBe(sessionId);
// Revocation likewise gets the current environment on the same session.
await execute({ ...ctx, runtimeTools: undefined });
expect(invocations.at(-1)!.env.PAPERCLIP_RUNTIME_TOOLS_MCP_URL).toBeUndefined();
expect(getFreshSessionHandoff).not.toHaveBeenCalled();
});
});