mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local adapters are the bridge between Paperclip's control plane and provider CLIs such as Claude Code and Codex. > - Those adapters can run either on the host machine or inside a remote/sandbox execution target. > - Sandbox probes need to validate the same auth/config path that real sandbox execution will use. > - The previous probe paths could surface misleading Claude errors, rely on host-only Codex state, or upload far more Codex home state than the probe needed. > - This pull request fixes the Claude and Codex sandbox probe/runtime behavior together while keeping provider-specific sandbox image work out of scope. > - The benefit is faster, clearer adapter health checks that better match real sandbox execution. ## Linked Issues or Issue Description No public GitHub issue was found for this exact bug during duplicate search. Bug report: **What happened?** Sandboxed Claude/Codex adapter tests could diverge from real runtime auth/config behavior. Claude sandbox probes could show the leading stream init line instead of the real final error, and Codex sandbox probes could upload full managed home state or mask a sandbox-local login with an empty uploaded `CODEX_HOME`. **Expected behavior** Sandbox probes should exercise the remote runtime contract, preserve useful sandbox credentials, avoid relying on unrelated host state, and report actionable probe failures. **Steps to reproduce** 1. Configure a remote/sandbox execution target for `claude_local` or `codex_local`. 2. Run the environment Test/probe path where host credentials differ from the sandbox's runtime credentials or the managed Codex home contains session history. 3. Observe that probe behavior can differ from the actual sandbox runtime path or surface an unhelpful Claude stream initialization line. **Paperclip version or commit** Current `master` before this PR, based on `4a2447da3`. **Deployment mode** Local development/control-plane deployment with remote sandbox execution targets. Related search performed: - Public issues: `Claude sandbox probe`, `Codex CODEX_HOME sandbox` returned no matches. - Public PRs: `Claude Codex sandbox probe`, `codex home sandbox`, `claude auth sandbox` returned no matches. ## What Changed - Made Claude sandbox Test probes materialize the same Paperclip-managed Claude config seed path used by sandbox execution. - Preserved sandbox-local Claude credentials when materializing remote Claude config and expanded auth-required detection for `/login` API-key failures. - Improved Claude hello-probe diagnostics so the final result/error is surfaced instead of the unhelpful stream init event, with transient upstream failures downgraded to warnings. - Changed Codex probe behavior to upload only minimal auth/config files instead of the full managed `CODEX_HOME`. - Let Codex sandbox probes leave `CODEX_HOME` unset when the host has no credentials, so pre-authenticated sandbox images can be tested directly. - Excluded bulky host-local Codex session/shell state from sandbox runtime home uploads. - Switched the Codex local default model away from the ChatGPT-unsupported `gpt-5.3-codex` option. - Added regression coverage for Claude parsing/probe paths, Codex adapter metadata/argument/probe behavior, and server-level Claude sandbox environment behavior. ## Verification Passed locally: - `pnpm install --frozen-lockfile` - `pnpm vitest run packages/adapters/claude-local/src/server/parse.test.ts packages/adapters/claude-local/src/server/test.probe.test.ts server/src/__tests__/claude-local-adapter-environment.test.ts` - `pnpm vitest run packages/adapters/codex-local/src/index.test.ts packages/adapters/codex-local/src/server/codex-args.test.ts packages/adapters/codex-local/src/server/test.remote.test.ts` - `pnpm --filter @paperclipai/adapter-claude-local typecheck` - `pnpm --filter @paperclipai/adapter-codex-local typecheck` - `pnpm --filter @paperclipai/server typecheck` - `git diff --check` ## Risks - Adapter configuration behavior is sensitive to local vs sandboxed execution mode, so review should focus on environment detection, argument construction, and any state written during probe/test runs. - The Codex default-model change may affect newly created agents that rely on the adapter default instead of an explicit model. - Excluding Codex session/shell state from sandbox uploads should be safe for fresh sandbox runs, but reviewers should confirm no runtime resume path depends on that host-local state. - Provider-specific setup/capture behavior is intentionally left to separate work. ## Model Used OpenAI GPT-5 Codex via Paperclip `codex_local`; tool-enabled local coding session with terminal access. Context window size was not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
266 lines
10 KiB
TypeScript
266 lines
10 KiB
TypeScript
import fs from "node:fs/promises";
|
|
import os from "node:os";
|
|
import { afterEach, describe, expect, it, vi } from "vitest";
|
|
import type { AdapterExecutionTarget } from "@paperclipai/adapter-utils/execution-target";
|
|
|
|
const {
|
|
ensureAdapterExecutionTargetDirectory,
|
|
ensureAdapterExecutionTargetCommandResolvable,
|
|
maybeRunSandboxInstallCommand,
|
|
runAdapterExecutionTargetProcess,
|
|
describeAdapterExecutionTarget,
|
|
resolveAdapterExecutionTargetCwd,
|
|
prepareAdapterExecutionTargetRuntime,
|
|
prepareManagedCodexHome,
|
|
restoreWorkspace,
|
|
capturedHomeAssetFiles,
|
|
} = vi.hoisted(() => {
|
|
const restoreWorkspace = vi.fn(async () => {});
|
|
// Records the files staged in the uploaded "home" asset at call time, before
|
|
// the probe's cleanup deletes the temp dir. Lets tests assert the upload is a
|
|
// minimal credentials-only home and not the full managed CODEX_HOME.
|
|
const capturedHomeAssetFiles: { value: string[] | null } = { value: null };
|
|
return {
|
|
capturedHomeAssetFiles,
|
|
ensureAdapterExecutionTargetDirectory: vi.fn(async () => {}),
|
|
ensureAdapterExecutionTargetCommandResolvable: vi.fn(async () => {}),
|
|
maybeRunSandboxInstallCommand: vi.fn(async () => null),
|
|
runAdapterExecutionTargetProcess: vi.fn(async () => ({
|
|
exitCode: 0,
|
|
signal: null,
|
|
timedOut: false,
|
|
stdout: [
|
|
"{\"type\":\"thread.started\",\"thread_id\":\"thread-1\"}",
|
|
"{\"type\":\"item.completed\",\"item\":{\"type\":\"agent_message\",\"text\":\"hello\"}}",
|
|
"{\"type\":\"turn.completed\",\"usage\":{\"input_tokens\":1,\"cached_input_tokens\":0,\"output_tokens\":1}}",
|
|
].join("\n"),
|
|
stderr: "",
|
|
pid: 123,
|
|
startedAt: new Date().toISOString(),
|
|
})),
|
|
describeAdapterExecutionTarget: vi.fn(() => "QA SSH"),
|
|
resolveAdapterExecutionTargetCwd: vi.fn((target, configuredCwd, fallbackCwd) => {
|
|
if (typeof configuredCwd === "string" && configuredCwd.trim().length > 0) return configuredCwd;
|
|
if (target && typeof target === "object" && "remoteCwd" in target && typeof target.remoteCwd === "string") {
|
|
return target.remoteCwd;
|
|
}
|
|
return fallbackCwd;
|
|
}),
|
|
prepareAdapterExecutionTargetRuntime: vi.fn(async (input: { assets?: Array<{ key: string; localDir: string }> }) => {
|
|
const homeAsset = input?.assets?.find((asset) => asset.key === "home");
|
|
if (homeAsset) {
|
|
capturedHomeAssetFiles.value = (await fs.readdir(homeAsset.localDir)).sort();
|
|
}
|
|
return {
|
|
target: null,
|
|
workspaceRemoteDir: "/remote/workspace/.paperclip-runtime/runs/test/workspace",
|
|
runtimeRootDir: "/remote/workspace/.paperclip-runtime/runs/test/workspace/.paperclip-runtime/codex",
|
|
assetDirs: {
|
|
home: "/remote/workspace/.paperclip-runtime/runs/test/workspace/.paperclip-runtime/codex/home",
|
|
},
|
|
restoreWorkspace,
|
|
};
|
|
}),
|
|
prepareManagedCodexHome: vi.fn(async () => {
|
|
// Return a real managed home seeded with credentials so the probe's
|
|
// minimal-home copy step (auth.json/config.toml) has something to read.
|
|
const dir = await fs.mkdtemp(`${os.tmpdir()}/paperclip-managed-codex-home-`);
|
|
await fs.writeFile(`${dir}/auth.json`, JSON.stringify({ OPENAI_API_KEY: "sk-managed" }));
|
|
await fs.writeFile(`${dir}/config.toml`, "model = \"gpt-5\"\n");
|
|
return dir;
|
|
}),
|
|
restoreWorkspace,
|
|
};
|
|
});
|
|
|
|
vi.mock("@paperclipai/adapter-utils/execution-target", async () => {
|
|
const actual = await vi.importActual<typeof import("@paperclipai/adapter-utils/execution-target")>(
|
|
"@paperclipai/adapter-utils/execution-target",
|
|
);
|
|
return {
|
|
...actual,
|
|
ensureAdapterExecutionTargetDirectory,
|
|
ensureAdapterExecutionTargetCommandResolvable,
|
|
maybeRunSandboxInstallCommand,
|
|
runAdapterExecutionTargetProcess,
|
|
describeAdapterExecutionTarget,
|
|
resolveAdapterExecutionTargetCwd,
|
|
prepareAdapterExecutionTargetRuntime,
|
|
};
|
|
});
|
|
|
|
vi.mock("./codex-home.js", async () => {
|
|
const actual = await vi.importActual<typeof import("./codex-home.js")>("./codex-home.js");
|
|
return {
|
|
...actual,
|
|
prepareManagedCodexHome,
|
|
};
|
|
});
|
|
|
|
import { testEnvironment } from "./test.js";
|
|
|
|
describe("codex remote environment diagnostics", () => {
|
|
afterEach(() => {
|
|
vi.clearAllMocks();
|
|
delete process.env.OPENAI_API_KEY;
|
|
});
|
|
|
|
it("stages managed CODEX_HOME in an isolated runtime dir and keeps the probe cwd on the original remote workspace", async () => {
|
|
const remoteTarget: AdapterExecutionTarget = {
|
|
kind: "remote",
|
|
transport: "ssh",
|
|
remoteCwd: "/remote/workspace",
|
|
spec: {
|
|
host: "127.0.0.1",
|
|
port: 22,
|
|
username: "agent",
|
|
privateKey: "PRIVATE KEY",
|
|
knownHosts: "KNOWN HOSTS",
|
|
remoteCwd: "/remote/workspace",
|
|
remoteWorkspacePath: "/remote/workspace",
|
|
strictHostKeyChecking: false,
|
|
},
|
|
};
|
|
|
|
const result = await testEnvironment({
|
|
companyId: "company-1",
|
|
adapterType: "codex_local",
|
|
config: {
|
|
command: "codex",
|
|
},
|
|
executionTarget: remoteTarget,
|
|
environmentName: "QA SSH",
|
|
});
|
|
|
|
expect(result.status).toBe("pass");
|
|
expect(result.checks.some((check) => check.code === "codex_hello_probe_passed")).toBe(true);
|
|
expect(prepareManagedCodexHome).toHaveBeenCalledTimes(1);
|
|
expect(prepareAdapterExecutionTargetRuntime).toHaveBeenCalledTimes(1);
|
|
const runtimeCalls = prepareAdapterExecutionTargetRuntime.mock.calls as unknown as Array<[
|
|
{
|
|
workspaceLocalDir: string;
|
|
target?: { remoteCwd?: string };
|
|
workspaceRemoteDir?: string;
|
|
assets?: Array<{ key: string; localDir: string }>;
|
|
},
|
|
]>;
|
|
const runtimeInput = runtimeCalls[0]?.[0];
|
|
// The probe must upload only a minimal credentials-only home, never the
|
|
// full managed CODEX_HOME (which can be hundreds of MB of session history).
|
|
const homeAsset = runtimeInput?.assets?.find((asset) => asset.key === "home");
|
|
expect(homeAsset?.localDir).toContain(`${os.tmpdir()}/paperclip-codex-probe-home-`);
|
|
expect(capturedHomeAssetFiles.value).toEqual(["auth.json", "config.toml"]);
|
|
expect(runtimeInput?.workspaceLocalDir).toContain(`${os.tmpdir()}/paperclip-codex-envtest-`);
|
|
expect(runtimeInput?.workspaceLocalDir).not.toBe("/remote/workspace");
|
|
expect(await fs.stat(runtimeInput!.workspaceLocalDir).catch(() => null)).toBeNull();
|
|
expect(runtimeInput?.target?.remoteCwd).toBe("/remote/workspace");
|
|
// `workspaceRemoteDir` is the base path passed to the runtime; the
|
|
// helper's per-run subdirectory is appended internally inside
|
|
// `prepareRemoteManagedRuntime`. Pre-building a per-run prefix here
|
|
// would double-nest the run id in the final path.
|
|
expect(runtimeInput?.workspaceRemoteDir).toBe("/remote/workspace");
|
|
expect(runAdapterExecutionTargetProcess).toHaveBeenCalledTimes(1);
|
|
const probeCall = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as
|
|
| [string, { kind: string; remoteCwd: string }, string, string[], { cwd: string; env: Record<string, string> }]
|
|
| undefined;
|
|
expect(probeCall?.[1]).toMatchObject({
|
|
kind: "remote",
|
|
remoteCwd: "/remote/workspace",
|
|
});
|
|
expect(probeCall?.[4]).toMatchObject({
|
|
cwd: "/remote/workspace",
|
|
env: expect.objectContaining({
|
|
CODEX_HOME: "/remote/workspace/.paperclip-runtime/runs/test/workspace/.paperclip-runtime/codex/home",
|
|
}),
|
|
});
|
|
expect(restoreWorkspace).toHaveBeenCalledTimes(1);
|
|
});
|
|
|
|
it("avoids /tmp CODEX_HOME for remote API-key hello probes", async () => {
|
|
const remoteTarget: AdapterExecutionTarget = {
|
|
kind: "remote",
|
|
transport: "sandbox",
|
|
providerKey: "cloudflare",
|
|
remoteCwd: "/remote/workspace",
|
|
runner: {
|
|
execute: async () => ({
|
|
exitCode: 0,
|
|
signal: null,
|
|
timedOut: false,
|
|
stdout: "",
|
|
stderr: "",
|
|
pid: null,
|
|
startedAt: new Date().toISOString(),
|
|
}),
|
|
},
|
|
};
|
|
|
|
const result = await testEnvironment({
|
|
companyId: "company-1",
|
|
adapterType: "codex_local",
|
|
config: {
|
|
command: "codex",
|
|
env: {
|
|
OPENAI_API_KEY: "sk-test",
|
|
},
|
|
},
|
|
executionTarget: remoteTarget,
|
|
environmentName: "QA Cloudflare",
|
|
});
|
|
|
|
expect(result.status).toBe("pass");
|
|
const probeCall = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as
|
|
| [string, AdapterExecutionTarget, string, string[], { cwd: string; env: Record<string, string> }]
|
|
| undefined;
|
|
expect(probeCall?.[4].env.CODEX_HOME).toContain("/remote/workspace/.paperclip-runtime/codex/probe-home-codex-envtest-");
|
|
expect(probeCall?.[4].env.CODEX_HOME?.startsWith("/tmp/")).toBe(false);
|
|
expect(probeCall?.[3]).toContain("--skip-git-repo-check");
|
|
});
|
|
|
|
it("does not override CODEX_HOME when the host has no credentials to seed", async () => {
|
|
// Pre-authenticated sandbox flow: the login lives inside the sandbox image,
|
|
// and the host has no Codex auth.json. The probe must not upload an empty
|
|
// home or set CODEX_HOME, so Codex falls back to the sandbox's baked-in login.
|
|
prepareManagedCodexHome.mockImplementationOnce(async () => {
|
|
const dir = await fs.mkdtemp(`${os.tmpdir()}/paperclip-managed-codex-home-noauth-`);
|
|
// No auth.json — only a config file.
|
|
await fs.writeFile(`${dir}/config.toml`, "model = \"gpt-5\"\n");
|
|
return dir;
|
|
});
|
|
|
|
const remoteTarget: AdapterExecutionTarget = {
|
|
kind: "remote",
|
|
transport: "sandbox",
|
|
providerKey: "daytona",
|
|
remoteCwd: "/remote/workspace",
|
|
runner: {
|
|
execute: async () => ({
|
|
exitCode: 0,
|
|
signal: null,
|
|
timedOut: false,
|
|
stdout: "",
|
|
stderr: "",
|
|
pid: null,
|
|
startedAt: new Date().toISOString(),
|
|
}),
|
|
},
|
|
};
|
|
|
|
const result = await testEnvironment({
|
|
companyId: "company-1",
|
|
adapterType: "codex_local",
|
|
config: { command: "codex" },
|
|
executionTarget: remoteTarget,
|
|
environmentName: "QA Daytona",
|
|
});
|
|
|
|
expect(result.status).toBe("pass");
|
|
// No managed-home upload, so the full-runtime staging is skipped entirely.
|
|
expect(prepareAdapterExecutionTargetRuntime).not.toHaveBeenCalled();
|
|
const probeCall = runAdapterExecutionTargetProcess.mock.calls[0] as unknown as
|
|
| [string, AdapterExecutionTarget, string, string[], { cwd: string; env: Record<string, string> }]
|
|
| undefined;
|
|
expect(probeCall?.[4].env.CODEX_HOME).toBeUndefined();
|
|
});
|
|
});
|