Files
PaperClipAI/packages/adapters/grok-local/src/server/test.ts
T
Devin Foley 5054c9ef9b fix(grok): stage the environment test from a host directory, and survive an absent workspace (#13416)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent runs through an adapter. Before you use an environment, the
adapter's environment test probes it and reports structured checks
> - The grok adapter's environment test stages the managed account into
a remote environment first. It gave the runtime its resolved `cwd` as
the host workspace directory
> - For a remote target that `cwd` is a path inside the sandbox. It does
not exist on the host that runs the test
> - A new workspace ignore scan reads that host directory. The scan
failed on the missing directory, and the error escaped the environment
test
> - The user got a server error instead of a result with checks
> - This pull request stages from a host temporary directory, sends the
remote path separately, and reports a staging failure as a check
> - The benefit is that the grok environment test always answers with
checks, and a workspace directory that does not exist no longer stops a
runtime preparation

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
The grok environment test failed with `Error: Workspace ignore scan
failed: git-ignore-scan-failed`. The route returned a server error, so
the user saw no checks at all.

**Expected behavior**
The environment test always returns `{status, checks}`. Every other
problem it finds (an invalid working directory, a command it cannot
resolve, a probe that times out) becomes a check with a level. A
credential staging problem must do the same.

**Steps to reproduce**
1. Configure a grok agent with a managed AI connection.
2. Point the agent at a remote (sandbox) environment.
3. Run the environment test for that environment.

**Paperclip version or commit**
Present on master. Both halves landed on 2026-09-12: the call site in
#13247, and the scan that it trips in #13353.

## What Changed

- `packages/adapters/grok-local/src/server/test.ts` stages the managed
account through a fresh host temporary directory and passes the remote
path as `workspaceRemoteDir`. This is the split `codex-local` and
`opencode-local` already use.
- A failure while staging becomes a `grok_environment_unprepared` check
with level `error`, instead of an exception that escapes the function.
The probe does not run after it, because there is no prepared
environment to probe.
- The temporary directory is removed in the existing `finally` block.
- `packages/adapter-utils/src/sandbox-managed-runtime.ts` treats a
`workspaceLocalDir` that does not exist as "nothing to sync". A
directory that is not there has no files for ignore rules to govern,
nothing to stage, and nothing to restore.
- Tests: the grok environment test now covers the managed-connection
remote branch, which had no coverage. `sandbox-file-sync.test.ts` covers
a preparation whose workspace directory does not exist.

## Verification

```
npx vitest run packages/adapters/grok-local/src/server/test.test.ts     # 9 passed
npx vitest run packages/adapter-utils/src/sandbox-file-sync.test.ts \
              packages/adapter-utils/src/sandbox-managed-runtime.test.ts # 90 passed
cd packages/adapter-utils && npx tsc --noEmit                            # clean
cd packages/adapters/grok-local && npx tsc --noEmit                      # clean
```

The two new grok tests fail against the old call site: the first asserts
the runtime never receives the remote path as its host workspace
directory, and the second asserts a staging failure becomes a check
instead of an exception.

## Risks

Low risk, and limited to environment preparation.

- The adapter change only affects the managed-connection remote branch
of one adapter's environment test.
- The `adapter-utils` change makes a preparation that used to throw now
continue with no workspace sync. A real workspace is unaffected, because
the directory exists in that case and the scan runs exactly as before.
- No migration. No API change.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-14 09:34:50 -07:00

407 lines
15 KiB
TypeScript

import type {
AdapterEnvironmentCheck,
AdapterEnvironmentTestContext,
AdapterEnvironmentTestResult,
} from "@paperclipai/adapter-utils";
import {
asNumber,
asString,
asStringArray,
ensurePathInEnv,
parseObject,
} from "@paperclipai/adapter-utils/server-utils";
import {
describeAdapterExecutionTarget,
prepareAdapterExecutionTargetRuntime,
ensureAdapterExecutionTargetCommandResolvable,
ensureAdapterExecutionTargetDirectory,
resolveAdapterExecutionTargetCwd,
runAdapterExecutionTargetProcess,
} from "@paperclipai/adapter-utils/execution-target";
import { mkdtemp, rm } from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { stageGrokHomeForSync } from "./grok-home.js";
import { copyBackGrokAuth } from "./grok-auth-copyback.js";
import { DEFAULT_GROK_LOCAL_MODEL } from "../index.js";
import { parseGrokJsonl } from "./parse.js";
import { ADAPTER_AUTH_MISSING_CHECK_CODE } from "@paperclipai/shared";
export interface GrokModelsProbe {
authenticated: boolean;
defaultModel: string | null;
models: string[];
}
function summarizeStatus(checks: AdapterEnvironmentCheck[]): AdapterEnvironmentTestResult["status"] {
if (checks.some((check) => check.level === "error")) return "fail";
if (checks.some((check) => check.level === "warn")) return "warn";
return "pass";
}
function firstNonEmptyLine(text: string): string {
return (
text
.split(/\r?\n/)
.map((line) => line.trim())
.find(Boolean) ?? ""
);
}
function summarizeProbeDetail(stdout: string, stderr: string, parsedError: string | null): string | null {
const raw = parsedError?.trim() || firstNonEmptyLine(stderr) || firstNonEmptyLine(stdout);
if (!raw) return null;
const clean = raw.replace(/\s+/g, " ").trim();
const max = 240;
return clean.length > max ? `${clean.slice(0, max - 3)}...` : clean;
}
function normalizeEnv(input: unknown): Record<string, string> {
if (typeof input !== "object" || input === null || Array.isArray(input)) return {};
const env: Record<string, string> = {};
for (const [key, value] of Object.entries(input as Record<string, unknown>)) {
if (typeof value === "string") env[key] = value;
}
return env;
}
const GROK_AUTH_REQUIRED_RE =
/(?:not\s+logged\s+in|login\s+required|run\s+`?grok\s+login`?|authentication\s+required|unauthorized|invalid\s+credentials)/i;
export function parseGrokModelsOutput(stdout: string): GrokModelsProbe {
const trimmedLines = stdout
.split(/\r?\n/)
.map((line) => line.trim())
.filter(Boolean);
const models: string[] = [];
let defaultModel: string | null = null;
let authenticated = false;
let inModelsBlock = false;
for (const line of trimmedLines) {
if (/logged in/i.test(line)) authenticated = true;
const defaultMatch = /^Default model:\s*(.+)$/i.exec(line);
if (defaultMatch?.[1]) {
defaultModel = defaultMatch[1].trim();
continue;
}
if (/^Available models:/i.test(line)) {
inModelsBlock = true;
continue;
}
if (!inModelsBlock) continue;
const bulletMatch = /^[*-]\s*(.+?)(?:\s+\(default\))?$/.exec(line);
if (bulletMatch?.[1]) {
models.push(bulletMatch[1].trim());
continue;
}
if (line.length > 0) {
models.push(line.replace(/\s+\(default\)$/, "").trim());
}
}
return {
authenticated,
defaultModel,
models: Array.from(new Set(models.filter(Boolean))),
};
}
export async function testEnvironment(
ctx: AdapterEnvironmentTestContext,
): Promise<AdapterEnvironmentTestResult> {
const checks: AdapterEnvironmentCheck[] = [];
const config = parseObject(ctx.config);
const command = asString(config.command, "grok");
const target = ctx.executionTarget ?? null;
const targetIsRemote = target?.kind === "remote";
const targetIsSandbox = target?.kind === "remote" && target.transport === "sandbox";
const cwd = resolveAdapterExecutionTargetCwd(target, asString(config.cwd, ""), process.cwd());
const targetLabel = targetIsRemote
? ctx.environmentName ?? describeAdapterExecutionTarget(target)
: null;
const runId = `grok-envtest-${Date.now()}-${Math.random().toString(16).slice(2)}`;
if (targetLabel) {
checks.push({
code: "grok_environment_target",
level: "info",
message: `Probing inside environment: ${targetLabel}`,
});
}
try {
await ensureAdapterExecutionTargetDirectory(runId, target, cwd, {
cwd,
env: {},
createIfMissing: true,
});
checks.push({
code: "grok_cwd_valid",
level: "info",
message: `Working directory is valid: ${cwd}`,
});
} catch (err) {
checks.push({
code: "grok_cwd_invalid",
level: "error",
message: err instanceof Error ? err.message : "Invalid working directory",
detail: cwd,
});
}
const env = normalizeEnv(config.env);
let stagedHome: string | undefined;
let runtimeWorkspaceLocalDir: string | undefined;
let restore: (() => Promise<void>) | undefined;
try {
if (config.managedAiConnection && targetIsRemote) {
try {
const hostHome = env.GROK_HOME;
stagedHome = await stageGrokHomeForSync(hostHome, { runId });
// `cwd` is the remote target path here, never a host directory, so the
// runtime gets an empty host workspace to stage from and the remote
// path separately — the same split codex-local and opencode-local
// draw. Handing it the remote path as `workspaceLocalDir` made the
// host-side ignore scan fail on a directory that does not exist.
runtimeWorkspaceLocalDir = await mkdtemp(
path.join(os.tmpdir(), `paperclip-grok-envtest-${runId}-`),
);
const prepared = await prepareAdapterExecutionTargetRuntime({
runId, target, adapterKey: "grok",
workspaceLocalDir: runtimeWorkspaceLocalDir,
workspaceRemoteDir: cwd,
assets: [{ key: "home", localDir: stagedHome, followSymlinks: true,
restore: async ({ assetDir, readFile }) => { await copyBackGrokAuth({
readSandboxAuth: () => readFile(path.posix.join(assetDir, "auth.json")),
hostHomeDir: hostHome, log: () => {},
}); },
}],
});
env.GROK_HOME = prepared.assetDirs.home;
restore = () => prepared.restoreWorkspace(() => {});
} catch (err) {
// The environment test reports what is wrong with the environment; a
// credential-staging failure is a finding, not a crash.
checks.push({
code: "grok_environment_unprepared",
level: "error",
message: err instanceof Error ? err.message : "Could not stage the managed account into the environment",
});
}
}
const runtimeEnv = ensurePathInEnv({ ...process.env, ...env });
try {
await ensureAdapterExecutionTargetCommandResolvable(command, target, cwd, runtimeEnv);
checks.push({
code: "grok_command_resolvable",
level: "info",
message: `Command is executable: ${command}`,
});
} catch (err) {
checks.push({
code: "grok_command_unresolvable",
level: "error",
message: err instanceof Error ? err.message : "Command is not executable",
detail: command,
});
}
const canRunProbe =
checks.every((check) =>
check.code !== "grok_cwd_invalid" &&
check.code !== "grok_command_unresolvable" &&
check.code !== "grok_environment_unprepared");
const configuredModel = asString(config.model, DEFAULT_GROK_LOCAL_MODEL).trim();
if (canRunProbe) {
const modelsProbe = await runAdapterExecutionTargetProcess(
runId,
target,
command,
["models"],
{
cwd,
env,
timeoutSec: Math.max(1, asNumber(config.helloProbeTimeoutSec, 45)),
graceSec: 5,
onLog: async () => {},
},
);
const probeOutput = `${modelsProbe.stdout}\n${modelsProbe.stderr}`;
const parsedModels = parseGrokModelsOutput(modelsProbe.stdout);
const authRequired = GROK_AUTH_REQUIRED_RE.test(probeOutput);
if (modelsProbe.timedOut) {
checks.push({
code: "grok_models_probe_timed_out",
level: "warn",
message: "`grok models` timed out.",
hint: "Retry the probe. If this persists, run `grok models` manually from the target environment.",
});
} else if ((modelsProbe.exitCode ?? 1) !== 0) {
checks.push({
code: authRequired ? "grok_auth_required" : "grok_models_probe_failed",
level: authRequired ? "warn" : "error",
message: authRequired
? "Grok CLI is not authenticated."
: "`grok models` failed.",
detail: summarizeProbeDetail(modelsProbe.stdout, modelsProbe.stderr, null),
hint: authRequired ? "Run `grok login` on the target host, then retry." : undefined,
});
if (authRequired && targetIsSandbox) {
// Emit the neutral canonical check so the user interface can decide
// login eligibility from a stable code. The user interface does not
// read the message text or the top-level status.
checks.push({
code: ADAPTER_AUTH_MISSING_CHECK_CODE,
level: "warn",
message: "This environment has no ready authentication for this adapter.",
detail: summarizeProbeDetail(modelsProbe.stdout, modelsProbe.stderr, null),
hint: "Provide credentials for this adapter, or start login in the environment.",
});
}
} else {
checks.push({
code: "grok_models_probe_passed",
level: "info",
message: parsedModels.authenticated
? "Grok CLI authentication is configured."
: "`grok models` completed.",
detail: parsedModels.defaultModel ? `Default model: ${parsedModels.defaultModel}` : undefined,
});
if (parsedModels.models.length > 0) {
checks.push({
code: "grok_models_discovered",
level: "info",
message: `Discovered ${parsedModels.models.length} Grok model(s).`,
});
} else {
checks.push({
code: "grok_models_empty",
level: "warn",
message: "Grok returned no available models.",
hint: "Run `grok models` manually and verify the account has access to a model.",
});
}
if (configuredModel) {
// The default model id is a sentinel meaning "let the Grok CLI pick its
// own default" — execute.ts only passes `--model` when the value differs
// from it, so the sentinel is never sent to grok and must not be checked
// against the discovered list (real grok never lists "grok-build", which
// otherwise produced a spurious "not found" warning on every probe).
const usesCliDefault = configuredModel === DEFAULT_GROK_LOCAL_MODEL;
const available = usesCliDefault || parsedModels.models.includes(configuredModel);
checks.push({
code: available ? "grok_model_configured" : "grok_model_not_found",
level: available ? "info" : "warn",
message: usesCliDefault
? `Using the Grok CLI's default model${parsedModels.defaultModel ? ` (${parsedModels.defaultModel})` : ""}.`
: available
? `Configured model: ${configuredModel}`
: `Configured model "${configuredModel}" not found in available models.`,
hint: available ? undefined : "Run `grok models` and choose an available model id.",
});
}
}
}
if (canRunProbe) {
const probeArgs = [
"--output-format",
"streaming-json",
"--always-approve",
"--permission-mode",
"dontAsk",
"--disable-web-search",
];
if (configuredModel && configuredModel !== DEFAULT_GROK_LOCAL_MODEL) {
probeArgs.push("--model", configuredModel);
}
probeArgs.push("--single", "Respond with exactly hello.");
const helloProbe = await runAdapterExecutionTargetProcess(
runId,
target,
command,
probeArgs,
{
cwd,
env,
timeoutSec: Math.max(1, asNumber(config.helloProbeTimeoutSec, 45)),
graceSec: 5,
onLog: async () => {},
},
);
const parsed = parseGrokJsonl(helloProbe.stdout);
const detail = summarizeProbeDetail(helloProbe.stdout, helloProbe.stderr, parsed.errorMessage);
const authRequired = GROK_AUTH_REQUIRED_RE.test(`${helloProbe.stdout}\n${helloProbe.stderr}`);
if (helloProbe.timedOut) {
checks.push({
code: "grok_hello_probe_timed_out",
level: "warn",
message: "Grok hello probe timed out.",
hint: "Retry the probe. If this persists, verify Grok can run a simple `--single` prompt manually.",
});
} else if ((helloProbe.exitCode ?? 1) !== 0) {
checks.push({
code: authRequired ? "grok_hello_probe_auth_required" : "grok_hello_probe_failed",
level: authRequired ? "warn" : "error",
message: authRequired
? "Grok CLI could not answer the hello probe because authentication is missing."
: "Grok hello probe failed.",
...(detail ? { detail } : {}),
hint: authRequired ? "Run `grok login` on the target host, then retry." : undefined,
});
if (authRequired && targetIsSandbox) {
// Emit the neutral canonical check so the user interface can decide
// login eligibility from a stable code. The user interface does not
// read the message text or the top-level status.
checks.push({
code: ADAPTER_AUTH_MISSING_CHECK_CODE,
level: "warn",
message: "This environment has no ready authentication for this adapter.",
...(detail ? { detail } : {}),
hint: "Provide credentials for this adapter, or start login in the environment.",
});
}
} else if (/\bhello\b/i.test(parsed.summary)) {
checks.push({
code: "grok_hello_probe_passed",
level: "info",
message: "Grok hello probe succeeded.",
});
} else {
checks.push({
code: "grok_hello_probe_unexpected_output",
level: "warn",
message: "Grok hello probe succeeded but returned unexpected output.",
...(detail ? { detail } : {}),
});
}
}
return {
adapterType: "grok_local",
status: summarizeStatus(checks),
checks,
testedAt: new Date().toISOString(),
};
} finally {
try { await restore?.(); } finally {
// Both temporary directories are removed even when one removal fails,
// and neither failure turns an answered environment test into a thrown
// error: these are best-effort temp directories, and the caller asked
// for checks.
await Promise.allSettled([
stagedHome ? rm(stagedHome, { recursive: true, force: true }) : undefined,
runtimeWorkspaceLocalDir ? rm(runtimeWorkspaceLocalDir, { recursive: true, force: true }) : undefined,
]);
}
}
}