Files
PaperClipAI/server/src/__tests__/adapter-models.test.ts
T
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00

381 lines
17 KiB
TypeScript

import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { models as claudeFallbackModels } from "@paperclipai/adapter-claude-local";
import { resetClaudeModelsCacheForTests } from "@paperclipai/adapter-claude-local/server";
import { models as codexFallbackModels } from "@paperclipai/adapter-codex-local";
import { models as cursorFallbackModels } from "@paperclipai/adapter-cursor-local";
import { models as opencodeFallbackModels } from "@paperclipai/adapter-opencode-local";
import { resetOpenCodeModelsCacheForTests } from "@paperclipai/adapter-opencode-local/server";
import { listAdapterModels, listServerAdapters, refreshAdapterModels, registerServerAdapter, unregisterServerAdapter } from "../adapters/index.js";
import { resetCodexModelsCacheForTests } from "../adapters/codex-models.js";
import { resetCursorModelsCacheForTests, setCursorModelsRunnerForTests } from "../adapters/cursor-models.js";
vi.mock("acpx/runtime", () => ({
createAcpRuntime: vi.fn(),
createAgentRegistry: vi.fn(),
createRuntimeStore: vi.fn(),
isAcpRuntimeError: vi.fn(() => false),
}));
describe("adapter model listing", () => {
beforeEach(() => {
delete process.env.OPENAI_API_KEY;
delete process.env.ANTHROPIC_API_KEY;
delete process.env.ANTHROPIC_BASE_URL;
delete process.env.ANTHROPIC_BEDROCK_BASE_URL;
delete process.env.CLAUDE_CODE_USE_BEDROCK;
delete process.env.PAPERCLIP_OPENCODE_COMMAND;
resetClaudeModelsCacheForTests();
resetCodexModelsCacheForTests();
resetCursorModelsCacheForTests();
setCursorModelsRunnerForTests(null);
resetOpenCodeModelsCacheForTests();
vi.restoreAllMocks();
});
it("returns an empty list for unknown adapters", async () => {
const models = await listAdapterModels("unknown_adapter");
expect(models).toEqual([]);
});
it("does not expose models for the retired acpx_local tombstone", () => {
const adapter = listServerAdapters().find((candidate) => candidate.type === "acpx_local");
expect(adapter?.models).toEqual([]);
});
it("returns codex fallback models when no OpenAI key is available", async () => {
const fetchSpy = vi.spyOn(globalThis, "fetch");
const models = await listAdapterModels("codex_local");
expect(models).toEqual(codexFallbackModels);
// The bare gpt-5.6 alias is intentionally not advertised (Codex has no metadata for it).
expect(models.some((model) => model.id === "gpt-5.6")).toBe(false);
expect(models.some((model) => model.id === "gpt-5.6-sol")).toBe(true);
expect(models.some((model) => model.id === "gpt-5.6-terra")).toBe(true);
expect(models.some((model) => model.id === "gpt-5.6-luna")).toBe(true);
expect(models.some((model) => model.id === "gpt-5.3-codex-spark")).toBe(false);
expect(fetchSpy).not.toHaveBeenCalled();
});
it("returns claude fallback models including the latest Opus alias when no Anthropic key is available", async () => {
const fetchSpy = vi.spyOn(globalThis, "fetch");
const models = await listAdapterModels("claude_local");
expect(models).toEqual(claudeFallbackModels);
expect(models.some((model) => model.id === "claude-opus-4-8")).toBe(true);
// Newer flagship models are offered, but Opus 4.8 stays the default (first) option.
expect(models[0]?.id).toBe("claude-opus-4-8");
expect(models.some((model) => model.id === "claude-sonnet-5")).toBe(true);
expect(models.some((model) => model.id === "claude-fable-5-1")).toBe(true);
expect(models.some((model) => model.id === "claude-fable-5")).toBe(true);
expect(models.some((model) => model.id === "claude-mythos-5")).toBe(true);
// Opus 5 is a current GA flagship and must be offered even when live discovery is unavailable.
expect(models.some((model) => model.id === "claude-opus-5")).toBe(true);
expect(models).toContainEqual({ id: "claude-opus-5-5", label: "Claude Opus 5.5" });
expect(fetchSpy).not.toHaveBeenCalled();
});
it("loads claude models dynamically and merges fallback options", async () => {
process.env.ANTHROPIC_API_KEY = "sk-ant-test";
const fetchSpy = vi.spyOn(globalThis, "fetch").mockResolvedValue({
ok: true,
json: async () => ({
data: [
{ id: "claude-sonnet-4-20250514", display_name: "Claude Sonnet 4" },
{ id: "claude-opus-4-8-20260529", display_name: "Claude Opus 4.8" },
],
}),
} as Response);
const first = await listAdapterModels("claude_local");
const second = await listAdapterModels("claude_local");
expect(fetchSpy).toHaveBeenCalledTimes(1);
expect(first).toEqual(second);
expect(first.some((model) => model.id === "claude-opus-4-8-20260529")).toBe(true);
expect(first.some((model) => model.id === "claude-opus-4-8")).toBe(true);
expect(first.some((model) => model.id === "claude-opus-5-5")).toBe(true);
});
it("refreshes cached claude models on demand", async () => {
process.env.ANTHROPIC_API_KEY = "sk-ant-test";
const fetchSpy = vi.spyOn(globalThis, "fetch")
.mockResolvedValueOnce({
ok: true,
json: async () => ({
data: [{ id: "claude-sonnet-4-20250514", display_name: "Claude Sonnet 4" }],
}),
} as Response)
.mockResolvedValueOnce({
ok: true,
json: async () => ({
data: [{ id: "claude-opus-4-8-20260529", display_name: "Claude Opus 4.8" }],
}),
} as Response);
const initial = await listAdapterModels("claude_local");
const refreshed = await refreshAdapterModels("claude_local");
expect(fetchSpy).toHaveBeenCalledTimes(2);
expect(initial.some((model) => model.id === "claude-sonnet-4-20250514")).toBe(true);
expect(refreshed.some((model) => model.id === "claude-opus-4-8-20260529")).toBe(true);
});
it("falls back to static claude models when Anthropic model discovery fails", async () => {
process.env.ANTHROPIC_API_KEY = "sk-ant-test";
vi.spyOn(globalThis, "fetch").mockResolvedValue({
ok: false,
status: 401,
json: async () => ({}),
} as Response);
const models = await listAdapterModels("claude_local");
expect(models).toEqual(claudeFallbackModels);
});
it.each([
["claude-fable-5-1", "Claude Fable 5.1"],
["claude-opus-5-5", "Claude Opus 5.5"],
])("does not duplicate %s when discovery returns the identical ID", async (id, displayName) => {
process.env.ANTHROPIC_API_KEY = "sk-ant-test";
vi.spyOn(globalThis, "fetch").mockResolvedValue({
ok: true,
json: async () => ({
data: [{ id, display_name: displayName }],
}),
} as Response);
const models = await listAdapterModels("claude_local");
expect(models.filter((model) => model.id === id)).toEqual([{ id, label: displayName }]);
// Curated fallbacks discovery did not return are still merged in.
expect(models.some((model) => model.id === "claude-fable-5")).toBe(true);
expect(models.some((model) => model.id === "claude-opus-4-8")).toBe(true);
});
it("exposes the Bedrock-native Fable 5.1 ID (never the direct ID) in Bedrock mode", async () => {
process.env.CLAUDE_CODE_USE_BEDROCK = "1";
const fetchSpy = vi.spyOn(globalThis, "fetch");
const models = await listAdapterModels("claude_local");
// Keep Opus 4.8 first, using its documented dateless Bedrock ID.
expect(models[0]?.id).toBe("us.anthropic.claude-opus-4-8");
expect(models.map((model) => model.id)).toEqual(expect.arrayContaining([
"us.anthropic.claude-opus-5-5", "us.anthropic.claude-opus-5", "us.anthropic.claude-sonnet-5",
"us.anthropic.claude-fable-5-1", "us.anthropic.claude-opus-4-7", "us.anthropic.claude-sonnet-4-6",
]));
expect(models.map((model) => model.id)).not.toEqual(expect.arrayContaining(["us.anthropic.claude-opus-4-8-v1"]));
expect(models.some((model) => model.id === "claude-fable-5-1")).toBe(false);
expect(fetchSpy).not.toHaveBeenCalled();
});
it.each([
["gemini_local", ["gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash", "gemini-3.5-flash-lite", "gemini-3.1-flash-lite", "gemini-3-flash-preview"]],
["grok_local", ["grok-build", "grok-4.7", "grok-4.6", "grok-4.5"]],
["kimi_local", ["kimi-code/kimi-for-coding", "kimi-code/k3", "kimi-code/k3-256k"]],
])("lists current %s models without a provider login", async (adapter, expectedIds) => {
const models = await listAdapterModels(adapter as string);
expect(models.map((model) => model.id)).toEqual(expect.arrayContaining(expectedIds as string[]));
expect(new Set(models.map((model) => model.id)).size).toBe(models.length);
if (adapter === "gemini_local") {
expect(models.some((model) => model.id.startsWith("gemini-2.0-"))).toBe(false);
}
if (adapter === "kimi_local") {
expect(models).toContainEqual({ id: "kimi-code/kimi-for-coding", label: "K2.8 Preview" });
}
});
it("includes current Cursor fallbacks when runtime discovery is unavailable", async () => {
setCursorModelsRunnerForTests(() => ({ status: 1, stdout: "", stderr: "", hasError: true }));
const models = await listAdapterModels("cursor");
expect(models.map((model) => model.id)).toEqual(expect.arrayContaining([
"composer-2.5", "claude-opus-5-5", "claude-fable-5-1", "claude-sonnet-5",
"gpt-5.6-sol", "gpt-5.6-terra", "gpt-5.6-luna", "grok-4.7", "gemini-3.8-flash", "muse-spark-1.3",
]));
});
it("keeps general OpenAI API models out of the Codex catalog", async () => {
process.env.OPENAI_API_KEY = "sk-test";
const fetchSpy = vi.spyOn(globalThis, "fetch").mockResolvedValue({
ok: true,
json: async () => ({
data: [
{ id: "gpt-image-1" },
{ id: "text-embedding-3-large" },
],
}),
} as Response);
const first = await listAdapterModels("codex_local");
const second = await listAdapterModels("codex_local");
expect(fetchSpy).not.toHaveBeenCalled();
expect(first).toEqual(codexFallbackModels);
expect(second).toEqual(codexFallbackModels);
});
it("keeps the curated Codex list when refreshing with an OpenAI key", async () => {
process.env.OPENAI_API_KEY = "sk-test";
const fetchSpy = vi.spyOn(globalThis, "fetch");
const initial = await listAdapterModels("codex_local");
const refreshed = await refreshAdapterModels("codex_local");
expect(fetchSpy).not.toHaveBeenCalled();
expect(initial).toEqual(codexFallbackModels);
expect(refreshed).toEqual(codexFallbackModels);
});
it("uses static Codex models without calling OpenAI model discovery", async () => {
process.env.OPENAI_API_KEY = "sk-test";
const fetchSpy = vi.spyOn(globalThis, "fetch").mockResolvedValue({
ok: false,
status: 401,
json: async () => ({}),
} as Response);
const models = await listAdapterModels("codex_local");
expect(models).toEqual(codexFallbackModels);
expect(fetchSpy).not.toHaveBeenCalled();
});
it("uses a custom Codex adapter's model and refresh hooks", async () => {
const builtin = listServerAdapters().find((adapter) => adapter.type === "codex_local")!;
const customModels = [{ id: "plugin-codex", label: "Plugin Codex" }];
const listModels = vi.fn(async () => customModels);
const refreshModels = vi.fn(async () => customModels);
registerServerAdapter({ ...builtin, models: [], listModels, refreshModels });
try {
await expect(listAdapterModels("codex_local")).resolves.toEqual(customModels);
await expect(refreshAdapterModels("codex_local")).resolves.toEqual(customModels);
expect(listModels).toHaveBeenCalledOnce();
expect(refreshModels).toHaveBeenCalledOnce();
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
codex_local: [{ id: "declared-codex", label: "Declared Codex" }],
});
const declared = [{ id: "declared-codex", label: "Declared Codex" }];
await expect(listAdapterModels("codex_local")).resolves.toEqual(declared);
await expect(refreshAdapterModels("codex_local")).resolves.toEqual(declared);
expect(listModels).toHaveBeenCalledOnce();
expect(refreshModels).toHaveBeenCalledOnce();
} finally {
delete process.env.PAPERCLIP_ADAPTER_MODELS;
unregisterServerAdapter("codex_local");
}
});
it("returns cursor fallback models when CLI discovery is unavailable", async () => {
setCursorModelsRunnerForTests(() => ({
status: null,
stdout: "",
stderr: "",
hasError: true,
}));
const models = await listAdapterModels("cursor");
expect(models).toEqual(cursorFallbackModels);
});
it("returns current provider-qualified OpenCode models when discovery is unavailable", async () => {
process.env.PAPERCLIP_OPENCODE_COMMAND = "__paperclip_missing_opencode_command__";
const models = await listAdapterModels("opencode_local");
expect(models).toEqual(opencodeFallbackModels);
expect(models.map((model) => model.id)).toEqual(expect.arrayContaining(["openai/gpt-6-astra", "openai/gpt-6-sol", "openai/gpt-6-luna", "openai/gpt-5.6-sol", "openai/gpt-5.6-terra", "openai/gpt-5.6-luna", "anthropic/claude-opus-5-5", "anthropic/claude-opus-5", "anthropic/claude-fable-5-1", "anthropic/claude-sonnet-5", "google/gemini-3.8-flash", "xai/grok-4.7"]));
});
it("loads cursor models dynamically and caches them", async () => {
const runner = vi.fn(() => ({
status: 0,
stdout: "Available models: auto, composer-1.5, gpt-5.3-codex-high, sonnet-4.6",
stderr: "",
hasError: false,
}));
setCursorModelsRunnerForTests(runner);
const first = await listAdapterModels("cursor");
const second = await listAdapterModels("cursor");
expect(runner).toHaveBeenCalledTimes(1);
expect(first).toEqual(second);
expect(first.some((model) => model.id === "auto")).toBe(true);
expect(first.some((model) => model.id === "gpt-5.3-codex-high")).toBe(true);
expect(first.some((model) => model.id === "composer-1")).toBe(true);
});
describe("PAPERCLIP_ADAPTER_MODELS declared models", () => {
afterEach(() => {
delete process.env.PAPERCLIP_ADAPTER_MODELS;
});
it("prefers declared env models over adapter discovery", async () => {
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
opencode_local: [
{ id: "tensorix/deepseek/deepseek-chat-v3.1", label: "DeepSeek v3.1" },
{ id: "tensorix/z-ai/glm-4.7" },
],
});
const models = await listAdapterModels("opencode_local");
expect(models).toEqual([
{ id: "tensorix/deepseek/deepseek-chat-v3.1", label: "DeepSeek v3.1" },
{ id: "tensorix/z-ai/glm-4.7", label: "tensorix/z-ai/glm-4.7" },
]);
});
it("uses declared Codex models for both listing and refresh", async () => {
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
codex_local: [{ id: "private-codex", label: "Private Codex" }],
});
const declared = [{ id: "private-codex", label: "Private Codex" }];
await expect(listAdapterModels("codex_local")).resolves.toEqual(declared);
await expect(refreshAdapterModels("codex_local")).resolves.toEqual(declared);
});
it("observes env changes between calls (memo keyed by raw env value)", async () => {
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
opencode_local: [{ id: "model-a" }],
});
expect(await listAdapterModels("opencode_local")).toEqual([
{ id: "model-a", label: "model-a" },
]);
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
opencode_local: [{ id: "model-b" }],
});
expect(await listAdapterModels("opencode_local")).toEqual([
{ id: "model-b", label: "model-b" },
]);
});
it("fails soft on malformed values: falls back to adapter models instead of throwing", async () => {
process.env.PAPERCLIP_ADAPTER_MODELS = "{not json";
process.env.PAPERCLIP_OPENCODE_COMMAND = "__paperclip_missing_opencode_command__";
const errorSpy = vi.spyOn(console, "error").mockImplementation(() => {});
const models = await listAdapterModels("opencode_local");
expect(models).toEqual(opencodeFallbackModels);
// Parsing is memoized per raw value: a second call must not re-log.
const callsAfterFirst = errorSpy.mock.calls.length;
expect(callsAfterFirst).toBeGreaterThan(0);
await listAdapterModels("opencode_local");
expect(errorSpy.mock.calls.length).toBe(callsAfterFirst);
});
it("ignores declared models for adapters not in the map", async () => {
process.env.PAPERCLIP_ADAPTER_MODELS = JSON.stringify({
opencode_local: [{ id: "model-a" }],
});
const models = await listAdapterModels("codex_local");
expect(models).toEqual(codexFallbackModels);
});
});
});