mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-07 16:11:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
217 lines
7.9 KiB
TypeScript
217 lines
7.9 KiB
TypeScript
import { randomUUID } from "node:crypto";
|
||
import { test, expect, type Route } from "@playwright/test";
|
||
|
||
const fulfill = (route: Route, body: unknown, status = 200) =>
|
||
route.fulfill({
|
||
status,
|
||
contentType: "application/json",
|
||
body: JSON.stringify(body),
|
||
});
|
||
|
||
test("AgentMail setup and email work through the normal task conversation", async ({
|
||
page,
|
||
request,
|
||
}) => {
|
||
const created = await request.post("/api/companies", {
|
||
data: { name: `AgentMail browser ${Date.now()}` },
|
||
});
|
||
expect(created.ok()).toBeTruthy();
|
||
const company = await created.json();
|
||
const agentResponse = await request.post(
|
||
`/api/companies/${company.id}/agents`,
|
||
{
|
||
data: {
|
||
name: "Mail agent",
|
||
role: "qa",
|
||
adapterType: "process",
|
||
adapterConfig: {
|
||
command: process.execPath,
|
||
args: ["-e", "process.exit(0)"],
|
||
},
|
||
},
|
||
},
|
||
);
|
||
expect(agentResponse.ok()).toBeTruthy();
|
||
const agent = await agentResponse.json();
|
||
const taskResponse = await request.post(
|
||
`/api/companies/${company.id}/issues`,
|
||
{ data: { title: "Customer email", status: "backlog" } },
|
||
);
|
||
expect(taskResponse.ok()).toBeTruthy();
|
||
const task = await taskResponse.json();
|
||
const inbox = {
|
||
id: randomUUID(),
|
||
companyId: company.id,
|
||
connectionId: randomUUID(),
|
||
assignedAgentId: agent.id,
|
||
address: "agent@agentmail.to",
|
||
status: "active",
|
||
receiveMode: "websocket",
|
||
lastError: null,
|
||
lastSyncAt: new Date().toISOString(),
|
||
};
|
||
let connected = false;
|
||
const sends: any[] = [];
|
||
const conversationId = randomUUID();
|
||
const thread = {
|
||
conversationId,
|
||
issueId: task.id,
|
||
endpoint: inbox,
|
||
subject: "Customer email",
|
||
messages: [
|
||
{
|
||
id: randomUUID(),
|
||
providerMessageId: "incoming-message",
|
||
from: "Customer <customer@example.test>",
|
||
to: [inbox.address],
|
||
cc: ["visible@example.test"],
|
||
bcc: ["private@example.test"],
|
||
subject: "Customer email",
|
||
direction: "inbound",
|
||
text: "Can you help?",
|
||
fullText: "Can you help?\nEarlier quoted context",
|
||
commentId: null,
|
||
attachmentIds: [],
|
||
timestamp: new Date().toISOString(),
|
||
automatic: false,
|
||
},
|
||
],
|
||
publications: [] as any[],
|
||
};
|
||
await page.route("**/api/instance/settings/experimental", (route) =>
|
||
fulfill(route, { enableChatConnectors: true }),
|
||
);
|
||
await page.route("**/api/**/email/**", async (route) => {
|
||
const url = new URL(route.request().url()),
|
||
method = route.request().method();
|
||
if (url.pathname.endsWith("/inspect"))
|
||
return fulfill(route, {
|
||
scope: { scope_type: "organization" },
|
||
inboxes: [{ inbox_id: inbox.address }],
|
||
domains: [
|
||
{
|
||
domain_id: "domain-id",
|
||
domain: "verified.example.test",
|
||
status: "VERIFIED",
|
||
},
|
||
],
|
||
});
|
||
if (url.pathname.endsWith("/inboxes") && method === "GET")
|
||
return fulfill(route, connected ? [inbox] : []);
|
||
if (url.pathname.endsWith("/inboxes") && method === "POST") {
|
||
const body = route.request().postDataJSON();
|
||
expect(body.receiveMode).toBe("websocket");
|
||
expect(body.assignedAgentId).toBe(agent.id);
|
||
connected = true;
|
||
return fulfill(route, inbox, 201);
|
||
}
|
||
if (url.pathname.endsWith(`/tasks/${task.id}`))
|
||
return fulfill(route, thread);
|
||
if (url.pathname.endsWith("/send")) {
|
||
const input = route.request().postDataJSON();
|
||
sends.push(input);
|
||
const publication = {
|
||
id: input.idempotencyKey,
|
||
issueId: input.parentIssueId ? randomUUID() : task.id,
|
||
conversationId,
|
||
outcome: "queued",
|
||
error: null,
|
||
providerMessageId: null,
|
||
};
|
||
thread.publications.push(publication);
|
||
return fulfill(route, publication, 202);
|
||
}
|
||
return fulfill(route, null);
|
||
});
|
||
await page.route(`**/api/chat-endpoints/${inbox.id}`, (route) =>
|
||
fulfill(route, {
|
||
...inbox,
|
||
provider: "agentmail",
|
||
setup: { step: "complete" },
|
||
capabilities: {},
|
||
botExternalId: inbox.address,
|
||
}),
|
||
);
|
||
await page.goto(
|
||
`/${company.issuePrefix}/apps/chat/connect?provider=agentmail&connectionId=${inbox.connectionId}`,
|
||
);
|
||
await expect(
|
||
page.getByRole("heading", { name: "Give an agent an email address" }),
|
||
).toBeVisible();
|
||
await page.getByRole("combobox").click();
|
||
await page.getByPlaceholder("Search all agents…").fill("Mail agent");
|
||
await page.getByRole("option", { name: "Mail agent" }).click();
|
||
await expect(
|
||
page.getByText("Mail agent is not a low-trust agent"),
|
||
).toBeVisible();
|
||
await page.getByRole("button", { name: "Configure low trust" }).click();
|
||
const trustDialog = page.getByRole("dialog");
|
||
await trustDialog
|
||
.getByRole("combobox")
|
||
.first()
|
||
.selectOption("low_trust_review");
|
||
await trustDialog.getByRole("combobox").nth(1).selectOption("root_issue");
|
||
await trustDialog.getByRole("combobox").nth(2).selectOption(task.id);
|
||
await trustDialog
|
||
.getByRole("button", { name: "Save trust settings" })
|
||
.click();
|
||
await expect(page.getByText("Low-trust review configured")).toBeVisible();
|
||
await page.getByRole("button", { name: "Review trust settings" }).click();
|
||
await page.getByRole("dialog").getByRole("combobox").first().selectOption("standard");
|
||
await page.getByRole("button", { name: "Save trust settings" }).click();
|
||
await expect(page.getByText("Mail agent is not a low-trust agent")).toBeVisible();
|
||
const savedAgent = await (await request.get(`/api/agents/${agent.id}`)).json();
|
||
expect(savedAgent.permissions.authorizationPolicy).toEqual({});
|
||
await page.getByRole("button", { name: "Continue", exact: true }).click();
|
||
await page.getByText("Advanced options", { exact: true }).click();
|
||
await expect(
|
||
page
|
||
.getByLabel("Domain", { exact: true })
|
||
.locator("option", { hasText: "verified.example.test" }),
|
||
).toHaveCount(1);
|
||
await page.getByRole("radio", { name: "Use an existing inbox" }).click();
|
||
await page.getByLabel("Available inbox").selectOption(inbox.address);
|
||
await page.getByRole("button", { name: "Review email address" }).click();
|
||
await expect(
|
||
page.getByText("Anyone can email an unrestricted inbox"),
|
||
).toBeVisible();
|
||
await expect(
|
||
page.getByRole("link", { name: "Set up allowlists ↗" }),
|
||
).toHaveAttribute(
|
||
"href",
|
||
"https://docs.agentmail.to/knowledge-base/allowlists-blocklists",
|
||
);
|
||
await page
|
||
.getByRole("button", { name: "Connect email address", exact: true })
|
||
.click();
|
||
await expect(
|
||
page.getByRole("heading", { name: "Your agent’s email is ready" }),
|
||
).toBeVisible();
|
||
await page.goto(`/${company.issuePrefix}/issues/${task.identifier}`);
|
||
const email = page.getByRole("article", { name: "Email received", exact: true });
|
||
await expect(email).toBeVisible();
|
||
await expect(email.getByText("Can you help?", { exact: true })).toBeVisible();
|
||
await expect(page.getByRole("button", {
|
||
name: /^(Internal comment|Email reply|Start email child task)$/,
|
||
})).toHaveCount(0);
|
||
await expect(email.getByText("Bcc: private@example.test")).not.toBeVisible();
|
||
await email.getByText("Email details", { exact: true }).click();
|
||
await expect(email.getByText("Bcc: private@example.test")).toBeVisible();
|
||
|
||
const composer = page.locator('[contenteditable="true"]').last();
|
||
await expect(composer).toBeEditable();
|
||
const instruction = "Please reply to the customer and confirm Friday delivery.";
|
||
await composer.fill(instruction);
|
||
await page.getByRole("button", { name: "Send", exact: true }).click();
|
||
await expect.poll(async () => {
|
||
const comments = await (await request.get(`/api/issues/${task.id}/comments`)).json();
|
||
return comments.some((comment: { body: string }) => comment.body.includes(instruction));
|
||
}).toBe(true);
|
||
// Task instructions persist normally; only an explicit agent action sends mail.
|
||
expect(sends).toHaveLength(0);
|
||
await page.screenshot({
|
||
path: test.info().outputPath("email-task-conversation.png"),
|
||
fullPage: true,
|
||
});
|
||
});
|