Files
PaperClipAI/server/src/services/chat-publication-errors.ts
T
DottaandPaperclip 4d317274ce feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Channels connect external conversations to company tasks and agent
execution.
> - Slack, Discord, and AgentMail already provide durable delivery and
access controls.
> - People also need to reach an agent from Apple Messages and send
photos.
> - Photon provides shared Pro DMs, dedicated numbers, and authenticated
event recovery.
> - This pull request connects Photon to the existing channel services.
> - People can message an agent while Paperclip retains task ownership
and approval authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: channel services, shared contracts, database constraints,
Apps, and agent Channels UI.

**Problem or motivation**

Paperclip has no iMessage channel. A person cannot use Apple Messages to
start a task, send a photo, or answer an agent's pending question.

**Proposed solution**

Add experimental **iMessage Photon** with Pro-compatible shared DMs or a
dedicated Photon Cloud number per agent channel. Reuse channel
admission, identity links, task generations, publication, and
interaction continuation. Keep groups disabled for shared allocation.
Dedicated lines support groups that an operator explicitly enables.
Require a fresh linked message and a published agent response before
setup completes.

**Alternatives considered**

Shared allocation has no owned phone number, so it reserves one project
and allows DMs only. Dedicated allocation reserves one stable number.
Local Mac access needs a separate deployment model. The upstream Photon
Chat SDK adapter does not persist the poll mappings and send receipts
required here. This change uses the lower-level SDK without adding
another agent runtime.

**Roadmap alignment**

This extends Connected Apps and agent communication through the existing
channel subsystem. It does not add a parallel tool connection or agent
loop. GitHub searches for Photon and iMessage found no matching provider
implementation.

**Additional context**

This ships behind the existing experimental channel gate. Dedicated-line
release qualification remains incomplete. Real Photon Pro DMs passed
task/reply, native poll, text answers, confirmation rejection, media,
restart, pause, reconnect, revocation, and removal tests. An
operator-supplied iPhone camera HEIC also passed the full round trip.
Dedicated groups remain unqualified. See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the
implementation plan](doc/plans/2026-09-11-imessage-photon.md).

## What Changed

- Add the provider catalog entry, shared setup contracts, and a forward
migration. A global partial index reserves the dedicated number or
shared project until its endpoint is archived.
- Add Cloud project inspection, vaulted project credentials,
selected-line token renewal, and a leased receiver. Persist checkpoint
updates under the receiver lease. Shared project replay accepts sparse
increasing sequences only after a complete recovery barrier.
- Connect DMs and enabled groups to existing task generations, sender
authorization, ordered delivery, and publication services. Keep each
iMessage conversation on its task after completion; only explicit `/new`
or `/close` releases the binding. Publish committed inbound comments
live and label their human bubbles “Sent from iMessage” in both
task-chat renderers.
- Persist immutable text/file send identities, upload receipts, poll
IDs, option IDs, per-person drafts, and canonical interaction
continuation proofs.
- Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG
previews, and related Live Photo companion video retention.
- Add the three-step setup flow and channel management surfaces with
official branding. Preserve the experimental gate and existing
pause/disconnect behavior.
- Add interactive production-component Storybooks for setup, access,
recovery, and ongoing conversations. Add provider, integration, catalog,
and browser regression coverage. Document setup, recovery, supported
boundaries, and qualification gaps.

## Verification

- Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and
receive native Codex replies in Apple Messages. Unlinked senders cannot
start work.
- Three real follow-ups each reopened the same completed task. Incoming
bubbles appeared on its open page without reload and showed “Sent from
iMessage.” The third follow-up ran after restarting the server on
`4d7222110`; the agent correctly repeated its previous reply from before
the restart.
- Native polls after restart, sequential text drafts, required-field
correction, explicit submission, approval rejection with a required
reason, and native continuation passed against Photon.
- PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC
passed in both directions. The camera photo produced a 3024×4032 JPEG
preview. The native agent described it and returned the received HEIC
byte-for-byte.
- Pause/resume, reconnect, identity revocation, removal, `/status`,
`/new`, `/close`, and stale answers after close passed live. Messages
suppressed by pause did not become work on resume. Removal stopped
intake and removed credential bindings.
- All 304 focused tests passed on `4d7222110`. These cover Photon
unit/integration behavior, both task-chat renderers, live comment
hydration, completed-task continuity after restart, enabled groups,
duplicate delivery, and explicit reset/close. The selected Teams
completion-boundary regression also passed. Full workspace
typecheck/build and token gates passed for the conversation fix; the
final UI changes passed their affected typecheck/build and tests.
- All 26 new Photon Storybook Playwright cases passed in light and dark
themes, including the complete shared-DM setup journey and 390px mobile
follow-ups. UI typecheck and the Storybook build passed. These stories
use simulated Photon responses and do not replace the live evidence
above.
- The full chat-adapters browser suite previously passed all 39 cases.
Migration checks passed, and migration 0275 applied to the isolated live
instance with the earlier Photon migration already applied.
- The local full Vitest run was previously interrupted by the host's
embedded-Postgres shared-memory limit; it is not a full-suite pass. All
30 applicable CI checks passed on preceding head `7a5419cac`, with two
skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit
required-story discovery guard to the 26 passing Storybook cases.
Greptile rates this final head 5/5 with no unresolved review threads.
All 30 applicable CI checks passed, with two optional checks skipped.
- A repeated live send key suppressed the duplicate but returned gRPC 6
/ SDK `internalError` without an original receipt. Paperclip keeps
unknown delivery unresolved. This provider behavior is covered by a
regression test.
- See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package
versions, redacted live evidence, deterministic coverage, and remaining
qualification gaps.

## Risks

- Dedicated group qualification remains unrun; groups are disabled for
the approved Pro scope. Real iPhone camera HEIC passed transport,
preview generation, agent inspection, and return. Keep the channel
experimental; the dedicated-line release matrix remains incomplete.
- Shared recovery and attachment aliases were verified against the live
gateway. Duplicate writes currently return an error without the original
receipt; unresolved sends require operator resolution. The
implementation fails visibly on invalid replay ordering, a reset cursor,
or changed identity.
- The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF
binaries have not been executed in this work. Linux musl has no packaged
converter. Unsupported conversion retains the original and reports the
missing preview.
- The migration adds a global reservation across companies for Photon
numbers and shared projects. Paused and revoked endpoints keep that
reservation until removal.
- Integration touches shared channel services. Existing provider browser
coverage passes; broad repository verification is recorded above.
- `pnpm-lock.yaml` is intentionally excluded under repository policy.
The repository bot owns lockfile updates. The additional Superagent
supply-chain scan is neutral/inconclusive because these new dependencies
are not yet in the committed lockfile. Its security scan passed; all
required CI checks pass.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code
execution, browser testing, and tool use. The exact served model
identifier and context-window size are not exposed in this session. No
sub-agents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 15:23:50 -05:00

394 lines
13 KiB
TypeScript

export type ChatPublicationErrorDisposition =
| {
kind: "retry";
retryAfterMs: number;
/** True when the provider explicitly asked Paperclip to slow down. */
providerRateLimit?: boolean;
reason: string;
}
| {
kind: "delivery_unknown";
reason: string;
}
| {
kind: "endpoint_attention";
reason: string;
}
| {
kind: "resource_unavailable";
reason: string;
}
| {
kind: "failed";
reason: string;
};
function finitePositive(value: unknown): number | null {
return typeof value === "number" && Number.isFinite(value) && value > 0
? value
: null;
}
function text(error: unknown): string {
return error instanceof Error ? error.message : String(error);
}
function responseHeader(
headers: Headers | Record<string, unknown> | undefined,
name: string,
): string | null {
if (!headers) return null;
if (headers instanceof Headers) return headers.get(name);
const target = name.toLowerCase();
const entry = Object.entries(headers).find(
([key]) => key.toLowerCase() === target,
);
const value = entry?.[1];
return typeof value === "string" || typeof value === "number"
? String(value)
: null;
}
function positiveNumber(value: string | null): number | null {
if (!value) return null;
const parsed = Number(value);
return Number.isFinite(parsed) && parsed > 0 ? parsed : null;
}
// Node timers accept delays through 2^31-1 milliseconds. Keep the durable
// provider hint intact up to that boundary instead of turning an hour-long
// flood-control response into a rapid series of fifteen-minute retries.
const MAX_PROVIDER_RETRY_AFTER_MS = 2_147_000_000;
function retryAfterHeaderMilliseconds(value: string | null): number | null {
if (!value) return null;
const seconds = positiveNumber(value);
if (seconds !== null) return seconds * 1000;
const absolute = Date.parse(value);
if (!Number.isFinite(absolute)) return null;
return Math.max(1, absolute - Date.now());
}
type StructuredProviderError = {
name?: unknown;
adapter?: unknown;
code?: unknown;
retryAfter?: unknown;
retryAfterMs?: unknown;
retry_after?: unknown;
status?: unknown;
statusCode?: unknown;
subCode?: unknown;
providerCodes?: unknown;
data?: { error?: unknown };
response?: {
status?: unknown;
headers?: Headers | Record<string, unknown>;
};
cause?: unknown;
original?: unknown;
originalError?: unknown;
details?: {
code?: unknown;
providerStatus?: unknown;
providerSubCode?: unknown;
providerCodes?: unknown;
};
innerHttpError?: { statusCode?: unknown };
};
/**
* Adapters sometimes wrap the provider SDK error before it reaches the
* durable outbox. Follow only the documented structured wrapper properties,
* with a small depth and cycle guard, so provider codes survive without ever
* classifying on free-form error text.
*/
function structuredProviderErrors(error: unknown): StructuredProviderError[] {
const records: StructuredProviderError[] = [];
const pending: Array<{ value: unknown; depth: number }> = [
{ value: error, depth: 0 },
];
const seen = new Set<object>();
while (pending.length > 0) {
const current = pending.shift();
if (
!current ||
current.depth > 4 ||
!current.value ||
typeof current.value !== "object" ||
seen.has(current.value)
) {
continue;
}
seen.add(current.value);
const record = current.value as StructuredProviderError;
records.push(record);
for (const nested of [
record.cause,
record.original,
record.originalError,
]) {
pending.push({ value: nested, depth: current.depth + 1 });
}
}
return records;
}
/**
* Classify a provider send failure by whether an external side effect could
* already have happened. Only ambiguous network failures stop automatic
* retry. Explicit provider rejections are safe to retry or repair without
* risking a duplicate message.
*/
export function classifyChatPublicationError(
error: unknown,
attempt: number,
): ChatPublicationErrorDisposition {
const values = structuredProviderErrors(error);
const firstNumber = (select: (value: StructuredProviderError) => unknown) =>
values
.map(select)
.find((candidate): candidate is number => typeof candidate === "number");
const names = values
.map((value) => value.name)
.filter((candidate): candidate is string => typeof candidate === "string");
const adapters = values
.map((value) => value.adapter)
.filter((candidate): candidate is string => typeof candidate === "string")
.map((candidate) => candidate.toLowerCase());
const codes = values
.map((value) => value.code)
.filter((candidate): candidate is string => typeof candidate === "string");
const platformCodes = values
.map((value) => value.data?.error)
.filter((candidate): candidate is string => typeof candidate === "string");
const detailsCodes = values
.map((value) => value.details?.code)
.filter((candidate): candidate is string => typeof candidate === "string");
const discordProviderCode =
values
.map((value) => value.code)
.find(
(candidate): candidate is number => typeof candidate === "number",
) ?? null;
const status = values
.flatMap((value) => [
value.status,
value.statusCode,
value.response?.status,
value.innerHttpError?.statusCode,
value.details?.providerStatus,
])
.find((candidate): candidate is number => typeof candidate === "number");
const teamsProviderCodes = values
.flatMap((value) => [
value.subCode,
value.details?.providerSubCode,
...(Array.isArray(value.providerCodes) ? value.providerCodes : []),
...(Array.isArray(value.details?.providerCodes)
? value.details.providerCodes
: []),
])
.filter((candidate): candidate is string => typeof candidate === "string")
.map((candidate) => candidate.toLowerCase());
const reason = text(error);
if (names.includes("PhotonError")) {
if (codes.includes("delivery_unknown")) return { kind: "delivery_unknown", reason };
if (codes.includes("credentials") || codes.includes("line_unavailable") || codes.includes("history_gap")) {
return { kind: "endpoint_attention", reason };
}
if (codes.includes("quota") || codes.includes("network") || codes.includes("attachment_not_ready")) {
return {
kind: "retry",
retryAfterMs: values
.map((value) => finitePositive(value.retryAfterMs))
.find((value) => value !== null) ?? Math.min(60_000, 1000 * 2 ** Math.min(attempt, 6)),
providerRateLimit: codes.includes("quota"),
reason,
};
}
return { kind: "failed", reason };
}
const responseHeaders = values
.map((value) => value.response?.headers)
.find((headers) => headers !== undefined);
const retryAfterHeaderMs = retryAfterHeaderMilliseconds(
responseHeader(responseHeaders, "retry-after"),
);
const rateLimitRemaining = responseHeader(
responseHeaders,
"x-ratelimit-remaining",
);
const rateLimitReset = positiveNumber(
responseHeader(responseHeaders, "x-ratelimit-reset"),
);
const githubRateLimit =
status === 403 &&
(retryAfterHeaderMs !== null ||
rateLimitRemaining === "0" ||
reason.toLowerCase().includes("secondary rate limit") ||
reason.toLowerCase().includes("rate limit exceeded"));
// Endpoint management owns the same lease as provider transport. Losing a
// short contention race is a definite local pre-transport outcome, so it is
// safe to retry automatically and must never be presented as an ambiguous
// provider delivery.
if (detailsCodes.includes("chat_endpoint_credentials_busy")) {
return { kind: "retry", retryAfterMs: 1_000, reason };
}
if (
names.includes("AdapterRateLimitError") ||
names.includes("RateLimitError") ||
codes.includes("RATE_LIMITED") ||
codes.includes("slack_webapi_rate_limited_error") ||
platformCodes.includes("ratelimited") ||
status === 429 ||
githubRateLimit
) {
const seconds = finitePositive(firstNumber((value) => value.retryAfter));
const milliseconds = finitePositive(
firstNumber((value) => value.retryAfterMs),
);
const telegramSeconds = finitePositive(
firstNumber((value) => value.retry_after),
);
const resetMilliseconds = rateLimitReset
? Math.max(1_000, rateLimitReset * 1_000 - Date.now())
: null;
const structuredSeconds = seconds ?? telegramSeconds;
return {
kind: "retry",
retryAfterMs: Math.min(
MAX_PROVIDER_RETRY_AFTER_MS,
milliseconds ??
(structuredSeconds !== null
? structuredSeconds * 1000
: retryAfterHeaderMs !== null
? retryAfterHeaderMs
: (resetMilliseconds ?? 2 ** Math.max(0, attempt) * 1000)),
),
providerRateLimit: true,
reason,
};
}
if (
// Discord distinguishes a destination the bot cannot access (50001) or
// cannot write to (50013) from invalid credentials and app-wide failures.
// The pinned adapter preserves the bounded numeric API code on its nested
// DiscordApiError. Quarantine only the affected channel/conversation;
// generic Discord 403s and every 401 remain endpoint-scoped below.
adapters.includes("discord") &&
status === 403 &&
(discordProviderCode === 50001 || discordProviderCode === 50013)
) {
return { kind: "resource_unavailable", reason };
}
if (
// Telegram uses 403 for destination-local conditions such as a user
// blocking the bot or removing it from a chat. Its adapter deliberately
// represents those responses as PermissionError while keeping an invalid
// bot token as AuthenticationError (401). Quarantine only that
// conversation/resource; the same bot may still serve every other chat.
adapters.includes("telegram") &&
names.includes("PermissionError") &&
codes.includes("PERMISSION_DENIED")
) {
return { kind: "resource_unavailable", reason };
}
if (
// A Teams app can remain healthy while Microsoft rejects writes to one
// conversation after the bot is blocked, uninstalled, or loses access.
// The pinned adapter preserves only bounded provider code tokens from the
// 403 response body; never use the free-form message for this distinction.
adapters.includes("teams") &&
names.includes("PermissionError") &&
codes.includes("PERMISSION_DENIED") &&
status === 403 &&
teamsProviderCodes.some((providerCode) =>
["messagewritesblocked", "forbiddenoperationexception"].includes(
providerCode,
),
)
) {
return { kind: "resource_unavailable", reason };
}
if (
names.includes("AuthenticationError") ||
names.includes("PermissionError") ||
codes.includes("AUTH_FAILED") ||
codes.includes("PERMISSION_DENIED") ||
status === 401 ||
status === 403 ||
[
"account_inactive",
"invalid_auth",
"missing_scope",
"not_authed",
"no_permission",
"token_revoked",
].some((platformCode) => platformCodes.includes(platformCode))
) {
return { kind: "endpoint_attention", reason };
}
if (
names.includes("ResourceNotFoundError") ||
codes.includes("NOT_FOUND") ||
status === 404 ||
status === 410 ||
[
"channel_not_found",
"channel_is_archived",
"is_archived",
"message_not_found",
"not_in_channel",
"thread_not_found",
].some((platformCode) => platformCodes.includes(platformCode)) ||
(names.includes("NetworkError") &&
reason.toLowerCase().includes("resource not found during"))
) {
return { kind: "resource_unavailable", reason };
}
if (
names.includes("ValidationError") ||
names.includes("NotImplementedError") ||
codes.includes("VALIDATION_ERROR") ||
codes.includes("NOT_IMPLEMENTED") ||
// Paperclip rejected the destination locally before opening a provider
// request, so delivery is definitively impossible rather than ambiguous.
codes.includes("CHAT_PROVIDER_PRETRANSPORT_REJECTED") ||
// Adapter-contract drift is detected during runtime construction, before
// any provider request can have been attempted.
codes.includes("CHAT_ADAPTER_COMPATIBILITY_ERROR")
) {
return { kind: "failed", reason };
}
// Slack WebClient platform errors represent a completed, structured API
// rejection even though the SDK does not expose an HTTP status. Unknown
// platform codes are therefore definite failures, not ambiguous delivery.
if (codes.includes("slack_webapi_platform_error")) {
return { kind: "failed", reason };
}
// Any remaining 4xx response is a definite provider rejection: the
// provider returned an HTTP response and did not accept the operation.
// Keep it out of delivery_unknown, which is reserved for requests whose
// external side effect cannot be determined. Provider-specific repairable
// cases (rate limits, auth, and missing destinations) were handled above.
if (status !== undefined && status >= 400 && status < 500) {
return { kind: "failed", reason };
}
// Adapter NetworkError and ordinary fetch/transport errors are ambiguous:
// the request may have reached the provider even when no response reached
// Paperclip. An operator must inspect the native conversation before replay.
return { kind: "delivery_unknown", reason };
}