fix(adapters): record unpriced CLI usage (#9505)

## Thinking Path

> - Paperclip is the open source control plane people use to manage
AI-agent companies
> - Budgets and spend telemetry are control-plane safety features, not
just reporting
> - Local Codex and Claude adapters can execute through either ACP or
their native CLI engines
> - The ACP lane records usage and reported cost, but CLI JSON output
often reports tokens without a price
> - The CLI lane was either losing per-run usage semantics or coercing
missing cost to zero, making real usage indistinguishable from a
genuinely free run
> - This pull request preserves CLI usage as per-run totals and records
token-bearing runs without a reported price as explicitly unpriced
ledger events
> - The benefit is accurate usage accounting and a visible pricing gap
instead of silently misleading zero-cost telemetry

## Linked Issues or Issue Description

Refs #9471
Refs #9230

**Bug description**

A `codex_local` run using the CLI engine can emit a final
`turn.completed` event with millions of input tokens and tens of
thousands of output tokens while the agent's spend ledger remains
indistinguishable from a true zero-usage, zero-cost run. Claude CLI
output has the same missing-price edge case.

**Expected behavior**

Token-bearing CLI runs should persist their usage. If the adapter
reports a price, the ledger should record it as reported; if the CLI
reports usage but no price, the ledger should explicitly mark the event
as unpriced rather than silently treating missing price data as a
reported `$0` cost.

**Reproduction shape**

1. Configure `codex_local` with `engine: cli`.
2. Run a task that produces a `turn.completed` usage payload.
3. Observe token usage in the run stream.
4. Before this change, missing price data is represented as ordinary
zero-cost spend and the CLI usage basis is not consistently propagated.

## What Changed

- Mark Codex and Claude native CLI usage totals as `per_run` and
propagate that basis through success and failure results.
- Stop coercing missing Claude CLI cost to `0`.
- Add `cost_status` to cost events with `reported` and `unpriced`
values, including an idempotent migration and shared validation/types.
- Persist token-bearing runs without a reported price as `unpriced`
ledger events while retaining zero cents until an authoritative price
exists.
- Add parser, execute-path, heartbeat-accounting, and cost-service
regression coverage for both local CLI adapters.
- Document the cost-status invariant and CLI accounting behavior.

## Verification

- `pnpm exec vitest run
packages/adapters/codex-local/src/server/parse.test.ts
packages/adapters/claude-local/src/server/parse.test.ts
server/src/__tests__/codex-local-execute.test.ts
server/src/__tests__/claude-local-execute.test.ts
server/src/__tests__/heartbeat-cost-accounting.test.ts
server/src/__tests__/costs-service.test.ts` — 6 files / 102 tests
passed.
- `pnpm --filter @paperclipai/shared typecheck`
- `pnpm --filter @paperclipai/db typecheck` — includes migration
numbering and safety checks.
- `pnpm --filter @paperclipai/adapter-codex-local typecheck`
- `pnpm --filter @paperclipai/adapter-claude-local typecheck`
- `pnpm --filter @paperclipai/server typecheck`

## Risks

- Existing cost rows default to `reported`, preserving current
interpretation; only new token-bearing events with absent cost are
marked `unpriced`.
- This change does not invent model pricing. Budget hard stops still
cannot charge an unknown amount, but operators and evals can now
distinguish missing pricing from a genuinely reported zero cost.
- Consumers that enumerate cost-event fields should tolerate the
additive `costStatus` field.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model `gpt-5.3-codex`, with repository tool use
and code execution; default reasoning mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip authored and GitHub committed 2026-07-13 20:44:38 -05:00
1 parent ce7dedf33d
commit efcce9cc8e
20 files changed
+122 -6

No files matched your search

@@ -1067,7 +1067,7 @@ export async function execute(ctx: AdapterExecutionContext): Promise<AdapterExec
biller: isBedrockAuth(effectiveEnv) ? "aws_bedrock" : "anthropic",
model: parsedStream.model || asString(parsed.model, model),
billingType,
costUsd: parsedStream.costUsd ?? asNumber(parsed.total_cost_usd, 0),
costUsd: parsedStream.costUsd,
resultJson: mergedResultJson,
summary: parsedStream.summary || asString(parsed.result, ""),
clearSession:
@@ -420,6 +420,6 @@ describe("parseClaudeStreamJson usage extraction", () => {
outputTokens: 1_800,
cachedInputTokens: 20,
});
expect(parsed.usageBasis).toBeNull();
expect(parsed.usageBasis).toBe("per_run");
});
});
@@ -111,7 +111,7 @@ export function parseClaudeStreamJson(stdout: string) {
usage,
// modelUsage covers exactly this CLI invocation, so mark it per-run to
// keep the server from applying its session-cumulative delta heuristic.
usageBasis: modelUsageTotals ? ("per_run" as const) : null,
usageBasis: "per_run" as const,
summary,
resultJson: finalResult,
};
@@ -976,6 +976,7 @@ export async function execute(ctx: AdapterExecutionContext): Promise<AdapterExec
errorCode: "codex_output_inactivity_monitor",
errorFamily: null,
usage: attempt.parsed.usage,
usageBasis: attempt.parsed.usageBasis,
sessionId: null,
sessionParams: null,
sessionDisplayId: null,
@@ -1074,6 +1075,7 @@ export async function execute(ctx: AdapterExecutionContext): Promise<AdapterExec
errorFamily,
retryNotBefore: transientRetryNotBefore ? transientRetryNotBefore.toISOString() : null,
usage: attempt.parsed.usage,
usageBasis: attempt.parsed.usageBasis,
sessionId: resolvedSessionId,
sessionParams: resolvedSessionParams,
sessionDisplayId: resolvedSessionId,
@@ -30,6 +30,7 @@ describe("parseCodexJsonl", () => {
cachedInputTokens: 2,
outputTokens: 4,
},
usageBasis: "per_run",
errorMessage: "resume failed",
});
});
@@ -63,6 +64,7 @@ describe("parseCodexJsonl", () => {
cachedInputTokens: 2,
outputTokens: 4,
},
usageBasis: "per_run",
errorMessage: null,
});
});
@@ -70,6 +70,7 @@ export function parseCodexJsonl(stdout: string) {
sessionId,
summary: finalMessage?.trim() ?? "",
usage,
usageBasis: "per_run" as const,
errorMessage,
};
}
@@ -0,0 +1 @@
ALTER TABLE "cost_events" ADD COLUMN IF NOT EXISTS "cost_status" text DEFAULT 'reported' NOT NULL;
@@ -1016,6 +1016,13 @@
"when": 1783822632557,
"tag": "0146_routine_activity_gate",
"breakpoints": true
},
{
"idx": 147,
"version": "7",
"when": 1783953514660,
"tag": "0147_cost_event_status",
"breakpoints": true
}
]
}
+1
View File
@@ -20,6 +20,7 @@ export const costEvents = pgTable(
provider: text("provider").notNull(),
biller: text("biller").notNull().default("unknown"),
billingType: text("billing_type").notNull().default("unknown"),
costStatus: text("cost_status").notNull().default("reported"),
model: text("model").notNull(),
inputTokens: integer("input_tokens").notNull().default(0),
cachedInputTokens: integer("cached_input_tokens").notNull().default(0),
+3
View File
@@ -687,6 +687,9 @@ export const BILLING_TYPES = [
] as const;
export type BillingType = (typeof BILLING_TYPES)[number];
export const COST_STATUSES = ["reported", "unpriced"] as const;
export type CostStatus = (typeof COST_STATUSES)[number];
export const FINANCE_EVENT_KINDS = [
"inference_charge",
"platform_fee",
+1
View File
@@ -373,6 +373,7 @@ export {
type SecretScope,
type StorageProvider,
type BillingType,
type CostStatus,
type FinanceEventKind,
type FinanceDirection,
type FinanceUnit,
+2 -1
View File
@@ -1,4 +1,4 @@
import type { BillingType } from "../constants.js";
import type { BillingType, CostStatus } from "../constants.js";
export interface CostEvent {
id: string;
@@ -12,6 +12,7 @@ export interface CostEvent {
provider: string;
biller: string;
billingType: BillingType;
costStatus: CostStatus;
model: string;
inputTokens: number;
cachedInputTokens: number;
+2 -1
View File
@@ -1,5 +1,5 @@
import { z } from "zod";
import { BILLING_TYPES } from "../constants.js";
import { BILLING_TYPES, COST_STATUSES } from "../constants.js";
export const createCostEventSchema = z.object({
agentId: z.string().uuid(),
@@ -11,6 +11,7 @@ export const createCostEventSchema = z.object({
provider: z.string().min(1),
biller: z.string().min(1).optional(),
billingType: z.enum(BILLING_TYPES).optional().default("unknown"),
costStatus: z.enum(COST_STATUSES).optional().default("reported"),
model: z.string().min(1),
inputTokens: z.number().int().nonnegative().optional().default(0),
cachedInputTokens: z.number().int().nonnegative().optional().default(0),