06-reference

mostlymetrics audit ai bill for savings

2026-09-01·reference·source: Mostly Metrics·by CJ Gustafson
mostly-metricstoken-costsai-roiharness-thesisagent-cost-attributioncfo-role

How to Audit Your AI Bill for Savings — Mostly Metrics

Part 1 of a two-part CJ Gustafson series: a Claude-driven walkthrough for turning an opaque Anthropic invoice (four line items, no detail) plus the console usage export (workspace/API-key/model/token-count granularity) into a dollarized list of five leakage patterns — stale models, missing prompt caching, standard-rate jobs that could run on batch, over-provisioned models on simple tasks, and single-job cost concentration.

Why this is in the vault

A concrete, reusable prompt sequence for LLM spend auditing — worth keeping as a template RDCO could run against its own Anthropic usage the day cost visibility becomes actionable.

Mapping against Ray Data Co

RDCO doesn't have an Anthropic invoice to audit yet — Max-subscription flat pricing means no per-call cost line, per the standing API-cost-budget-controlled memory — but the five leakage categories CJ names map almost one-to-one onto open RDCO harness-cost questions that current dispatch practice doesn't check systematically: (1) stale-model drift — RDCO's delegation-model-effort-pairing memory sets Fable at high/xhigh for meaningfully large tasks, but nothing re-verifies that a dispatched sub-agent's model choice is still the cheapest one that clears the bar as new models ship; (2) missing caching — every newsletter-processing and deep-research sub-agent re-reads the same vault context (schema files, CLAUDE.md, memory) per invocation with no verified cache-hit discipline; (3) sync vs. batch — crons like acquisition-rescan, vault-health, and sync-contacts run on standard latency when none are time-sensitive, the textbook "fast lane for background work" leak; (4) over-provisioned model on simple work — the general-purpose/Explore agent split exists but isn't audited against actual task complexity; (5) single-job concentration — no current instrument would show which RDCO skill or cron consumes the plurality of token spend. The actionable takeaway isn't to build this today (no invoice, no motive) but to bank the five-category checklist and CJ's "confirm the two files agree before trusting any number" discipline as the audit method for whenever RDCO's Anthropic usage moves off flat pricing — this directly extends the instrumentation gap flagged in the 2026-08-04 CFO-token-cost-gate note ("no seat count, no PO, no renewal event... worth flagging as a future instrumentation need") and gives it an actual runnable procedure instead of just a stated need.

The core argument

CJ's method: pull the Anthropic console usage export (not just the invoice, which is stripped of workspace/model/rate-type detail) and hand both files to Claude with a staged prompt sequence — (1) reconcile the export total against the invoice total before trusting any downstream number, (2) have Claude rewrite the analyst's own question into a sharper one before running it, (3) run three specific checks (actual blended cost/million tokens vs. list price by model; spend share by category — regular input/output/cached/batch; concentration in the single biggest job), (4) convert findings into a short non-blaming Slack message for engineering, (5) turn the whole thing into a five-line recurring monthly checklist. He frames this as usable on any consumption-based bill (Snowflake, Datadog, AWS), not just Anthropic — general applicability to usage-priced infra, not an Anthropic-specific hack. Explicit claim throughout: none of the five fixes degrade output quality, only cost.

⚠️ Sponsorship

Brex-sponsored issue (known recurring sponsor, brex.com/metrics affiliate link, opening promo block pitching Brex corporate cards + AI-native expense workflows). The sponsorship is adjacent to but distinct from the article's content — Brex's own cross-platform AI-spend data was cited as source material in the prior 2026-08-04 CFO-token-cost-gate issue from this same sender, but this issue's tutorial content (Anthropic invoice + usage export) is CJ's own demo data, not Brex's. No new sponsor.

Related