06-reference

mostly metrics llm costs pl

2026-07-09·reference·source: Mostly Metrics·by CJ Gustafson
llm-costscogsrd-expensegross-marginsaas-metricscfo-mailbagai-accountingrule-of-40arrusage-based-billing

"How to Classify Your LLM Costs on the P&L" — @cjgustafson222

Why this is in the vault

The headline question — where do AI inference costs belong on the P&L — is not abstract. It determines reported gross margin for every SaaS company shipping customer-facing AI features, and it's the first accounting question a CFO asks when AI spend crosses a materiality threshold. The $180K/month example in Q2 is where RDCO's own toolchain is heading; the R&D vs COGS split applies today. phData engagements with enterprise clients will surface this question as those clients ship AI features to their own customers.

Issue contents

This week's CFO Mailbag features two operating finance leaders:

Questions covered:

  1. Adjusting Metrics Post RIF — severance and real estate exit costs distorting Rule of 40 for two quarters; adjust, footnote, or take the hit?
  2. Classifying LLM Costs on the P&L (headline) — $180K/month LLM API spend growing 20% MoM; currently in R&D but CRO argues customer-facing inference should be COGS
  3. Reporting Usage Based Contracts as ARR — how to frame usage-based contract value in ARR reporting
  4. Controller or FP&A? Sequencing the Finance Hire — which function to staff first as a scaling startup
  5. Monthly KPI Packets for Investors — format, frequency, and what to include

Note: Only Q1 is available from the email extract; Q2–Q5 answers are behind the Mostly Metrics paywall. The canonical article URL above contains the full issue.

The core argument

Q1 — Adjusting Metrics Post RIF (full content):

Both CFOs lean toward not creating an "adjusted" series. Vanna's argument: adjusting numbers damages credibility, and two quarters of noise is short enough to let the trend speak. Mitzi's approach: report as-is, then include a footnote showing what the metric looks like stripped of one-time severance and real estate exit costs — transparent without manufacturing an alternative GAAP.

The implicit principle: footnoting preserves investor trust better than headline-adjusted numbers, especially when the distortion has a clear endpoint.

Q2 — LLM Cost Classification (question only; answer paywalled):

The question frames the live accounting debate cleanly:

"Our engineering team is burning through $180K/month on LLM API costs, growing 20% MoM. Right now it sits in R&D, but our CRO is arguing the customer-facing AI features should hit COGS since they scale with usage. How are you thinking about where AI inference costs belong, and does the answer change how you talk about gross margin to the Street?"

The R&D vs COGS split hinges on whether the LLM call is part of the product delivery to the customer (COGS) or the development of the product (R&D). Customer-facing inference that scales per-transaction is a strong candidate for COGS reclassification. The gross margin implication is significant — at $180K/month COGS-classified, gross margin compression becomes a Board-level narrative item.

Mapping against Ray Data Co

The R&D→COGS reclassification question applies to RDCO's own AI API stack today. Claude API calls powering the always-on COO agent are internal tooling (R&D / G&A); ElevenLabs calls that produce client-deliverable audio are COGS. The boundary is the same one Q2 surfaces — does the LLM call land inside a product the customer receives, or inside RDCO's own operations? At current scale this is an internal bookkeeping question, but as RDCO productizes any AI-delivered service to phData or external clients, this classification becomes P&L-material.

For phData engagements: when a DSA is helping a client CFO understand AI ROI, the COGS vs R&D framing determines whether AI investment shows up as margin compression (bad for the gross margin story) or as operating leverage (the R&D narrative). Knowing which classification a client is using — and whether it's consistent with usage-scaling behavior — is a credibility move in CFO conversations.

The Q1 content (RIF adjustment discipline) is also directly applicable: RDCO should default to footnoting rather than headline-adjusting any one-time AI infrastructure costs when reporting to advisors or investors.

Related